Neural Network Model Training Method, Audio Background Music Method and Related Devices
By training the neural network model to extract the deep features of audio and background music, the inefficiency problem caused by manual intervention in the long audio soundtrack is solved, and automated and accurate background music recommendations are achieved.
Patent Information
- Application Number
- CN202211699711.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-12-28
AI Technical Summary
The existing long audio soundtracking process requires manual participation, resulting in inefficiency and subjective results, making it difficult for users to quickly find the right background music.
By training the neural network model, the deep features of audio and background music are extracted using the convolutional structure and fully connected layer, the adaptation degree is calculated using the sigmoid layer, and the model parameters are adjusted until converge, so as to achieve automated soundtracks.
It improves the efficiency of the adaptability judgment between audio and background music, reduces manual intervention, and improves the automation and accuracy of the soundtrack.
Smart Images

Figure CN116013358B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of audio processing, and particularly to a neural network model training method, an audio scoring method and related devices. Background Art
[0002] With the development of the times, people are gradually getting rid of reading books. Especially for those who love novels, since reading novels takes a lot of time, in order to maintain their hobbies during the breaks in life, people will choose to listen to books. When driving, during a short lunch break, etc., long audio is a good choice for users. Long audio includes stories, novels, cross talks, interviews, etc., with rich and diverse forms, and is deeply loved by users. For a piece of long audio, scoring is an important factor in attracting listeners and assisting in expression. Although the human voice is the main content of long audio, scoring is often an important guide for perfecting the description of the scene and giving people room for imagination. Adding appropriate background music to long audio can make the audio content more colorful and also make listeners more involved and immersive. Currently, most long audio is scored manually, or a certain selection of background music is given for users to independently choose the background music for scoring.
[0003] The existing long audio scoring is often selected by professional production personnel according to the content and mood of the long audio, or a certain category of background music is provided for users to choose. No matter which scheme is used, it requires manual participation, and the scoring result is relatively subjective, that is, the scoring result of a long audio that a certain listener likes may not be liked by other listeners. Moreover, whether it is to select the favorite background music from a large amount of background music or to score a large number of long audio works, it is very time-consuming, resulting in slow manual scoring efficiency. Summary of the Invention
[0004] The embodiments of the present application provide a neural network model training method, an audio scoring method and related devices, which can be used to produce audio scoring works with better adaptability.
[0005] In the first aspect of the embodiments of the present application, a neural network model training method is provided. The method is applied to a computer device, and the method includes:
[0006] Obtain multiple groups of training samples, each group of the training samples including a fusion feature vector corresponding to multiple dimensional features of an audio sample and a fusion feature vector corresponding to multiple dimensional features of a background music sample, and a true label used to represent the degree of adaptation between the audio sample and the background music sample;
[0007] Obtain an initial neural network model, the initial neural network model including a convolutional structure, a fully connected layer connected to the convolutional structure, and a sigmoid layer connected to the fully connected layer;
[0008] Input each group of the training samples into the initial neural network model, so that the initial neural network model performs the following operations:
[0009] Use the convolutional structure to extract features from the fused feature vector of the audio sample to obtain the deep features of the audio sample, and extract features from the fused feature vector of the background music sample to obtain the deep features of the background music sample;
[0010] Use the fully connected layer to splice the deep features of the audio sample and the deep features of the background music sample to obtain deep fusion features;
[0011] Use the sigmoid layer to calculate the prediction label for the deep fusion features;
[0012] Construct a loss function according to the true label and the prediction label, adjust the model parameters of the initial neural network model according to the loss value of the loss function, and stop training until the loss function satisfies the convergence condition to obtain the target neural network model.
[0013] The second aspect of the embodiments of the present application provides an audio background music matching method, which is applied to a computer device. The method includes:
[0014] Obtain the fused feature vectors corresponding to the multi-dimensional features of the target audio to be background music matched, and obtain the fused feature vectors corresponding to the multi-dimensional features of each alternative background music;
[0015] Input the fused feature vector of the target audio into the pre-trained target neural network model to obtain the audio feature vector output by the target neural network model;
[0016] Input the fused feature vector of the alternative background music into the target neural network model to obtain the background music feature vector output by the target neural network model;
[0017] Calculate the distance between the audio feature vector and the background music feature vector of each alternative background music, and use the alternative background music corresponding to the background music feature vector closest to the audio feature vector as the background music of the target audio.
[0018] The third aspect of the embodiments of the present application provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the method of the first aspect or the second aspect described above.
[0019] A fourth aspect of the embodiments of the present application provides a computer storage medium storing instructions that, when executed on a computer, cause the computer to execute the methods in the foregoing first aspect or second aspect.
[0020] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages:
[0021] The convolutional structure of the initial neural network model extracts features from the fused feature vectors of audio samples to obtain the deep features of audio samples, and extracts features from the fused feature vectors of background music samples to obtain the deep features of background music samples. The fully connected layer is used to splice the deep features of audio samples and the deep features of background music samples to obtain deep fusion features. The sigmoid layer is used to calculate the prediction labels for the deep fusion features. Furthermore, the model parameters can be adjusted according to the prediction labels and the true labels until the convergence condition is met and the model training stops. Since the model is trained with various audio samples and background music samples with different degrees of adaptation, during the training process, it can learn the similarity of the feature vectors between highly adapted audio and background music, and can also learn the differences in the feature vectors between poorly adapted and non - adapted audio and background music. Therefore, when the model is applied, it can amplify the adaptability of the features between audio and background music, making the adapted audio and background music closer in features, while the non - adapted audio and background music are farther apart in features, thus making it easier for users to judge whether the background music is suitable for the audio and improving the efficiency of audio scoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a schematic flowchart of a neural network model training method in an embodiment of the present application;
[0023] Figure 2 It is another schematic flowchart of an audio scoring method in an embodiment of the present application;
[0024] Figure 3 It is a schematic diagram of an application scenario of a target neural network model in an embodiment of the present application;
[0025] Figure 4 It is another schematic diagram of the structure of a computer device in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The embodiments of the present application provide a neural network model training method, an audio scoring method and related devices, which can be used to produce audio scoring works with better adaptability.
[0027] Please refer to Figure 1 , an embodiment of the neural network model training method in the embodiments of the present application includes:
[0028] 101. Obtain multiple groups of training samples. Each group of the training samples includes a fusion feature vector corresponding to multiple-dimensional features of an audio sample and a fusion feature vector corresponding to multiple-dimensional features of a background music sample, as well as a ground truth label for representing the degree of adaptation between the audio sample and the background music sample.
[0029] The method of this embodiment can be applied to a computer device, which can be a server, a terminal, or other computer devices capable of performing data processing. When the computer device is a terminal, it can be a personal computer (PC), a desktop computer, or other terminal devices; when the computer device is a server, it can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud databases, cloud computing, and big data and artificial intelligence platforms.
[0030] Before model training, the labels of each group of training samples can be annotated, that is, a ground truth label is added to each group of training samples. This ground truth label represents the degree of adaptation between the audio sample and the background music sample in each group of training samples. Among them, the ground truth label can be any form of text that can be recognized by the computer device and can be used for loss function calculation.
[0031] For example, for an unadapted combination in the training samples, its ground truth label can be recorded as 0, while for an adapted combination, its ground truth label can be recorded as 1. In subsequent steps, the loss function can be calculated according to the ground truth labels of each training sample.
[0032] The ground truth labels of multiple groups of training samples can include various degrees of adaptation between the audio sample and the background music sample, such as taking values between 0 and 1. The ground truth labels corresponding to various degrees of adaptation can include multiple values such as 0, 0.1, 0.2, 0.3, 0.4... 0.7, 0.8, 0.9, 1, etc., so that the neural network model to be trained can learn the features of training samples with various different degrees of adaptation.
[0033] Among them, the audio sample can be audio data such as human voice audio that needs to be accompanied by music. The multiple-dimensional features of the audio sample refer to the auditory experience and effects brought by the audio sample to the user. For example, it can include multiple-dimensional features such as the rhythm, reading speed, and emotion of the audio sample. The computer device can obtain a fusion feature vector of the multiple-dimensional features of the audio sample. The fusion feature vector is a data representation of the multiple-dimensional features of the audio sample, that is, it represents the fusion status of the multiple-dimensional features of the audio sample in the form of data.
[0034] Similarly, the multiple-dimensional features of the background music sample refer to the characteristics of the background music and the auditory experience and effects brought to the user. For example, they may include features of multiple dimensions such as the genre, soothing degree, and emotion of the background music. The computer device can obtain the fusion feature vector of the multiple-dimensional features of the background music. The fusion feature vector is the data representation of the multiple-dimensional features of the background music, that is, it represents the fusion status of the multiple-dimensional features of the background music in the form of data.
[0035] 102. Input each group of the training samples into the initial neural network model so that the initial neural network model performs the following operations: extract the deep features of the audio sample by performing feature extraction on the fusion feature vector of the audio sample, and extract the deep features of the background music sample by performing feature extraction on the fusion feature vector of the background music sample, splice the deep features of the audio sample and the deep features of the background music sample to obtain the deep fusion feature, and calculate the predicted label for the deep fusion feature;
[0036] The computer device can obtain the initial neural network model. The initial neural network model includes a convolutional structure, a fully connected layer connected to the convolutional structure, and a sigmoid layer connected to the fully connected layer. The convolutional structure of the initial neural network model is used to perform deep feature extraction on the fusion feature vector of the audio sample and the fusion feature vector of the background music sample. The fully connected layer is used to splice the deep features of the audio sample extracted by the convolutional structure and the deep features of the background music sample. The sigmoid layer is used to calculate the predicted probability according to the sigmoid function for the spliced features output by the fully connected layer.
[0037] The fusion feature vectors of the audio samples and the background music samples of multiple groups of training samples are input into the initial neural network model. The initial neural network model performs iterative training on multiple groups of training samples. The convolutional structure of the initial neural network model can further extract the deep features of the fusion feature vector of the audio sample and the deep features of the fusion feature vector of the background music sample respectively. The deep features of the audio sample and the deep features of the background music sample are input into the fully connected layer for feature splicing. Then the fully connected layer outputs the deep fusion feature. The sigmoid layer calculates the predicted label according to the sigmoid function for the deep fusion feature output by the fully connected layer. This predicted label can represent the probability that the audio sample and the background music sample are adapted to each other.
[0038] 103. Construct a loss function according to the true label and the predicted label, adjust the model parameters of the initial neural network model according to the loss value of the loss function, and stop training until the loss function meets the convergence condition to obtain the target neural network model;
[0039] After the feature extraction and adaptability calculation are completed for each group of training samples, a loss function can be constructed based on the predicted labels and true labels corresponding to each group of training samples output by the initial neural network model, and the loss value of this loss function can be calculated. Then, according to this loss value, the model parameters of the initial neural network model are adjusted until the training stops when the loss function meets the convergence condition, and a target neural network model is obtained. Among them, the loss function meeting the convergence condition can be that its loss value has been stable within a numerical range during multiple rounds of model training. For example, it fluctuates slightly around a preset numerical level, and the fluctuation range can be a preset numerical amplitude, or its loss value decreases to a certain numerical range, which is not limited here.
[0040] Since the training samples can include audio samples and background music samples with various different degrees of adaptability, through the model training process, the target neural network model can learn the similarity of the feature vectors between highly adaptable audio and background music, and can also learn the differences in the feature vectors between lowly adaptable and non - adaptable audio and background music. Thus, when extracting the feature vectors of audio and background music subsequently, the extracted audio feature vectors and background music feature vectors are closer in distance when the audio and background music are adaptable, and farther in distance when the audio and background music are not adaptable. This is equivalent to amplifying the adaptability between the audio and the background music, making the adaptability between the audio and the background music more obvious and facilitating the adaptability judgment.
[0041] In this embodiment, the convolutional structure of the initial neural network model extracts the deep features of the audio samples from the fused feature vectors of the audio samples, and extracts the deep features of the background music samples from the fused feature vectors of the background music samples. The fully - connected layer is used to splice the deep features of the audio samples and the deep features of the background music samples to obtain the deep fusion features. The sigmoid layer is used to calculate the predicted labels from the deep fusion features, and then the model parameters can be adjusted according to the predicted labels and the true labels until the model training stops when the convergence condition is met. Since the model is trained with audio samples and background music samples with various different degrees of adaptability, during the training process, it can learn the similarity of the feature vectors between highly adaptable audio and background music, and can also learn the differences in the feature vectors between lowly adaptable and non - adaptable audio and background music. Furthermore, when the model is applied, it can amplify the adaptability of the features between the audio and the background music, making the adaptable audio and background music closer in features, while the non - adaptable audio and background music are farther apart in features, thus making it more convenient for users to judge whether the background music is adaptable to the audio and improving the efficiency of audio scoring.
[0042] Based on Figure 1In the embodiment shown, in a preferred implementation manner, a fused feature vector corresponding to multiple-dimensional features of an audio sample is obtained. Specifically, multiple pre-trained target audio feature extraction models are obtained, where the multiple target audio feature extraction models extract feature vectors for different-dimensional features of the audio sample, the spectral features of the audio sample are extracted, and the spectral features of the audio sample are input into the multiple target audio feature extraction models. Each target audio feature extraction model outputs a feature vector of one-dimensional feature of the audio sample, and feature vectors of multiple-dimensional features of the audio sample are obtained, that is, multiple feature vectors of the audio sample correspond to different feature dimensions of the audio sample. Further, the feature vectors of multiple-dimensional features of the audio sample are fused to obtain a fused feature vector corresponding to multiple-dimensional features of the audio sample.
[0043] Among them, the training steps of each target audio feature extraction model include:
[0044] Obtain multiple groups of audio training samples, where each group of audio training samples includes an audio sample and a feature label for representing one-dimensional feature of the audio sample;
[0045] Obtain an initial audio feature extraction model, where the initial audio feature extraction model includes a convolutional structure and a Softmax layer connected to the convolutional structure;
[0046] Extract the spectral features of the audio sample, input the spectral features of the audio sample into the initial audio feature extraction model, and the initial audio feature extraction model uses the convolutional structure to extract feature vectors from the spectral features of the audio sample and uses the Softmax layer to calculate the predicted feature label for the feature vectors;
[0047] When the relationship between the predicted feature label and the true feature label satisfies the convergence condition, stop training to obtain the target audio feature extraction model.
[0048] The multiple-dimensional features of the audio sample can be features in dimensions such as prosody, reading speed, and emotion. The feature label of each dimension feature can be used to describe the specific content of this dimension feature. For example, for prosody or reading speed, grading can be performed. For example, the prosody of the audio sample is divided into 5 levels according to the number of occurrences of stress, and the reading speed of the human voice audio is divided into 5 levels from slow to fast or from fast to slow; for the emotion feature included in the audio sample, the emotion feature can be divided into 5 dimensions: normal, excited, angry, happy, and sad. The feature vector of the audio sample corresponds to one level in one dimension feature of the audio sample, such as corresponding to one level in the prosody feature of the audio sample or corresponding to one level in the reading speed of the audio sample. Therefore, the feature label can be marked for each dimension feature of each audio sample according to the specific content and level of the above multiple dimension features to represent the specific level of the audio sample in any dimension feature.
[0049] After multiple groups of audio training samples are input into the initial audio feature extraction model, the initial audio feature extraction model learns the specific content and characteristics of one-dimensional features of each group of training samples, continuously performs iterative training based on the audio training samples to update the model parameters, and stops training when the relationship between the output feature labels and the pre-annotated feature labels meets the convergence condition, obtaining the target audio feature extraction model. Since each target audio feature extraction model learns the specific content and characteristics of one-dimensional features of audio samples, each target audio feature extraction model can be used to extract the feature vectors of one-dimensional features of audio samples.
[0050] Among them, the relationship between the feature labels output by the initial audio feature extraction model and the pre-annotated feature labels can be represented by a loss function, that is, a loss function is constructed based on the output feature labels and the pre-annotated feature labels. When the loss value of this loss function reaches a certain preset numerical range, or the loss value tends to be stable, it is determined that the model training meets the convergence condition. The initial audio feature extraction model can specifically be a convolutional neural network model or other feature extraction models.
[0051] Based on Figure 1 In a preferred implementation manner of the embodiment shown, a fusion feature vector corresponding to multiple-dimensional features of a background music sample is obtained. The specific method may be to obtain multiple pre-trained target background music feature extraction models, where the multiple target background music feature extraction models extract feature vectors for different-dimensional features of the background music, extract the spectral features of the background music sample, input the spectral features of the background music sample into the multiple target background music feature extraction models, and each target background music feature extraction model outputs the feature vector of one-dimensional features of the background music sample, obtaining the feature vectors of multiple-dimensional features of the background music sample, that is, the multiple feature vectors of the background music sample correspond to different feature dimensions of the background music sample, and further fusing the feature vectors of multiple-dimensional features of the background music sample to obtain the fusion feature vector corresponding to multiple-dimensional features of the background music sample.
[0052] Among them, the training steps of each target background music feature extraction model include:
[0053] Obtain multiple groups of background music training samples, where each group of background music training samples includes a background music sample and a feature label used to represent one-dimensional features of the background music sample;
[0054] Extract the spectral features of the background music samples, and input the spectral features of the background music samples into the initial background music feature extraction model. The initial background music feature extraction model uses a convolutional structure to extract features from the spectral features of the background music samples to obtain feature vectors, and uses a Softmax layer to calculate the predicted feature labels for the feature vectors;
[0055] When the relationship between the predicted feature labels and the true feature labels meets the convergence condition, stop training to obtain the target background music feature extraction model.
[0056] The multi-dimensional features of the background music samples can be features in dimensions such as genre, soothing degree, emotion, etc. The feature labels for each dimension feature can be used to describe the specific content of that dimension feature. For example, for the genre feature of the background music, feature labels can be used to indicate that the genre feature of the background music is pop, classical, electronic music, rock, other types, etc.; for the soothing degree feature of the background music, feature labels can be used to represent different soothing degree levels of the background music, such as using 5 feature labels to represent 5 different degrees of soothing of the background music; for the emotion feature of the background music, feature labels can be used to indicate that the emotion feature of the background music is normal, excited, angry, joyful, sad, etc. The feature vector of the background music sample corresponds to a level in a dimension feature of the background music sample, such as corresponding to the "pop" level in the genre feature of the background music sample, or corresponding to the "joyful" level in the emotion feature of the background music sample. Therefore, the feature labels can be annotated for each dimension feature of each background music sample according to the specific content and level of the above multi-dimensional features to characterize the specific content of the background music sample in any dimension feature.
[0057] After multiple groups of background music training samples are input into the initial background music feature extraction model, the initial background music feature extraction model learns the specific content and characteristics of a dimension feature of each group of training samples, and continuously performs iterative training according to the background music training samples to update the model parameters. When the relationship between the output feature labels and the pre-annotated feature labels meets the convergence condition, stop training to obtain the target background music feature extraction model. Since each target background music feature extraction model learns the specific content and characteristics of a dimension feature of the background music sample, each target background music feature extraction model can be used to extract the feature vector of a dimension feature of the background music sample.
[0058] Among them, the relationship between the feature labels output by the initial background music feature extraction model and the pre-annotated feature labels can be represented by a loss function, that is, a loss function is constructed based on the output feature labels and the pre-annotated feature labels. When the loss value of this loss function reaches a certain preset numerical range, or the loss value tends to be stable, it is determined that the model training meets the convergence condition. The initial background music feature extraction model can specifically be a convolutional neural network model or other feature extraction models.
[0059] Based on Figure 1 In the embodiment shown, in a preferred implementation manner, the initial neural network model can be a siamese network model. The siamese network model includes a first convolutional neural network and a second convolutional neural network. The first convolutional neural network and the second convolutional neural network are respectively connected to a fully connected layer, and the fully connected layer is connected to a sigmoid layer. Therefore, when training the siamese network model, the fused feature vectors corresponding to the audio samples of each group of training samples can be input into the first convolutional neural network of the siamese network model. The first convolutional neural network extracts features from the fused feature vectors of the audio samples to obtain the deep features of the audio samples, and the fused feature vectors corresponding to the background music samples are input into the second convolutional neural network of the siamese network model. The second convolutional neural network extracts features from the fused feature vectors of the background music samples to obtain the deep features of the background music samples.
[0060] The outputs of the first convolutional neural network and the second convolutional neural network are transmitted to the fully connected layer. The fully connected layer performs feature splicing on the output of the first convolutional neural network and the output of the second convolutional neural network to obtain deep fused features. The features output by the fully connected layer are transmitted to the sigmoid layer. The sigmoid layer maps the output of the fully connected layer to a target value within a preset numerical range. This target value is used as the prediction label for the adaptability of the audio and background music in the training samples by the siamese network model. Therefore, a loss function can be constructed based on the target value and the true label, and the model parameters of the siamese network model are adjusted according to the loss value of the loss function until the training stops when the loss function meets the convergence condition, and the target neural network model is obtained. The target neural network model is the siamese network model that has completed the above training process.
[0061] Among them, the preset numerical range based on which the sigmoid layer maps the output of the fully connected layer to a target value can be, for example, a numerical range between 0 and 1, or it can also be other numerical ranges. The loss function can specifically be a cross-entropy loss function, and its expression can be represented as:
[0062] L(y,f(x))=-[ylog(f(x))+(1-y)log(1-f(x))]
[0063] Wherein, y represents the true label marked by the training sample. For example, when the audio sample is adapted to the background music sample, its value can be 1, and when it is not adapted, its value can be 0; f(x) is the output of the sigmoid layer, that is, the sigmoid layer maps the output of the fully connected layer to a target value.
[0064] It should be noted that this loss function can also be other types of loss functions, such as the perceptron loss function, the square loss function, and so on.
[0065] The above describes the training steps of the target neural network model. The following will further describe the application process of the target neural network model.
[0066] The following will be based on the foregoing Figure 1 On the basis of the shown embodiments, the embodiments of the present application will be further described in detail. Based on the above Figure 1 On the basis of the shown embodiments, the embodiments of the present application also propose an audio scoring method. Please refer to Figure 2 , and the steps of the audio scoring method include:
[0067] 201. Obtain the fusion feature vector corresponding to the multiple-dimensional features of the target audio to be scored, and obtain the fusion feature vector corresponding to the multiple-dimensional features of each alternative background music;
[0068] The method of this embodiment can be applied to a computer device, which can be a server, a terminal, or other computer devices capable of performing data processing. When the computer device is a terminal, it can be a personal computer (PC), a desktop computer, or other terminal devices; when the computer device is a server, it can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud databases, cloud computing, and big data and artificial intelligence platforms.
[0069] The target audio can be audio data such as human voice audio that needs to be scored. The computer device provides multiple background musics, and the user can select a background music from them and use the computer device to score the target audio. The multiple-dimensional features of the target audio refer to the auditory experience and effects brought by the target audio to the user, such as features in multiple dimensions including the rhythm, reading speed, emotion, etc. of the target audio. The computer device can obtain the fusion feature vector of the multiple-dimensional features of the target audio, and the fusion feature vector is the data representation of the multiple-dimensional features of the target audio, that is, it represents the fusion status of the multiple-dimensional features of the target audio in the form of data.
[0070] Similarly, the multiple-dimensional features of the alternative background music also refer to the characteristics of the alternative background music and the auditory experience and effects brought to the user. For example, they may include features in multiple dimensions such as the genre, soothing degree, and emotion of the alternative background music. The computer device can obtain the fusion feature vector of the multiple-dimensional features of the alternative background music. The fusion feature vector is the data representation of the multiple-dimensional features of the alternative background music, that is, it represents the fusion status of the multiple-dimensional features of the alternative background music in the form of data.
[0071] 202. Input the fusion feature vector of the target audio into the pre-trained target neural network model to obtain the audio feature vector output by the target neural network model;
[0072] 203. Input the fusion feature vector of the alternative background music into the target neural network model to obtain the background music feature vector output by the target neural network model;
[0073] In this embodiment, the target neural network model can be the target neural network model that has completed the model training steps of the neural network model training method shown in Figure 1 That is to say, the computer device can deploy the target neural network model that has completed the model training steps of the neural network model training method shown in the foregoing Figure 1 to use the target neural network model to assist the user in audio scoring.
[0074] After obtaining the fusion feature vector of the target audio, the fusion feature vector of the target audio can be input into the pre-trained target neural network model, and the fusion feature vector of the alternative background music can be input into the target neural network model to obtain the audio feature vector corresponding to the target audio and the background music feature vector corresponding to the alternative background music.
[0075] During the training process of the target neural network model, combinations of suitable audio and background music are trained, as well as combinations of unsuitable audio and background music are trained, so that the distance between the audio feature vector and the background music feature vector extracted by the target neural network model for the combination of suitable audio and background music is closer, and the distance between the audio feature vector and the background music feature vector extracted by the target neural network model for the combination of unsuitable audio and background music is farther, thus making it more convenient to determine whether the target audio and the alternative background music are suitable according to the distance between the audio feature vector and the background music feature vector.
[0076] 204. Calculate the distance between the audio feature vector and the background music feature vector of each alternative background music, and use the alternative background music corresponding to the background music feature vector closest to the audio feature vector as the background music of the target audio;
[0077] The computer device can calculate the distance between the audio feature vector of the target audio and the background music feature vector of each alternative background music using the distance calculation method of feature vectors. This distance represents the degree of adaptation between the target audio and the alternative background music, that is, the smaller the distance, the higher the degree of adaptation; the larger the distance, the lower the degree of adaptation. Therefore, the alternative background music corresponding to the background music feature vector with the closest distance to the audio feature vector of the target audio can be used as the background music of the target audio.
[0078] In this embodiment, the computer device obtains the fused feature vector of multiple-dimensional features of the target audio and the fused feature vector of multiple-dimensional features of the alternative background music. The fused feature vector of the target audio and the fused feature vector of the alternative background music are respectively input into the target neural network model to obtain the audio feature vector and the background music feature vector output by the target neural network model. Calculate the distance between the audio feature vector and the background music feature vector of each alternative background music, and use the alternative background music corresponding to the background music feature vector with the closest distance to the audio feature vector as the background music of the target audio. Since the audio feature vector fuses multiple-dimensional features of the audio, and the background music feature vector fuses multiple-dimensional features of the background music, the audio and background music corresponding to the closest audio feature vector and background music feature vector are more adapted in style and auditory effect, making the music scoring result more acceptable and recognized by most listeners. At the same time, it realizes automatic music scoring for the target audio, eliminating the need for manual music scoring, reducing the effort and time, and also improving the efficiency of audio music scoring.
[0079] Based on Figure 2 In the embodiment shown, in a preferred implementation manner, feature fusion is performed on the feature vectors of multiple-dimensional features of the audio sample, or feature fusion is performed on the feature vectors of multiple-dimensional features of the background music sample. The method can be feature splicing, that is, use the feature splicing method to splice the feature vectors of multiple-dimensional features of the background music sample or the feature vectors of multiple-dimensional features of the audio sample to obtain the fused feature vector. In addition, feature fusion can also be achieved by adding the feature vectors of multiple-dimensional features of the background music sample or the feature vectors of multiple-dimensional features of the audio sample and taking the average. The specific manner of the above feature fusion is not limited in this embodiment.
[0080] Based on Figure 2 In the embodiment shown, in a preferred implementation manner, when calculating the distance between the audio feature vector of the target audio and the background music feature vector of each alternative background music, the method of calculating this distance can be to calculate the cosine distance, Euclidean distance, Hamming distance, etc. between the two. Any distance metric method for feature vectors can be applied in the embodiments of this application.
[0081] Combined with Figure 1 and Figure 2 in Figure 3 In an application scenario of an embodiment of the present application shown, three sub-convolutional networks corresponding to the long audio to be accompanied by music (i.e., the above-mentioned target audio feature extraction model) are respectively deployed to extract the feature vectors of the long audio, and three sub-convolutional networks corresponding to the background music (i.e., the above-mentioned target background music feature extraction model) are deployed to extract the feature vectors of the background music.
[0082] Extract the spectral features of the long audio, and input the spectral features of the long audio into its corresponding three sub-convolutional networks respectively. Each sub-convolutional network outputs the feature vectors of one-dimensional features of the long audio, and the feature vectors of multiple-dimensional features of the long audio are obtained, that is, embedding_1, embedding_2, and embedding_3. The feature vectors of multiple-dimensional features of the long audio are fused to obtain the fused feature vectors corresponding to multiple-dimensional features of the long audio.
[0083] Extract the spectral features of the background music, and input the spectral features of the background music into its corresponding three sub-convolutional networks. Each sub-convolutional network outputs the feature vectors of one-dimensional features of the background music, and the feature vectors of multiple-dimensional features of the background music are obtained, that is, embedding_a, embedding_b, and embedding_c. The feature vectors of multiple-dimensional features of the background music are fused to obtain the fused feature vectors corresponding to multiple-dimensional features of the background music.
[0084] In addition, the computer device can deploy a siamese network structure, which includes a convolutional module branch m and a convolutional module branch n. The fused feature vector of the long audio can be input into the convolutional module branch m, and the convolutional module branch m extracts the deep features of the long audio; the fused feature vector of the background music can be input into the convolutional module branch n, and the convolutional module branch n extracts the deep features of the background music. The deep features of the long audio and the deep features of the background music are feature-stitched through a fully connected layer (not shown in the figure) to obtain the deep fusion features, and the sigmoid layer calculates the deep fusion features according to the sigmoid function, that is, it is transformed into a binary classification problem of whether the long audio and the background music are compatible. Therefore, according to the audio feature vector of the long audio output by the siamese network structure and the background music feature vector of each alternative background music, the distance between the audio feature vector and the background music feature vector can be calculated, and the most compatible long audio and background music can be determined according to this distance to complete the audio accompaniment process.
[0085] The computer device in the embodiment of the present application will be described below. Please refer to Figure 4 An embodiment of the computer device in the embodiment of the present application includes:
[0086] The computer device 400 may include one or more central processing units (CPUs) 401 and a memory 405, and one or more applications or data are stored in the memory 405.
[0087] Among them, the memory 405 may be volatile storage or persistent storage. The programs stored in the memory 405 may include one or more modules, and each module may include a series of instruction operations on the computer device. Further, the central processing unit 401 may be configured to communicate with the memory 405 and execute a series of instruction operations in the memory 405 on the computer device 400.
[0088] The computer device 400 may further include one or more power supplies 402, one or more wired or wireless network interfaces 403, one or more input / output interfaces 404, and / or one or more operating systems, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0089] The central processing unit 401 may perform the operations performed by the computer device in the foregoing Figures 1 to 2 illustrated embodiments, and details are not described herein again.
[0090] An embodiment of the present application also provides a computer storage medium. In one embodiment, the computer storage medium stores instructions that, when executed on a computer, cause the computer to perform the operations performed by the computer device in the foregoing Figures 1 to 2 illustrated embodiments.
[0091] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and details are not described herein again.
[0092] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point, the displayed or discussed couplings or direct couplings or communication connections to each other may be through some interfaces, and the indirect couplings or communication connections of the devices or units may be in electrical, mechanical, or other forms.
[0093] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0094] In addition, each functional unit in various embodiments of the present application may be integrated into a processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0095] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs that can store program codes.
Claims
1. A method for training a neural network model, characterized in that, The method is applied to a computer device, and the method includes: Obtain multiple groups of training samples, where each group of the training samples includes a fusion feature vector corresponding to multiple-dimensional features of an audio sample and a fusion feature vector corresponding to multiple-dimensional features of a background music sample, and a true label for representing the adaptation degree between the audio sample and the background music sample; Input each group of the training samples into an initial neural network model, so that the initial neural network model performs the following operations: Extract the deep features of the audio sample by performing feature extraction on the fusion feature vector of the audio sample, and extract the deep features of the background music sample by performing feature extraction on the fusion feature vector of the background music sample; Concatenate the deep features of the audio sample and the deep features of the background music sample to obtain a deep fusion feature, and calculate a predicted label for the deep fusion feature; Construct a loss function according to the true label and the predicted label, and adjust the model parameters of the initial neural network model according to the loss value of the loss function, and stop training until the loss function meets the convergence condition, so as to obtain a target neural network model.
2. The method according to claim 1, wherein The obtaining of the fusion feature vector corresponding to the multiple-dimensional features of the audio sample includes: Obtain multiple pre-trained target audio feature extraction models, where the multiple target audio feature extraction models extract feature vectors for different-dimensional features of the audio sample; Extract the spectral features of the audio sample, input the spectral features of the audio sample into the multiple target audio feature extraction models, and each target audio feature extraction model outputs a feature vector of one-dimensional feature of the audio sample, so as to obtain the feature vectors of the multiple-dimensional features of the audio sample; Fuse the feature vectors of the multiple-dimensional features of the audio sample to obtain the fusion feature vector corresponding to the multiple-dimensional features of the audio sample.
3. The method according to claim 2, characterized in that, The training steps of each target audio feature extraction model include: Obtain multiple groups of audio training samples, where each group of the audio training samples includes an audio sample and a true feature label for representing one-dimensional feature of the audio sample; Obtain an initial audio feature extraction model, where the initial audio feature extraction model includes a convolutional structure and a Softmax layer connected to the convolutional structure; Extract the spectral features of the audio sample, input the spectral features of the audio sample into the initial audio feature extraction model, so that the initial audio feature extraction model uses the convolutional structure to extract feature vectors from the spectral features of the audio sample and uses the Softmax layer to calculate a predicted feature label for the feature vectors; Stop training when the relationship between the predicted feature label and the true feature label meets the convergence condition, so as to obtain the target audio feature extraction model.
4. The method according to claim 1, wherein The obtaining of the fusion feature vector corresponding to the multiple-dimensional features of the background music sample includes: Obtain multiple pre-trained target background music feature extraction models, where the multiple target background music feature extraction models extract feature vectors for different-dimensional features of the background music; Extract the spectral features of the background music sample, and input the spectral features of the background music sample into the multiple target background music feature extraction models. Each of the target background music feature extraction models outputs a feature vector of a dimensional feature of the background music sample, and obtain the feature vectors of multiple dimensional features of the background music sample; Fuse the feature vectors of multiple dimensional features of the background music sample to obtain a fused feature vector corresponding to multiple dimensional features of the background music sample.
5. The method according to claim 4, wherein The training steps of each of the target background music feature extraction models include: Obtain multiple groups of background music training samples, where each group of background music training samples includes a background music sample and a true feature label for representing a dimensional feature of the background music sample; Obtain an initial background music feature extraction model, where the initial background music feature extraction model includes a convolutional structure and a Softmax layer connected to the convolutional structure; Extract the spectral features of the background music sample, and input the spectral features of the background music sample into the initial background music feature extraction model, so that the initial background music feature extraction model uses the convolutional structure to extract features from the spectral features of the background music sample to obtain a feature vector, and uses the Softmax layer to calculate a predicted feature label for the feature vector; Stop training when the relationship between the predicted feature label and the true feature label satisfies the convergence condition, and obtain the target background music feature extraction model.
6. The method according to claim 1, wherein The initial neural network model is a siamese network model, where the siamese network model includes a first convolutional neural network and a second convolutional neural network. The first convolutional neural network and the second convolutional neural network are respectively connected to a fully connected layer, and the fully connected layer is connected to a sigmoid layer; The step of inputting the fused feature vector corresponding to the audio sample of each group of the training samples and the fused feature vector corresponding to the background music sample into the initial neural network model includes: Use the first convolutional neural network to extract features from the fused feature vector of the audio sample to obtain the deep feature of the audio sample, and use the second convolutional neural network to extract features from the fused feature vector of the background music sample to obtain the deep feature of the background music sample; Use the fully connected layer to splice the deep feature of the audio sample and the deep feature of the background music sample to obtain a deep fused feature; Use the sigmoid layer to calculate a predicted label for the deep fused feature.
7. The method according to any one of claims 1 to 6, characterized in that The fused feature vector of the audio sample is obtained by fusing multiple feature vectors of the audio sample, and the multiple feature vectors of the audio sample correspond to different feature dimensions of the audio sample; among them, the multiple dimensional features of the audio sample include prosody, reading speed, and emotion, each dimensional feature has multiple levels, and the feature vector of the audio sample corresponds to one level in a dimensional feature of the audio sample; The fused feature vector of the background music sample is obtained by fusing multiple feature vectors of the background music sample, and the multiple feature vectors of the background music sample correspond to different feature dimensions of the background music sample; wherein, the multiple dimensional features of the background music sample include genre, soothing degree, and emotion, each dimensional feature has multiple levels, and the feature vector of the background music sample corresponds to one level in one dimensional feature of the background music sample.
8. An audio background music method, characterized in that, The method is applied to a computer device, and the method includes: Obtaining fused feature vectors corresponding to multiple dimensional features of a target audio to be scored, and obtaining fused feature vectors corresponding to multiple dimensional features of each alternative background music; Inputting the fused feature vector of the target audio into a pre-trained target neural network model to obtain an audio feature vector output by the target neural network model; the target neural network model is trained by the neural network model training method according to any one of claims 1 to 7; Inputting the fused feature vector of the alternative background music into the target neural network model to obtain a background music feature vector output by the target neural network model; Calculating the distance between the audio feature vector and the background music feature vector of each alternative background music, and using the alternative background music corresponding to the background music feature vector closest to the audio feature vector as the background music of the target audio.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.
10. A computer storage medium, characterized in that, Instructions are stored in the computer storage medium, and when the instructions are executed on the computer, the computer is caused to execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Background music volume adjusting method and device, electronic equipment and storage medium
CN115277935A
Audio reconstruction method and device for image-assisted audio completion
CN115440251A