Training methods for audio matching models, audio matching methods, and computer equipment

By training an audio matching model and adjusting it using residual networks and loss functions, the problem of low matching degree between background music and audio content theme in audio matching was solved, achieving higher audio matching accuracy and listener experience.

CN115481278BActive Publication Date: 2026-07-31TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
Filing Date
2022-08-31
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing audio matching methods often result in low matching rates between background music and audio content themes, leading to poor audio content richness and listener experience.

Method used

By training an audio matching model, features of dry audio and melody audio are obtained. Residual networks are used for multi-dimensional downsampling and feature extraction. A loss function is constructed to adjust the model parameters, thereby achieving the matching of dry audio and melody.

Benefits of technology

It improves the accuracy of audio matching, ensures that the background music matches the theme of the audio content, and enhances the listener's experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115481278B_ABST
    Figure CN115481278B_ABST
Patent Text Reader

Abstract

This application relates to a training method for an audio matching model, an audio matching method, a computer device, and a program product. The method involves obtaining audio type identification results for each melody audio based on audio features using an audio matching model to be trained. A pre-trained audio matching model is then trained using a first loss function constructed from the audio type identification results and the audio types of the melody audio. The pre-trained audio matching model outputs melody audio that matches the audio type corresponding to the dry voice features, based on the dry voice features and melody features, as the audio matching result. Finally, a second loss function constructed from the audio matching result is used to train an audio matching model for matching dry voice and melody. Compared to traditional manual audio matching, this method, through pre-training based on audio type and retraining based on dry voice and melody matching, obtains an audio matching model that can be used for audio matching, thus improving the accuracy of audio matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to a training method for an audio matching model, an audio matching method, an apparatus, a computer device, a storage medium, and a computer program product. Background Technology

[0002] With the development of computer technology, people can now listen to audio through mobile phones and other computer devices. Furthermore, as people gradually learn to slow down and abandon the fast-paced, consumerist approach of short videos, they are choosing to enrich themselves with longer audio content such as audiobooks. Therefore, long-form audio technology is an important research direction now and in the future. Although human voice is the main content of long-form audio, background music often plays a crucial role in describing scenes and providing room for the listener's imagination. Adding suitable background music to long-form audio can enrich the audio content and allow listeners to become more engaged and immersive. Currently, the method of adding background music to long-form audio is usually through manual selection. However, everyone has different musical preferences, so the selected background music may not match the theme of the audio content, resulting in a low degree of matching.

[0003] Therefore, current audio matching methods suffer from low matching accuracy. Summary of the Invention

[0004] Therefore, it is necessary to provide a training method, audio matching method, apparatus, computer device, computer-readable storage medium, and computer program product for an audio matching model that can improve audio matching accuracy, in order to address the above-mentioned technical problems.

[0005] Firstly, this application provides a method for training an audio matching model, the method comprising:

[0006] The dry sound features of the dry audio are obtained, and the melody features and audio types of each melody audio in a plurality of melody audios are obtained; the plurality of melody audios include a target melody audio that matches the dry audio and other melody audios other than the target melody audio.

[0007] The melody features of each melody audio are input into the audio matching model to be trained, and the audio matching model to be trained identifies the audio type of each melody audio.

[0008] A first loss function is constructed based on the acquired audio type and the identified audio type of the melody audio. The model parameters of the audio matching model to be trained are adjusted based on the first loss function until the pre-training conditions are met to obtain the pre-trained audio matching model.

[0009] The dry features of the dry audio, the melody features of the target melody audio, and the melody features of other melody audios are input into the pre-trained audio matching model. The pre-trained audio matching model identifies the audio types of the dry features, the melody features of the target melody audio, and the melody features of other melody audios, and then outputs the audio matching result for the dry audio.

[0010] Based on the second loss function constructed from the audio matching results, the model parameters of the pre-trained audio matching model are adjusted until the training conditions are met to obtain the trained audio matching model.

[0011] In one embodiment, obtaining the audio type of each melody audio among multiple melody audios includes:

[0012] For each melody audio, obtain the audio name of that melody audio;

[0013] The audio database is queried based on the audio name to obtain the target audio information corresponding to the audio name; the audio database includes audio information of multiple melody audios;

[0014] Extract type information from the target audio information to determine the audio type of the melody audio.

[0015] In one embodiment, the audio matching model to be trained is a residual model to be trained; the step of inputting the melody features of each melody audio into the audio matching model to be trained, and having the audio matching model to be trained identify the audio type of each melody audio, includes:

[0016] The melody features of each melody audio are input into the residual model to be trained. The residual model to be trained performs multi-dimensional downsampling on the melody features of each melody audio based on multiple convolutional layers, and outputs the audio type of each melody audio based on the downsampled audio features. The number of multiple dimensions increases with the number of convolutional layers.

[0017] In one embodiment, the step of constructing a first loss function based on the acquired audio type and the identified audio type corresponding to the melody audio, and adjusting the model parameters of the audio matching model to be trained based on the first loss function until the pre-training conditions are met to obtain the pre-trained audio matching model includes:

[0018] Obtain the posterior probability of the melody features output by the activation layer of the audio matching model to be trained;

[0019] The system detects whether the acquired audio type of the melody audio matches the identified audio type of the melody audio, and determines the value of the parameter corresponding to the identified audio type in the first loss function based on the detection result.

[0020] The output value of the first loss function is determined based on the numerical value and the posterior probability.

[0021] If the output value of the first loss function is greater than the threshold of the first loss function, the model parameters of the audio matching model to be trained are adjusted according to the output value of the first loss function until the output value of the first loss function is less than or equal to the threshold of the first loss function, thus obtaining the pre-trained audio matching model.

[0022] In one embodiment, obtaining the dry sound features of the dry audio and obtaining the melody features of each melody audio among multiple melody audios includes:

[0023] Acquire dry audio and multiple melody audio; divide the dry audio and the multiple melody audio into multiple dry audio segments and multiple melody audio segments respectively;

[0024] For each dry audio segment, obtain the dry audio segment features corresponding to that dry audio segment; obtain the dry audio features based on multiple dry audio segment features;

[0025] For each melody audio segment corresponding to each melody audio, obtain the melody audio segment features corresponding to that melody audio segment; obtain the melody features corresponding to that melody audio based on multiple melody audio segment features.

[0026] In one embodiment, the step of inputting the dry audio features of the dry audio, the melody features of the target melody audio, and the melody features of other melody audios into the pre-trained audio matching model, and then having the pre-trained audio matching model identify the audio types of the dry audio features, the melody features of the target melody audio, and the melody features of the other melody audios, and outputting an audio matching result for the dry audio, includes:

[0027] The features of multiple dry audio segments corresponding to the dry audio, the features of multiple target melody audio segments corresponding to the target melody audio, and the features of multiple other melody audio segments corresponding to other melody audio are respectively input into the pre-trained audio matching model. The pre-trained audio matching model determines the dry audio vector corresponding to the audio type of the dry audio based on the average value of the features of the multiple dry audio segments, determines the target melody audio vector corresponding to the audio type of the target melody audio based on the average value of the features of the multiple target melody audio segments, and determines the other melody audio vector corresponding to the audio type of the other melody audio based on the average value of the features of the multiple other melody audio segments.

[0028] The audio matching result is output based on the distances between the dry audio vector and the target melody audio vector and the other melody audio vectors.

[0029] In one embodiment, adjusting the model parameters of the pre-trained audio matching model based on the second loss function constructed from the audio matching results until the training conditions are met to obtain the trained audio matching model includes:

[0030] Obtain the first cosine distance between the dry audio vector and the target melody audio vector;

[0031] Obtain the second cosine distance between the dry audio vector and the other melodic audio vectors;

[0032] The output value of the second loss function is determined based on the first cosine distance and the second cosine distance;

[0033] If the output value of the second loss function is greater than the threshold of the second loss function, the model parameters of the pre-trained audio matching model are adjusted according to the output value of the second loss function until the output value of the second loss function is less than or equal to the threshold of the second loss function, and the trained audio matching model is obtained.

[0034] Secondly, this application provides an audio matching method, the method comprising:

[0035] In response to an audio matching command, the dry characteristics of the dry audio are obtained;

[0036] The dry vocal features are input into an audio matching model. After the audio matching model identifies the dry vocal features and the audio type of each melody audio in the melody library, it outputs a target melody audio corresponding to the audio type of the dry vocal audio based on the matching degree between the dry vocal features and the audio types of each melody audio. This target melody audio has the highest matching degree with the audio type of the dry vocal audio. The audio matching model is trained based on the method described above.

[0037] Thirdly, this application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.

[0038] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.

[0039] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.

[0040] The aforementioned audio matching model training method, audio matching method, device, computer equipment, storage medium, and computer program product involve inputting multiple melody features into the audio matching model to be trained. The model then obtains audio type recognition results for each melody audio. A pre-trained audio matching model is trained based on a first loss function constructed from the audio type recognition results and the audio types of the melody audio. This pre-trained model outputs melody audio matching the corresponding audio type based on the dry vocal features, target melody features, and other melody features, serving as the audio matching result. Finally, a second loss function constructed from the audio matching result is used to train the final audio matching model for matching dry vocals and melodies. Compared to traditional manual audio matching, this scheme, through pre-training based on audio type and re-training based on dry vocal and melody matching, yields an audio matching model suitable for audio matching, thus improving the accuracy of audio matching. Attached Figure Description

[0041] Figure 1 This is a flowchart illustrating the training method of an audio matching model in one embodiment;

[0042] Figure 2 This is a schematic diagram of the residual network structure in one embodiment;

[0043] Figure 3 This is a flowchart illustrating the pre-training steps in one embodiment;

[0044] Figure 4 This is a flowchart illustrating the training steps of an audio matching model in one embodiment;

[0045] Figure 5 This is a schematic diagram of the structure of an audio matching model in one embodiment;

[0046] Figure 6 This is a flowchart illustrating an audio matching method in one embodiment;

[0047] Figure 7 This is a flowchart illustrating the audio matching method in another embodiment;

[0048] Figure 8 This is a schematic diagram of the interface for the audio matching step in one embodiment;

[0049] Figure 9 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0051] In one embodiment, such as Figure 1 As shown, a training method for an audio matching model is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and implemented through the interaction between the terminal and the server, including the following steps:

[0052] Step S202: Obtain the dry sound features of the dry sound audio, and obtain the melody features and audio type of each melody audio in the multiple melody audios; the multiple melody audios include the target melody audio that matches the dry sound audio and other melody audios besides the target melody audio.

[0053] The dry audio can be audio recorded by the user using the terminal, specifically the user's voice audio, such as the user reading aloud. The melody audio can be a type of music audio, serving as background music for the dry audio; that is, the dry audio can be matched with the melody audio. When matching dry audio and melody audio, the terminal can perform the matching based on audio features. The terminal can obtain the dry audio and its corresponding audio features, as well as multiple melody audios, their corresponding melody features, and their corresponding audio types. The audio features of the dry audio can be features extracted from human voice, and the audio features of the melody audio can be features extracted from the melody music. The audio features of the dry audio can be simply referred to as dry audio features, and the audio features of the melody audio can be simply referred to as melody features. The audio type of the melody audio can be music genre information, and the melody audio includes target melody audio that matches the dry audio, as well as other melody audios that do not match the dry audio. When the dry audio matches the target melody audio, it means that the audio type of the target melody audio is consistent with the audio type of the dry audio. In this case, the terminal can consider the audio type of the dry audio to be the audio type of the target melody audio that matches the dry audio.

[0054] When acquiring the audio features of a melody audio, the terminal can obtain the audio type of the melody audio from the audio information of the melody audio in the database. For example, in one embodiment, acquiring the audio type of each melody audio among multiple melody audios includes: acquiring the audio name of each melody audio; querying the audio database based on the audio name to obtain the target audio information corresponding to the audio name; the audio database includes the audio information of multiple melody audios; and extracting type information from the target audio information as the audio type of the melody audio. In this embodiment, there can be multiple melody audios, and each melody audio can have a corresponding audio name, such as a song name. For each melody audio, the terminal can acquire the audio name of the melody audio and query the audio database based on the audio name to obtain the target audio information corresponding to the audio name. The audio database includes the audio information of multiple melody audios. The audio information includes various types of information, and the terminal can extract the type information from the target audio information as the audio type of the melody audio. Specifically, the melody audio can be a song, and the audio type can be a song genre. The audio information of the song can include multiple fields, and the genre field can be one of the fields. When a terminal obtains the audio type of a song, it can query the audio database based on the song's name to retrieve the song's audio information and extract the genre field from the audio information as the audio type. At the same time, these audio types can be used as label information for training the audio matching model to be trained.

[0055] Step S204: Input the melody features of each melody audio into the audio matching model to be trained, and let the audio matching model to be trained identify the audio type of each melody audio.

[0056] The audio matching model to be trained can be a model that needs to be trained based on melody audio and the corresponding audio type. The terminal can acquire melody features corresponding to multiple melody audios and input these features into the audio matching model to be trained. The audio matching model then determines the audio type recognition result for each melody audio. Specifically, the audio matching model to be trained can identify the audio type of each of the multiple melody features, thereby allowing the terminal to determine the audio type of the melody audio corresponding to each melody feature and output the audio type recognition result. The training of the audio matching model by the terminal can be a pre-training process, used to train the audio matching model to be trained into a pre-trained audio matching model capable of recognizing the audio type of melody audio.

[0057] Step S206: Construct a first loss function based on the acquired audio type and the identified audio type of the melody audio, and adjust the model parameters of the audio matching model to be trained based on the first loss function until the pre-training conditions are met to obtain the pre-trained audio matching model.

[0058] The audio type identification result can be the audio type output by the audio matching model to be trained based on the audio features of the melody audio. The terminal can detect whether the audio type identification result matches the audio type of the melody audio input to the audio matching model to be trained, and construct a first loss function based on the audio type identification result and the audio type of the melody audio. The terminal can adjust the model parameters of the audio matching model to be trained based on the first loss function until the pre-training conditions are met, thus obtaining the pre-trained audio matching model. The pre-training conditions can be that the output value of the first loss function meets a certain numerical range, for example, when the output value of the first loss function is less than or equal to the threshold of the first loss function, it is determined that the audio matching model to be trained meets the pre-training conditions. Furthermore, the audio matching model to be trained can be a residual network model. The terminal can perform multi-level feature extraction and identification through the residual network to determine the audio type of the melody audio. Specifically, taking a song as an example, the pre-trained audio matching model can identify the genre information of the song.

[0059] Step S208: Input the dry features of the dry audio, the melody features of the target melody audio, and the melody features of other melody audios into the pre-trained audio matching model. The pre-trained audio matching model identifies the audio types of the dry features, the melody features of the target melody audio, and the melody features of other melody audios, and then outputs the audio matching results for the dry audio.

[0060] After the terminal trains the pre-trained audio matching model based on the audio features of the melody audio, the pre-trained audio matching model can identify the audio type represented by each audio feature. For example, it can identify the audio type corresponding to each melody audio. Among the multiple melody audios mentioned above, there is a target melody audio that matches the dry audio. The terminal can determine the audio type of the target melody audio through the pre-trained audio matching model, and the audio type of the dry audio is consistent with the audio type of its corresponding target melody audio.

[0061] To achieve audio type-based matching between dry audio and melodic audio, the terminal can perform secondary training on a pre-trained audio matching model to fine-tune the model parameters. The terminal can input the dry audio features, the target melodic features of the target melody audio, and other melodic features of other melody audio into the pre-trained audio matching model. The pre-trained audio matching model then identifies the audio type of each of the dry audio features, target melody features, and other melodic features, performs matching based on the audio type of each feature, and outputs the audio matching result for the dry audio based on the matching degree. Among them, other melodic features are the audio features of other melodic audios that do not match the dry audio among the above multiple melodic audios. The audio matching result can be a type of melodic audio. When the terminal uses the pre-trained audio matching model for recognition, it can identify the audio type to which each melodic feature belongs and the audio type of the dry audio feature, and determine the melodic audio that is closest to the audio type of the dry audio by determining the distance between the vectors corresponding to each audio type, and output the corresponding melodic audio. Furthermore, the terminal can compare the melodic audio output by the pre-trained audio matching model with the melodic audio that matches the input dry audio, and train the pre-trained audio matching model based on the comparison result to obtain the final audio matching model.

[0062] Step S210: Based on the second loss function constructed from the audio matching results, adjust the model parameters of the pre-trained audio matching model until the training conditions are met to obtain the trained audio matching model.

[0063] The audio matching result can be a matching relationship between a melody audio and a dry audio, output by the pre-trained audio matching model based on the input dry audio features and melody features. The terminal can construct a second loss function based on the audio matching result and adjust the model parameters of the pre-trained audio matching model based on the output value of the second loss function until the training conditions are met, thus obtaining a trained audio matching model. Specifically, the terminal can determine the second loss function based on the degree of matching between the dry audio and its target melody audio, as well as the degree of matching between the dry audio and other melody audio in the audio matching result. By adjusting the model parameters, the degree of matching between the dry audio and its target melody audio is maximized. This reduces the output value of the second loss function. When the output value of the second loss function is less than its threshold, the terminal obtains the trained audio matching model. The terminal can then use the trained audio matching model to identify the melody audio with the highest audio type match degree with the dry audio, as the audio matching result for the dry audio.

[0064] In the training and matching methods of the aforementioned audio matching model, multiple melody features are input into the audio matching model to be trained. The model then obtains the audio type recognition results for each melody audio. A pre-trained audio matching model is trained based on a first loss function constructed from the audio type recognition results and the audio types of the melody audio. The pre-trained model outputs melody audio that matches the audio type corresponding to the dry voice features, based on the dry voice features, target melody features, and other melody features. This output is the audio matching result. A second loss function constructed from the audio matching result is then used to train the final audio matching model for matching dry voice and melody. Compared to traditional manual audio matching, this approach, through pre-training based on audio type and re-training based on dry voice and melody matching, yields an audio matching model that can be used for audio matching, thus improving the accuracy of audio matching.

[0065] In one embodiment, the melody features of each melody audio are input into the audio matching model to be trained, and the audio matching model to be trained identifies the audio type of each melody audio. This includes: inputting the melody features of each melody audio into the residual model to be trained, the residual model to be trained performing multi-dimensional downsampling on the melody features of each melody audio based on multiple convolutional layers, and outputting the audio type of each melody audio based on the downsampled audio features; the number of multiple dimensions increases with the number of convolutional layers.

[0066] In this embodiment, the audio matching model to be trained can be a residual model to be trained. The residual model can include multiple convolutional layers, and each convolutional layer can contain multiple sampling dimensions. When the terminal trains the audio matching model, it can input the multiple melody features into the residual model. The residual model then downsamples the melody features corresponding to each melody audio in multiple dimensions based on the multiple convolutional layers. After multiple downsampling operations, the terminal can output the audio type recognition result of each melody audio based on the downsampled audio features. The number of dimensions increases with the number of convolutional layers. By downsampling the audio features with gradually increasing dimensions, the terminal can ultimately determine the audio type of the melody audio based on the sampled features.

[0067] Specifically, such as Figure 2 As shown, Figure 2This is a schematic diagram of a residual network structure in one embodiment. Each box in the diagram can represent a convolutional layer. Parameters within the boxes, such as 3*3conv and 128, indicate that the convolutional layer has 128 kernels with a kernel size of 3x3. `maxpool` represents max pooling, and `avgpool` represents average pooling. Parameters outside the boxes, such as 128*28*28, k=3, s=1, p=1, indicate that the current convolutional layer sample size is 128*28*28, the stride `s` is 1, the kernel size `k` is 3*3, and the padding scale `p` is 1. The residual network downsampling described above is a process of removing irrelevant information layer by layer. The closer to the output, the closer the network output is to the task. Therefore, the terminal can extract embedding features near the network output to determine the audio type of the melody audio based on these features. Through the convolutional neural network with the above residual structure, the terminal can avoid gradient problems when deepening the network, thus enhancing the model's learning ability and improving its transferability.

[0068] Furthermore, the terminal can construct a first loss function based on the aforementioned audio type recognition results and the audio type of the melody audio, and adjust the model parameters of the audio matching model to be trained, thereby obtaining a pre-trained audio matching model. For example, in one embodiment, a first loss function is constructed based on the acquired audio type and the recognized audio type corresponding to the melody audio, and the model parameters of the audio matching model to be trained are adjusted based on the first loss function until the pre-training conditions are met to obtain the pre-trained audio matching model. This includes: obtaining the posterior probability of the melody features output by the activation layer of the audio matching model to be trained; detecting whether the acquired audio type of the melody audio matches the recognized audio type of the melody audio, and determining the value of the parameter corresponding to the recognized audio type in the first loss function based on the detection result; determining the output value of the first loss function based on the value and the posterior probability; if the output value of the first loss function is greater than the threshold of the first loss function, adjusting the model parameters of the audio matching model to be trained based on the output value of the first loss function until the output value of the first loss function is less than or equal to the threshold of the first loss function, thereby obtaining the pre-trained audio matching model.

[0069] In this embodiment, the audio matching model to be trained may include an activation layer, and the terminal can obtain the posterior probability of the melody feature output by the activation layer in the audio matching model to be trained. Furthermore, the first loss function includes a parameter corresponding to the audio type recognition result and a posterior probability. The terminal can detect the match between the audio type recognition result and the audio type of the melody audio. If the audio type recognition result matches the audio type of the melody audio, the terminal can determine that the value of the parameter corresponding to the audio type recognition result in the first loss function is a first value, and the posterior probability of the melody feature can be recorded as the first posterior probability. If the terminal detects that the audio type recognition result does not match the audio type of the melody audio, the terminal can determine that the value of the parameter corresponding to the audio type recognition result in the first loss function is a second value, and the posterior probability of the melody feature can be recorded as the second posterior probability. Wherein, the first value and the second value are different, and the sum of the first value and the second value is one; the first posterior probability and the second posterior probability are different, and the sum of the first posterior probability and the second posterior probability is one.

[0070] The terminal can obtain the output value of the first loss function based on the above values ​​and posterior probabilities. If the output value of the first loss function is greater than the threshold of the first loss function, the terminal can adjust the model parameters of the audio matching model to be trained according to the first loss function until the output value of the first loss function is less than or equal to the threshold of the first loss function. At this point, the terminal can obtain the pre-trained audio matching model that has been pre-trained.

[0071] Specifically, such as Figure 3 As shown, Figure 3 This is a flowchart illustrating the pre-training steps in one embodiment. The multiple melody audios mentioned above can be multiple background music tracks. When training the audio matching model to be trained, the terminal can acquire multiple background music tracks, which can serve as background music for the long audio input by the user. The terminal can segment the background music into fixed-length sub-segments and obtain the genre fields of these songs from the audio database. It then trains the genre classification model using a ResNet residual neural network, i.e., the aforementioned pre-trained audio matching model. Based on the aforementioned residual network model and the first loss function, the terminal can adjust the model parameters of the audio matching model to be trained, so that the audio type output by the audio matching model to be trained matches the audio type corresponding to the input melody audio. Specifically, the aforementioned first loss function is as follows: Loss = -[y i logp+(1-y i)log(1-p)]. Where Loss represents the output value of the first loss function, yi represents the parameter corresponding to the audio type recognition result, which can be either the first or second value mentioned above, and p represents the posterior probability output by the softmax layer in the audio matching model to be trained; when the audio type recognition result output by the audio matching model to be trained matches the audio type of the input melody audio, y i The value is 1 if the audio type is true and 0 otherwise. Additionally, in some embodiments, when the audio type recognition result output by the audio matching model to be trained matches the audio type of the input melody audio, y... i It can also be other values. When the audio type recognition result output by the audio matching model to be trained does not match the audio type of the input melody audio, then it is (1-y i The corresponding value for ) and y i and (1-y i) The values ​​are different.

[0072] Through the above embodiments, the terminal can train a pre-trained audio matching model that can perform audio type recognition based on the residual neural network and the first loss function. Thus, the terminal can adjust the model parameters of the pre-trained audio matching model to obtain an audio matching model that can be used to match dry audio and melodic audio, thereby improving the accuracy of audio matching.

[0073] In one embodiment, obtaining the dry audio features of dry audio and obtaining the melody features of each melody audio in multiple melody audios includes: obtaining dry audio and multiple melody audio; segmenting the dry audio and multiple melody audio into multiple dry audio segments and multiple melody audio segments respectively; for each dry audio segment, obtaining the dry audio segment features corresponding to that dry audio segment; obtaining the dry audio features based on the multiple dry audio segment features; for each melody audio corresponding to each melody audio, obtaining the melody audio segment features corresponding to that melody audio segment; and obtaining the melody features corresponding to that melody audio based on the multiple melody audio segment features.

[0074] In this embodiment, the terminal can extract audio features from the dry audio and the melody audio respectively. When extracting audio features, the terminal can acquire dry audio and multiple melody audio segments, and segment both the dry audio and the multiple melody audio segments to obtain multiple dry audio segments and multiple melody audio segments of preset duration. The preset duration can be a value shorter than the duration of all the aforementioned audio segments. Thus, the terminal can obtain multiple dry audio segments and multiple melody audio segments. For each dry audio segment, the terminal can acquire the corresponding dry audio segment features. Furthermore, the terminal can extract audio features from the multiple dry audio segments to obtain multiple dry audio segment features. Based on the multiple dry audio segment features, the terminal can obtain the aforementioned dry audio features.

[0075] There can be multiple melody audio files, and each melody audio file can correspond to multiple melody audio segments. For each melody audio segment corresponding to each melody audio file, the terminal can obtain the melody audio segment features corresponding to that melody audio segment. The terminal can extract audio segment features for each melody audio segment, thereby obtaining the melody features corresponding to that melody audio file based on multiple melody audio segment features.

[0076] After obtaining the aforementioned audio features by segmenting the audio into segments, the terminal can train an audio matching model based on these features. For example, in one embodiment, the dry audio features of the dry audio, the melody features of the target melody audio, and the melody features of other melody audios are input into a pre-trained audio matching model. The pre-trained audio matching model identifies the audio types of the dry audio features, the melody features of the target melody audio, and the melody features of other melody audios, and outputs an audio matching result for the dry audio. This includes: inputting multiple dry audio segment features corresponding to the dry audio, multiple target melody audio segment features corresponding to the target melody audio, and multiple other melody audio segment features corresponding to other melody audios into the pre-trained audio matching model; determining the dry audio vector corresponding to the audio type of the dry audio based on the average value of the multiple dry audio segment features, determining the target melody audio vector corresponding to the audio type of the target melody audio based on the average value of the multiple target melody audio segment features, and determining the other melody audio vector corresponding to the audio type of other melody audio based on the average value of the multiple other melody audio segment features; and outputting the audio matching result based on the distance between the dry audio vector and the target melody audio vector and the other melody audio vectors.

[0077] In this embodiment, the aforementioned dry audio features may include multiple dry audio segment features, the aforementioned target melody features may include multiple melody audio segment features, and the aforementioned other melody features may include multiple other melody audio segment features. The terminal can input the aforementioned multiple dry audio segment features, multiple melody audio segment features, and multiple other melody audio segment features into a pre-trained audio matching model. The terminal can obtain the average value of the multiple dry audio segment features in the fully connected layer through the pre-trained audio matching model, and determine the audio type of the dry audio based on the average value of the multiple dry audio segment features. Thus, the terminal can obtain the dry audio vector corresponding to the audio type of the dry audio based on the average value. The terminal can also determine the audio type of the target melody audio based on the average value of the aforementioned multiple target melody audio segment features, and thus obtain the target melody audio vector corresponding to the target melody audio type. The terminal can also determine the audio type of other melody audio based on the average value of the aforementioned multiple other audio segment features, and thus obtain the other melody audio vector corresponding to the audio type of other melody audio based on the average value. The positions of the aforementioned dry audio vector, target melody audio vector, and other melody audio vectors in the neural network of the pre-trained audio matching model may be different. Therefore, there is a certain distance between the three vectors. The terminal can obtain the distance between the aforementioned dry audio vector and the target melody audio vector, as well as the distance between the dry audio vector and other melody audio vectors. Thus, the terminal can output the audio matching result based on the distance between the dry audio vector and the target melody audio vector and other melody audio vectors respectively.

[0078] Specifically, the audio matching process performed by the pre-trained audio matching model described above can be as follows: Figure 4 As shown, Figure 4This is a flowchart illustrating the training steps of an audio matching model in one embodiment. After the terminal trains the pre-trained audio matching model, it can fine-tune the network parameters of the pre-trained audio matching model based on the aforementioned dry audio and melody audio. The terminal can acquire multiple long audio works that have been set to music. These works include dry audio as training samples and corresponding target melody audio as training samples. These works need to have high playback volume and a certain level of popularity so that they can positively guide the pre-trained audio matching model. The terminal can also acquire other melody audio besides the aforementioned target melody audio, i.e., melody audio from other genres, and use them together with the aforementioned dry audio and target melody audio as training data. Among them, the aforementioned dry audio can be a long audio, and the aforementioned melody audio can be background music. The terminal can select background music 'a' and acquire the long audio that uses the background music as background music, i.e., background music 'a' matches the long audio, and background music 'a' is the target melody audio of the long audio. The terminal can also randomly select a background music 'b' from other genres besides the genre corresponding to background music 'a', i.e., the aforementioned other melody audio. The terminal can segment the aforementioned audio data into fixed-length segments, such as segments of preset duration, and then extract audio features based on each segment. The terminal can input these audio features into a ResNet neural network, i.e., the pre-trained audio matching model. The terminal can calculate a multivariate loss based on the embedding layer, extract the embedding (i.e., the aforementioned audio vectors) from the fully connected layer before the Softmax layer, and calculate the network loss based on the multivariate loss function.

[0079] Through the above embodiments, the terminal can extract audio features based on audio segments, and train the pre-trained audio matching model based on audio vectors using multiple audio segment features to obtain an audio matching model that can be used to match dry audio and melodic audio, thereby improving the accuracy of audio matching.

[0080] In one embodiment, the model parameters of a pre-trained audio matching model are adjusted based on a second loss function constructed from the audio matching results until the training conditions are met to obtain a trained audio matching model. This includes: obtaining a first cosine distance between the dry audio vector and the target melody audio vector; obtaining a second cosine distance between the dry audio vector and other melody audio vectors; determining the output value of the second loss function based on the first and second cosine distances; if the output value of the second loss function is greater than a second loss function threshold, adjusting the model parameters of the pre-trained audio matching model based on the output value of the second loss function until the output value of the second loss function is less than or equal to the second loss function threshold to obtain a trained audio matching model.

[0081] In this embodiment, the terminal can train the pre-trained audio matching model based on a second loss function. This second loss function includes parameters for the dry audio vector, the target melody audio vector, and other melody audio vectors. Since the dry audio vector, target melody audio vector, and other melody audio vectors have corresponding positions in the pre-trained audio matching model, the terminal can obtain the first cosine distance between the dry audio vector and the target melody audio vector, and the second cosine distance between the dry audio vector and other melody audio vectors. The terminal can construct a second loss function based on the first and second cosine distances and obtain its output value. If the terminal detects that the output value of the second loss function is greater than a threshold value, it can adjust the model parameters of the pre-trained audio matching model according to the second loss function until the output value of the second loss function is less than or equal to the threshold value. At this point, the terminal obtains the trained audio matching model.

[0082] Specifically, taking the example of the dry audio being a long audio file and the melody audio being background music, the positional distribution between the dry audio vector, the target melody audio vector, and other melody audio vectors can be as follows: Figure 5 As shown, Figure 5 This is a schematic diagram of the audio matching model in one embodiment. The aforementioned dry audio vector can be obtained based on the average of features from multiple dry audio segments. The aforementioned target melody audio vector can be obtained based on the average of features from multiple target melody audio segments. The aforementioned other melody audio vectors can be obtained based on the average of features from multiple other melody audio segments. It should be noted that audio segment features are also called embedding values. Figure 5 In the above long audio, Anchor is the average embedding of all sub-segments, Positive is the average embedding of sub-segments of background music a, and Negative is the average embedding of sub-segments of background music b. The goal of the terminal adjusting the pre-trained audio matching model is to make the Positive vector as close as possible to the Anchor vector.

[0083] The second loss function is as follows: L = max(d(a,p) - d(a,n) + margin, 0). Here, L represents the output value of the second loss function, a represents the Anchor, p represents the Positive, n represents the Negative, d(a,p) represents the cosine distance between the anchor and the positive, d(a,n) represents the cosine distance between the anchor and the negative, and margin is an adjustable coefficient. The terminal can adjust the model parameters of the pre-trained audio matching model to make the background song 'a' closer to its corresponding music and farther away from other non-corresponding music in the embedding space. When the pre-trained audio matching model is fine-tuned to convergence, for example, when the output value of the second loss function is less than or equal to the threshold of the second loss function, the terminal can obtain the trained audio matching model.

[0084] Through this embodiment, the terminal can train the pre-trained audio matching model based on the cosine distance between each vector and the second loss function to obtain an audio matching model for matching dry audio and melodic audio, thereby improving the accuracy of audio matching.

[0085] In one embodiment, such as Figure 6 As shown, an audio matching method is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server, including the following steps:

[0086] Step S302: In response to the audio matching command, obtain the dry characteristics of the dry audio.

[0087] The audio matching command can be a user-initiated command via the terminal, used to match dry audio and melodic audio. The user can input dry audio into the terminal and trigger the audio matching command. After receiving the command, the terminal can acquire the user-input dry audio and obtain its features. Specifically, the terminal can extract spectral features from the dry audio as its dry audio features.

[0088] Step S304: Input the dry vocal features into the audio matching model. After the audio matching model identifies the dry vocal features and the audio type of each melody audio in the melody library, the target melody audio corresponding to the dry vocal audio is output based on the matching degree between the dry vocal features and the audio type of each melody audio, and is used as the audio matching result. The matching degree between the target melody audio and the audio type corresponding to the dry vocal audio is the largest. The audio matching model is trained based on the above method.

[0089] After acquiring the dry sound features, the terminal can input these features into an audio matching model. The model then identifies the audio type of both the dry sound features and the melody features of various melody audios in a melody library. Based on the audio type of the dry sound features, the terminal identifies the melody audio that matches that type. This melody audio can be selected from a preset melody library, and each melody audio has a certain degree of matching. The terminal can output a target melody audio of the corresponding audio type as the audio matching result. The melody library includes various types of melody audio, and the target melody audio can be the melody audio with the highest matching degree among the matched melody audios. The terminal can obtain the audio matching model based on the training method described above.

[0090] In the aforementioned audio matching method, multiple melody features are input into the audio matching model to be trained. The model then obtains the audio type identification results for each melody audio. A pre-trained audio matching model is trained based on a first loss function constructed from the audio type identification results and the audio types of the melody audio. The pre-trained model outputs melody audio that matches the audio type corresponding to the dry voice features, based on the dry voice features, target melody features, and other melody features. This output serves as the audio matching result. Finally, a second loss function constructed from the audio matching result is used to train the final audio matching model, which is then used for matching dry voice and melody. Compared to traditional manual audio matching, this approach, through pre-training based on audio type and re-training based on dry voice and melody matching, yields an audio matching model that can be used for audio matching, thus improving the accuracy of audio matching.

[0091] In one embodiment, such as Figure 7 As shown, Figure 7 This is a flowchart illustrating the audio matching method in another embodiment. In this embodiment, the aforementioned melody audio can be background music. The terminal can acquire candidate background music to construct an audio vector embedding library as a retrieval library. The terminal can calculate the distance between the audio vector embedding of the long audio input by the user without background music and the audio vector embedding of the background music in the retrieval library, and determine the minimum distance, which represents the highest matching degree with the input long audio. Then, the terminal can use the background music corresponding to that audio vector as the best background music result for the input long audio.

[0092] The interface displayed by the terminal during audio matching can be as follows: Figure 8 As shown, Figure 8This is a schematic diagram of the audio matching step in one embodiment. Users can input long audio based on the text being read aloud, and by triggering an audio matching command, the terminal can determine the best background music to match the long audio through an audio matching model and output it, thereby realizing the addition of background music to long audio recordings of human voices.

[0093] Through the above embodiments, the terminal can match human voice audio and melody audio based on the audio matching model, thereby improving the accuracy of audio matching.

[0094] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0095] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 9 As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a training method for an audio matching model and an audio matching method. The display screen can be an LCD screen or an e-ink display screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0096] Those skilled in the art will understand that Figure 9The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0097] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-described training method for the audio matching model and the audio matching method.

[0098] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the above-described training method for the audio matching model and the audio matching method.

[0099] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described training method for the audio matching model and the audio matching method.

[0100] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0101] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0102] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0103] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A training method for an audio matching model, characterized in that, The method includes: The dry audio features are obtained, and the melody features and audio types of each melody audio in multiple melody audios are obtained; the multiple melody audios include a target melody audio that matches the dry audio and other melody audios besides the target melody audio; the dry audio represents human voice reading audio without background music; The melody features of each melody audio are input into the audio matching model to be trained, and the audio matching model to be trained identifies the audio type of each melody audio. A first loss function is constructed based on the acquired audio type and the identified audio type of the melody audio. The model parameters of the audio matching model to be trained are adjusted based on the first loss function until the pre-training conditions are met to obtain the pre-trained audio matching model. The dry audio features, the melody features of the target melody audio, and the melody features of other melody audios are input into the pre-trained audio matching model. The pre-trained audio matching model identifies the audio types of the dry audio features, the melody features of the target melody audio, and the melody features of the other melody audios, and outputs an audio matching result for the dry audio. The audio matching result includes the melody audio identified by the pre-trained audio matching model that is closest to the audio type of the dry audio. Based on the second loss function constructed from the audio matching results, the model parameters of the pre-trained audio matching model are adjusted until the training conditions are met to obtain the trained audio matching model. The trained audio matching model is used to identify the target melody audio with the highest audio type matching degree with the dry audio from a melody library including multiple types of melody audio, and use it as the audio matching result of the dry audio.

2. The method according to claim 1, characterized in that, The step of obtaining the audio type of each melody audio in multiple melody audios includes: For each melody audio, obtain the audio name of that melody audio; The audio database is queried based on the audio name to obtain the target audio information corresponding to the audio name; the audio database includes audio information of multiple melody audios; Extract type information from the target audio information to determine the audio type of the melody audio.

3. The method according to claim 1, characterized in that, The audio matching model to be trained is a residual model to be trained; the step of inputting the melody features of each melody audio into the audio matching model to be trained, and having the audio matching model to be trained identify the audio type of each melody audio, includes: The melody features of each melody audio are input into the residual model to be trained. The residual model to be trained performs multi-dimensional downsampling on the melody features of each melody audio based on multiple convolutional layers, and outputs the audio type of each melody audio based on the downsampled audio features. The number of multiple dimensions increases with the number of convolutional layers.

4. The method according to claim 1, characterized in that, The process of constructing a first loss function based on the acquired and recognized audio types corresponding to the melody audio, and adjusting the model parameters of the audio matching model to be trained based on the first loss function until the pre-training conditions are met to obtain the pre-trained audio matching model, includes: Obtain the posterior probability of the melody features output by the activation layer of the audio matching model to be trained; The system detects whether the acquired audio type of the melody audio matches the identified audio type of the melody audio, and determines the value of the parameter corresponding to the identified audio type in the first loss function based on the detection result. The output value of the first loss function is determined based on the numerical value and the posterior probability. If the output value of the first loss function is greater than the threshold of the first loss function, the model parameters of the audio matching model to be trained are adjusted according to the output value of the first loss function until the output value of the first loss function is less than or equal to the threshold of the first loss function, thus obtaining the pre-trained audio matching model.

5. The method according to claim 1, characterized in that, The process of obtaining the dry audio features and the melodic features of each melodic audio in multiple melodic audio sequences includes: Acquire dry audio and multiple melody audio; divide the dry audio and the multiple melody audio into multiple dry audio segments and multiple melody audio segments respectively; For each dry audio segment, obtain the dry audio segment features corresponding to that dry audio segment; obtain the dry audio features based on multiple dry audio segment features; For each melody audio segment corresponding to each melody audio, obtain the melody audio segment features corresponding to that melody audio segment; obtain the melody features corresponding to that melody audio based on multiple melody audio segment features.

6. The method according to claim 5, characterized in that, The process involves inputting the dry audio features of the dry audio, the melody features of the target melody audio, and the melody features of other melody audio into the pre-trained audio matching model. The pre-trained audio matching model then identifies the audio types of the dry audio features, the melody features of the target melody audio, and the melody features of the other melody audio, and outputs an audio matching result for the dry audio. This includes: The features of multiple dry audio segments corresponding to the dry audio, the features of multiple target melody audio segments corresponding to the target melody audio, and the features of multiple other melody audio segments corresponding to other melody audio are respectively input into the pre-trained audio matching model. The pre-trained audio matching model determines the dry audio vector corresponding to the audio type of the dry audio based on the average value of the features of the multiple dry audio segments, determines the target melody audio vector corresponding to the audio type of the target melody audio based on the average value of the features of the multiple target melody audio segments, and determines the other melody audio vector corresponding to the audio type of the other melody audio based on the average value of the features of the multiple other melody audio segments. The audio matching result is output based on the distances between the dry audio vector and the target melody audio vector and the other melody audio vectors.

7. The method according to claim 6, characterized in that, The step of adjusting the model parameters of the pre-trained audio matching model based on the second loss function constructed from the audio matching results until the training conditions are met to obtain the trained audio matching model includes: Obtain the first cosine distance between the dry audio vector and the target melody audio vector; Obtain the second cosine distance between the dry audio vector and the other melodic audio vectors; The output value of the second loss function is determined based on the first cosine distance and the second cosine distance; If the output value of the second loss function is greater than the threshold of the second loss function, the model parameters of the pre-trained audio matching model are adjusted according to the output value of the second loss function until the output value of the second loss function is less than or equal to the threshold of the second loss function, and the trained audio matching model is obtained.

8. An audio matching method, characterized in that, The method includes: In response to an audio matching command, the dry audio features are obtained; the dry audio represents human voice reading audio without background music; The dry vocal features are input into an audio matching model. After the audio matching model identifies the dry vocal features and the audio types of each melody audio in the melody library, it outputs the target melody audio corresponding to the dry vocal audio based on the matching degree between the dry vocal features and the audio types of each melody audio, as the audio matching result. The target melody audio has the highest matching degree with the audio type of the dry vocal audio. The audio matching model is trained based on the method described in any one of claims 1 to 7. The training object includes a pre-trained audio matching model. The audio matching result output by the pre-trained audio matching model includes the melody audio that is closest to the audio type of the dry vocal audio.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.