Sound effect matching method and device, storage medium and computer program product

The pre-trained tone prediction model extracts the tone characteristics of the singer's audio to be matched, selects the matching tone vector and determines its sound effect parameters, which solves the problem that ordinary users find it difficult to adapt to the singer's singing audio, and improves the convenience and auditory experience of the sound effect parameters.

CN120220707APending Publication Date: 2025-06-27TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510343359.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Ordinary users need to manually adjust the sound effects parameters when selecting sound effects mode, and lack professional knowledge of audio files, making it difficult to adapt to singer singing audio, reducing the auditory experience.

Method used

The pre-trained tone prediction model uses the pre-trained tone prediction model to extract the matching singer audio and alternative singer audio, determine the similarity of the tone vector, and select the matching target tone vector from the alternative tone vectors to determine its corresponding sound effect parameters.

Benefits of technology

It improves the convenience of determining the sound effects parameters of the singer's audio to be matched, avoids the difficulty of users setting sound effects parameters manually, and improves the hearing experience of ordinary users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220707A_ABST
    Figure CN120220707A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a sound effect matching method and device, a storage medium and a computer program product, which are used for improving the convenience of determining sound effect parameters of singer audio to be matched. The method provided by the embodiment of the invention comprises the following steps: performing timbre feature extraction on a to-be-matched singer audio and at least two alternative singer audios through a pre-trained timbre prediction model, obtaining a tone vector of a to-be-matched singer corresponding to the to-be-matched singer audio and at least two alternative tone vectors of alternative singers corresponding to the at least two alternative singer audios, and training a pre-trained tone prediction model by adopting comparative learning and a ternary loss function; respectively determining the similarity between the tone vector and each alternative tone vector; determining a target timbre vector matched with the timbre vector from the at least two alternative timbre vectors according to the similarity; and determining the sound effect parameter of the audio corresponding to the target timbre vector as the sound effect parameter of the audio of the singer to be matched.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio data processing, and in particular, to a sound effect matching method, device, storage medium, and computer program product. Background Art

[0002] In current audio playback technologies, selecting a suitable sound effect mode is crucial for enhancing the user's auditory experience.

[0003] In the prior art, when selecting a sound effect mode, users usually need to manually adjust various sound effect parameters, or select a sound effect mode from several default settings. However, the current methods for users to select a sound effect mode require users to have an in-depth understanding of the professional attributes of audio files, and ordinary users generally lack professional knowledge in audio files. Therefore, when singing an audio, ordinary users cannot manually adjust the sound effect parameters suitable for the current singer's audio file, and correspondingly, the auditory experience of ordinary users singing the audio is reduced. Summary of the Invention

[0004] Embodiments of the present invention provide a sound effect matching method, device, storage medium, and computer program product, which are used to determine the sound effect parameters suitable for the audio of the singer to be matched according to the audio of the singer to be matched and the voice color prediction model, thereby improving the convenience of determining the sound effect parameters of the audio of the singer to be matched.

[0005] The first aspect of the embodiments of the present application provides a sound effect matching method, including:

[0006] Performing voice color feature extraction on the audio of the singer to be matched and at least two alternative singer audios through a pre-trained voice color prediction model, to obtain the voice color vector of the singer to be matched corresponding to the audio of the singer to be matched and at least two alternative voice color vectors of the alternative singers corresponding to the at least two alternative singer audios, where the pre-trained voice color prediction model is trained using contrastive learning and a triplet loss function;

[0007] Respectively determining the similarity between the voice color vector and each of the alternative voice color vectors;

[0008] According to the similarity, determining a target voice color vector that matches the voice color vector from the at least two alternative voice color vectors;

[0009] Determining the sound effect parameters of the audio corresponding to the target voice color vector as the sound effect parameters of the audio of the singer to be matched.

[0010] The second aspect of the embodiments of the present application provides a sound effect matching method, including:

[0011] The timbre vectors of the singer audio to be matched and the timbre vectors of at least two alternative singer audios are subjected to timbre and sound effect conversion through a pre-trained timbre and sound effect conversion model, so as to obtain the timbre and sound effect vectors of the singer audio to be matched and at least two alternative timbre and sound effect vectors of the at least two alternative singer audios, wherein the pre-trained timbre and sound effect conversion model is trained by using contrastive learning and a triplet loss function;

[0012] The similarities between the timbre and sound effect vectors and each of the alternative timbre and sound effect vectors are determined respectively;

[0013] According to the similarities, a target timbre and sound effect vector that matches the timbre and sound effect vector is determined from the at least two alternative timbre and sound effect vectors;

[0014] The sound effect parameters of the audio corresponding to the target timbre and sound effect vector are determined as the sound effect parameters of the singer audio to be matched.

[0015] A third aspect of the embodiments of the present application provides a computer device, including a processor, which is configured to implement the sound effect matching method provided in the first aspect of the embodiments of the present application, or the sound effect matching method provided in the second aspect of the embodiments of the present application when executing a computer program stored in a memory.

[0016] A fourth aspect of the embodiments of the present application provides a computer-readable storage medium, on which a computer program is stored. The computer program, when executed by a processor, is configured to implement the sound effect matching method provided in the first aspect of the embodiments of the present application, or the sound effect matching method provided in the second aspect of the embodiments of the present application.

[0017] A fifth aspect of the embodiments of the present application provides a computer program product, on which a computer program is stored. The computer program, when executed by a processor, is configured to implement the sound effect matching method provided in the first aspect of the embodiments of the present application, or the sound effect matching method provided in the second aspect of the embodiments of the present application.

[0018] It can be seen from the above technical solutions that the embodiments of the present invention have the following advantages:

[0019] The sound effect matching method in the embodiments of the present application includes: extracting timbre features from the singer audio to be matched and at least two alternative singer audios through a pre-trained timbre prediction model, obtaining the timbre vector of the singer to be matched corresponding to the singer audio to be matched and at least two alternative timbre vectors of the alternative singers corresponding to the at least two alternative singer audios, wherein the pre-trained timbre prediction model is trained using contrastive learning and a triplet loss function; respectively determining the similarity between the timbre vector and each of the alternative timbre vectors; determining, according to the similarity, a target timbre vector that matches the alternative timbre vector from the at least two timbre vectors; and determining the sound effect parameters of the audio corresponding to the target timbre vector as the sound effect parameters of the singer audio to be matched.

[0020] Because the pre-trained timbre prediction model in the embodiments of the present application can, based on the similarity between the timbre vector of the singer audio to be matched and at least two timbre vectors of at least two alternative singers, determine a target timbre vector similar to the timbre vector of the singer to be matched from the at least two alternative timbre vectors of the at least two alternative singers, and finally determine the sound effect parameters corresponding to the target timbre vector as the sound effect parameters of the singer audio to be matched, thereby improving the convenience of setting the sound effect parameters of the singer audio to be matched and avoiding the problem of allowing the user to manually set the sound effect parameters of the singer audio to be matched. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a schematic diagram of the architecture of the sound effect matching system in the embodiments of the present application;

[0022] Figure 2 It is a schematic diagram of an embodiment of the audio matching method in the embodiments of the present application;

[0023] Figure 3 It is a schematic diagram of an embodiment of the training process of the timbre prediction model in the embodiments of the present application;

[0024] Figure 4 It is a schematic diagram of an embodiment of obtaining the timbre vector of the singer to be matched in the embodiments of the present application;

[0025] Figure 5 It is a schematic diagram of another embodiment of the audio matching method in the embodiments of the present application;

[0026] Figure 6 It is a schematic diagram of an embodiment of the training process of the timbre sound effect conversion model in the embodiments of the present application;

[0027] Figure 7 It is a schematic diagram of an embodiment of the process of obtaining the training samples of the timbre sound effect conversion model in the embodiments of the present application;

[0028] Figure 8It is a refinement step of step 501 in the embodiment of the present application;

[0029] Figure 9 It is a schematic diagram of an embodiment in the training process of the timbre prediction model in the application embodiment;

[0030] Figure 10 It is a schematic diagram of an embodiment for obtaining the timbre vector of the singer to be matched in the embodiment of the present application;

[0031] Figure 11 It is a schematic diagram of an embodiment of the computer device in the embodiment of the present application. Detailed implementation manners

[0032] The embodiment of the present invention provides a sound effect matching method, device, storage medium and computer program product, which are used to determine the sound effect parameters adapted to the audio of the singer to be matched according to the audio of the singer to be matched and the timbre prediction model, thereby improving the convenience of determining the sound effect parameters of the audio of the singer to be matched.

[0033] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0034] The terms "first", "second", "third", "fourth", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order different from that shown or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0035] The embodiment of the present application provides a sound effect matching method, and the general principle of this sound effect matching method is as follows:

[0036] When it is necessary to set the sound effect parameters of the audio for the singer to be matched, the timbre feature extraction can be performed on the audio of the singer to be matched and the audio of at least two alternative singers through a pre-trained timbre prediction model, so as to obtain the timbre vector of the singer to be matched corresponding to the audio of the singer to be matched and at least two alternative timbre vectors of the alternative singers corresponding to the at least two alternative singer audios. Among them, the pre-trained timbre prediction model is trained by using contrastive learning and a triplet loss function; then the similarity between the timbre vector and each alternative timbre vector is determined respectively; finally, according to the similarity, the target timbre vector that matches the alternative timbre vector is determined from the at least two timbre vectors, and the sound effect parameters of the audio corresponding to the target timbre vector are determined as the sound effect parameters of the audio of the singer to be matched, thereby improving the convenience of setting the sound effect parameters of the audio of the singer to be matched.

[0037] In order to better implement the above audio matching method, an embodiment of the present application provides a sound effect matching system. Please refer to Figure 1 , Figure 1 which is a schematic architecture diagram of a sound effect matching system provided by an embodiment of the present application. The sound effect matching system may include at least one terminal device 101 and a server 102; different types of application programs may be installed on the terminal device 101. For example, a KTV application program, an instant messaging application program, a live broadcast application program, a conference communication application program, etc. may be installed on the terminal device 101; the terminal device 101 may be a smart phone, a tablet computer, a notebook computer, a desktop computer, an intelligent vehicle, etc. The server 102 may be used to store the audio data and image data generated by different types of application programs of the terminal device 101. The server 102 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms, etc.

[0038] Among them, the above sound effect matching method is executed by the terminal device 101 or the server 102. When the sound effect matching method is executed by the terminal device 101, the alternative singer audio that the terminal device 101 needs to obtain may be included in the server 102. When the terminal device 101 needs to set sound effect parameters for the audio of the singer to be matched, the terminal device 101 can obtain the alternative singer audio from the server 102, and then extract the timbre features of the audio of the singer to be matched and at least two alternative singer audios through a pre-trained timbre prediction model, obtaining the timbre vector of the singer to be matched corresponding to the audio of the singer to be matched and at least two alternative timbre vectors of the alternative singers corresponding to the at least two alternative singer audios. Among them, the pre-trained timbre prediction model is trained using contrastive learning and a triplet loss function; respectively determine the similarity between the timbre vector and each of the alternative timbre vectors; according to the similarity, determine a target timbre vector that matches the timbre vector from the at least two alternative timbre vectors; determine the sound effect parameters of the audio corresponding to the target timbre vector as the sound effect parameters of the audio of the singer to be matched, thereby improving the convenience of setting sound effect parameters for the audio of the singer to be matched.

[0039] For ease of understanding, the audio matching method in the embodiments of the present application will be described below. Please refer to Figure 2 , an embodiment of the audio matching method in the embodiments of the present application, includes:

[0040] 201. Extract the timbre features of the audio of the singer to be matched and at least two alternative singer audios through a pre-trained timbre prediction model, obtaining the timbre vector of the singer to be matched corresponding to the audio of the singer to be matched and at least two alternative timbre vectors of the alternative singers corresponding to the at least two alternative singer audios. Among them, the pre-trained timbre prediction model is trained using contrastive learning and a triplet loss function;

[0041] Because during the process of singing a song, if the timbres of the audio of two singers are similar, the same sound effect parameters (the sound effect parameters here include but are not limited to equalizer adjustment hyperparameters, reverb timing, dynamic compression, etc.) can be set for the audio of these two singers when they are singing the song. Therefore, based on this principle, the embodiments of the present application can determine a target alternative singer whose timbre is similar to that of the audio of the singer to be matched from multiple alternative singers in the sound effect pool, and determine the sound effect parameters of the target alternative singer audio as the sound effect parameters of the audio of the singer to be matched.

[0042] Specifically, when obtaining the timbre vectors of the singer audio to be matched and at least two alternative singer audios, it can be based on a pre-trained timbre prediction model to generate the corresponding timbre vectors. For example, when inputting the singer audio to be matched into the timbre prediction model, the timbre vector of the singer audio to be matched output by the pre-trained timbre prediction model can be obtained. And when inputting multiple audios of the alternative singer audios into the pre-trained timbre prediction model, multiple timbre vectors corresponding to the multiple alternative singer audios output by the pre-trained timbre prediction model can be obtained. Among them, the pre-trained timbre prediction model in the embodiments of the present application is pre-trained based on contrastive learning and triplet loss function.

[0043] Further, in order to improve the processing speed of the pre-trained timbre prediction model for audio data, the embodiments of the present application can also perform the following steps:

[0044] Obtain the spectral feature vector of the singer audio to be matched and at least two spectral feature vectors of at least two alternative singer audios, and input the spectral feature vector of the singer audio to be matched and at least two spectral feature vectors of at least two alternative singer audios into the pre-trained timbre prediction model, so as to obtain the timbre vector of the singer audio to be matched output by the pre-trained timbre prediction model and at least two timbre vectors of at least two alternative singer audios.

[0045] Since the pre-trained timbre prediction model in the embodiments of the present application receives the spectral feature vector of the audio rather than the audio data itself, it is equivalent to reducing the amount of data processed by the pre-trained timbre prediction model and improving the data processing efficiency of the model. Among them, when extracting the spectral feature vector of the audio, it can be the Mel spectrum or the feature matrix of multiple songs extracted through audio vector technology.

[0046] Further, the process of training the timbre prediction model using contrastive learning and triplet loss function will be described in the following embodiments and will not be elaborated here.

[0047] 202. Determine the similarity between the timbre vector of the singer to be matched and the alternative timbre vectors of each alternative singer respectively;

[0048] After obtaining the timbre vector of the singer to be matched and at least two alternative timbre vectors of at least two alternative singers, the similarity between the timbre vector of the singer to be matched and at least two alternative timbre vectors of at least two alternative singers can be calculated. Among them, the method for calculating the similarity can be to calculate the similarity based on the cosine value between two vectors or to calculate the similarity based on the distance between two vectors. The method for calculating the similarity is not specifically limited here.

[0049] 203. Determine a target timbre vector that matches the timbre vector of the singer to be matched from at least two alternative timbre vectors according to the similarity.

[0050] After obtaining the parameter that characterizes the similarity between the timbre vector of the singer to be matched and at least two alternative timbre vectors of at least two alternative singers, a target timbre vector that is similar to the timbre vector of the singer to be matched can be determined from at least two alternative timbre vectors of at least two alternative singers based on this parameter and a preset threshold.

[0051] Specifically, the preset threshold here can be custom-set according to the number of target alternative singers required. For example, when the number of target alternative singers required is small, the preset threshold can be set larger, and when the number of target alternative singers required is large, the preset threshold can be set smaller. There is no specific limitation on the size of the preset threshold here.

[0052] 204. Determine the sound effect parameters of the audio corresponding to the target timbre vector as the sound effect parameters of the audio of the singer to be matched.

[0053] After obtaining the target timbre vector, based on the principle that similar timbre leads to similar sound effect parameters, the sound effect parameters of the audio corresponding to the target timbre vector can be directly determined as the sound effect parameters of the audio of the singer to be matched, thus improving the convenience of setting the matching sound effect parameters for the audio of the singer to be matched.

[0054] In the embodiments of the present application, the timbre vector of the audio of the singer to be matched and at least two timbre vectors of at least two alternative singers can be directly obtained through a timbre prediction model. Then, based on the similarity between the timbre vector of the audio of the singer to be matched and at least two alternative timbre vectors of at least two alternative singers, a target timbre vector similar to the timbre vector of the audio of the singer to be matched is determined. Then, the sound effect parameters of the audio corresponding to the target timbre vector are set as the sound effect parameters of the audio of the singer to be matched, thus avoiding the problem of manually adjusting the sound effect parameters in the prior art and improving the convenience of setting the sound effect parameters of the audio of the singer to be matched.

[0055] It is easy to understand that based on Figure 2 the embodiments of, before using the pre-trained timbre prediction model to extract the timbre features of the audio of the singer to be matched and at least two alternative singers, the initialized timbre prediction model also needs to be trained. The process of training the initialized timbre prediction model is described below. Please refer to Figure 3 :

[0056] 301. Obtain anchor samples, positive samples, and negative samples. Among them, the anchor samples include the spectral feature vectors of the target song audio sung by the singer to be matched, the positive samples include the spectral feature vectors of other song audios sung by the singer to be matched that are different from the target song, and the negative samples include the spectral feature vectors of any song sung by other singers different from the singer to be matched;

[0057] Since the timbre prediction model is trained using contrastive learning and triplet loss function, the process of training the timbre prediction model through contrastive learning and triplet loss function is described below:

[0058] Specifically, when training the initialized timbre prediction model through contrastive learning, it is generally necessary to first determine the anchor samples, positive samples, and negative samples, so that after the timbre prediction model is trained, the distance between the output anchor samples and positive samples is much smaller than the distance between the anchor samples and negative samples, thus completing the training of the initialized timbre prediction model.

[0059] Among them, the anchor samples in the embodiments of the present application include the spectral feature vectors of the target song audio sung by the singer to be matched, the positive samples include the spectral feature vectors of other song audios sung by the singer to be matched that are different from the target song, and the negative samples are the spectral feature vectors of any song sung by other singers. For example, when the anchor sample is the spectral feature vector of the audio of "Little Donkey" sung by Zhang San, the positive samples are the spectral feature vectors of the audios of "Cinderella" and "Autumn Porcelain" sung by Zhang San, and the negative samples can be the spectral feature vectors of the audios of "Little Donkey", "Cinderella", or "Autumn Porcelain" sung by Li Si. Of course, the above examples are only explanations of the anchor samples, positive samples, and negative samples rather than limitations.

[0060] 302. Input the anchor samples into the initialized timbre prediction model to obtain the timbre vectors of the anchor samples output by the initialized timbre prediction model;

[0061] For the convenience of distinction, in the embodiments of the present application, the timbre prediction model before training is called the initialized timbre prediction model, and the model after training is called the pre-trained timbre prediction model.

[0062] Specifically, when training the initialized timbre prediction model, the anchor samples are input into the initialized timbre prediction model to obtain the timbre vectors of the anchor samples output by the initialized timbre prediction model. Among them, the timbre prediction model here can be a convolutional neural network model, a recurrent neural network model, a Transformer model, etc. The type of the timbre prediction model is not specifically limited here.

[0063] 303. Input the positive samples into the initialized timbre prediction model to obtain the timbre vectors of the positive samples output by the initialized timbre prediction model;

[0064] Similar to step 302, when a positive sample is input into the initialized timbre prediction model, a timbre vector of the positive sample output by the initialized timbre prediction model is obtained.

[0065] 304. Input a negative sample into the initialized timbre prediction model to obtain a timbre vector of the negative sample output by the initialized timbre prediction model;

[0066] Similar to step 302, when a negative sample is input into the initialized timbre prediction model, a timbre vector of the negative sample output by the initialized timbre prediction model is obtained.

[0067] 305. Calculate a loss amount according to the timbre vector of the anchor sample, the timbre vector of the positive sample, the timbre vector of the negative sample, and the triplet loss function;

[0068] Specifically, when training the initialized timbre prediction model in the embodiments of the present application, the triplet loss function is used to calculate the loss amount, and the triplet loss function is as follows:

[0069] L(a,p,n)=max(d(a,p)-d(a,n)+margin,0);

[0070] Wherein, d(a,p) is the distance between the timbre vector of the anchor sample and the timbre vector of the positive sample, d(a,n) is the distance between the timbre vector of the anchor sample and the timbre vector of the negative sample, and margin is a hyperparameter used to adjust the threshold of the distance between the positive sample and the negative sample.

[0071] 306. Train the initialized timbre prediction model according to the loss amount and the backpropagation algorithm to obtain a pre-trained timbre prediction model.

[0072] After obtaining the loss amount calculated by the above triplet loss function, the initialized timbre prediction model is trained by using the loss amount and the backpropagation algorithm, that is, the parameters of the initialized timbre prediction model are adjusted until a pre-trained timbre prediction model is obtained.

[0073] In the embodiments of the present application, the initialized timbre prediction model is trained by using contrastive learning and the triplet loss function, and the process of training the initialized timbre prediction model by using contrastive learning has at least the following advantages:

[0074] 1. Efficiently utilize data: Because contrastive learning does not require a large number of labels, only a small amount of paired data is needed to train the initialized timbre prediction model, that is, the number of labels is reduced;

[0075] 2. Strong robustness: Since contrastive learning focuses on the relative relationships between samples, it has a high tolerance for small noises and small errors, improving the generalization ability of the model.

[0076] 3. Good generalization: Since the model does not need to directly predict the target value but learns how to distinguish samples, it helps the model understand and capture the essence of the data.

[0077] Furthermore, based on Figure 2 the embodiments described above, since the timbre of the singer to be matched is different when singing different audio, and in order to accurately obtain the overall timbre characteristics of the singer, the embodiments of the present application can perform the following steps to improve the accuracy of the timbre vector of the audio of the singer to be matched. Please refer to Figure 4 :

[0078] 401. Perform preprocessing on the audio of the singer to be matched to obtain the preprocessed audio of the singer to be matched, where the preprocessing includes extracting the human voice audio from the audio, removing the silent segments in the human voice audio, and splicing the human voice audio after removing the silent segments.

[0079] Since when extracting timbre characteristics, it is more desired to obtain the dry voice (i.e., the voice without accompaniment) of the audio of the singer to be matched, the embodiments of the present application can first perform preprocessing on the audio of the singer to be matched and the multiple alternative singer audios before obtaining the spectral feature vector of the audio of the singer to be matched and the multiple spectral feature vectors of the multiple alternative singer audios, to obtain the preprocessed audio of the singer to be matched and the preprocessed multiple alternative singer audios, where the preprocessing here can be extracting the human voice audio from the audio, removing the silent segments in the human voice audio, and splicing the human voice audio after removing the silent segments.

[0080] Specifically, the method for extracting the human voice audio from the audio can be: spectral subtraction or deep learning method. When using spectral subtraction to extract the human voice audio, first calculate the total spectrum of the song, and then subtract the frequency components of the accompaniment from the total spectrum to obtain the human voice part. When using a deep learning method to extract the human voice audio, it can be to use a neural network (such as models like U-net, Wave-U-Net, etc.) to train the audio so that when the entire song is input, the model outputs the human voice signal. The process and method for obtaining the human voice audio are not specifically limited here.

[0081] Further, when identifying the silent segments in the human voice audio, a noise threshold detection method or a short-time energy analysis method can be used. For example, when using the noise threshold detection method, the average value of the audio signal is calculated by this method, and a threshold is set. When the signal is lower than the threshold for a certain time length (such as 0.5 s), it is considered a silent segment. For the short-time energy analysis method, the energy of the audio frames is calculated in segments. If the energy of several consecutive frames is less than a certain threshold, it is identified as a silent segment. Here, there is no specific limitation on the method for identifying the silent segments in the human voice audio.

[0082] Further, after identifying the silent segments in the human voice audio, the silent segments in the human voice audio are removed, and the human voice audio after removing the silent segments is spliced together to form the human voice audio.

[0083] 402. Obtain multiple spectral feature vectors of multiple audios of the singer to be matched after preprocessing;

[0084] Further, in order to improve the integrity and accuracy of obtaining the timbre features of the singer to be matched, the embodiments of the present application can also obtain multiple spectral feature vectors of multiple audios of the singer to be matched after preprocessing. For example, the audios of m (such as 6, 7, or 10) songs sung by the singer to be matched can be obtained, and then the feature matrices of multiple songs are extracted by using the Mel spectrum or through audio vector technology.

[0085] As an optional embodiment, when extracting the feature matrix of the audio by using the Mel spectrum, the waveform of the audio can be transformed through short-time Fourier transform and Mel spectrum conversion to obtain the Mel spectrum of the audio.

[0086] 403. Input the multiple spectral feature vectors of multiple audios of the singer to be matched into the pre-trained timbre prediction model respectively to obtain multiple timbre vectors of multiple audios of the singer to be matched output by the pre-trained timbre prediction model;

[0087] After obtaining the multiple spectral feature vectors of multiple audios of the singer to be matched, the multiple spectral feature vectors of multiple audios of the singer to be matched are input into the pre-trained timbre prediction model respectively to obtain multiple timbre vectors of multiple audios of the singer to be matched output by the pre-trained timbre prediction model.

[0088] 404. Screen out the first number of outlier timbre vectors from the multiple timbre vectors of the singer to be matched to obtain the second number of timbre vectors;

[0089] In order to improve the overall integrity and accuracy of obtaining the timbre of the singer to be matched, the embodiments of the present application can screen out the first number of outlier timbre vectors from the multiple timbre vectors of the singer to be matched to obtain the second number of timbre vectors.

[0090] Specifically, when screening out the first quantity of outlier timbre vectors, the average value and standard deviation of multiple timbre vectors can be calculated, and then the Z-score of each timbre vector is calculated, where the Z-score represents the degree of its deviation from the average value. Then, all the Z-scores are sorted in descending order according to their absolute values, and the vectors corresponding to the two largest Z-scores are selected as relative outliers, and the timbre vectors corresponding to the relative outliers are deleted from the multiple timbre vectors.

[0091] 405. Calculate the timbre vector of the singer audio to be matched according to the second quantity of timbre vectors.

[0092] After obtaining the second quantity of timbre vectors, the timbre vector of the singer audio to be matched is calculated according to the second quantity of timbre vectors.

[0093] For example, after obtaining the second quantity of timbre vectors, the average value of the second quantity of timbre vectors can be calculated, and this average value is determined as the timbre vector of the singer audio to be matched.

[0094] In the embodiment of the present application, in order to improve the overall integrity and accuracy of the timbre of the singer audio to be matched, the human voice audio is first extracted from the singer audio to be matched, and then the timbre vectors of multiple songs sung by the singer to be matched are obtained. Then, the first quantity of discrete points are screened out from the timbre vectors of multiple songs, and the average value of the second quantity of timbre vectors after screening out the discrete points is determined as the timbre vector of the singer audio to be matched, thereby improving the integrity and accuracy of the timbre vector of the singer audio to be matched.

[0095] Next, another embodiment of the audio matching method in the embodiment of the present application will be described. Please refer to Figure 5 , another embodiment of the audio matching method in the embodiment of the present application, includes:

[0096] 501. Perform timbre and sound effect conversion on the timbre vector of the singer audio to be matched and the timbre vectors of at least two alternative singer audios through a pre-trained timbre and sound effect conversion model, to obtain the timbre and sound effect vector of the singer audio to be matched and at least two alternative timbre and sound effect vectors of the at least two alternative singer audios, where the pre-trained timbre and sound effect conversion model is trained using contrastive learning and a triplet loss function;

[0097] In Figure 2In an embodiment, when setting sound effect parameters for a singer's audio, it is generally considered that if the timbres of two singers' audios are similar, then the sound effect parameters of these two singers' audios are considered similar. However, in an actual scenario, the accuracy of this theory may not be high. In order to further improve the accuracy of setting sound effect parameters for the audio of a singer to be matched, the embodiment of the present application may also determine a target timbre-sound effect vector similar to the timbre-sound effect vector of the singer to be matched based on a timbre-volume vector (that is, an intermediate vector based on a timbre vector and a sound effect vector), and set the sound effect parameters of the audio corresponding to the target timbre-sound effect vector as the sound effect parameters of the audio of the singer to be matched.

[0098] Specifically, before generating the timbre-sound effect vectors of the audio of the singer to be matched and at least two alternative singer audios, the embodiment of the present application needs to first obtain the timbre vector of the audio of the singer to be matched and at least two alternative timbre vectors of at least two alternative singer audios in the sound effect pool, and based on the timbre vector of the audio of the singer to be matched and the at least two alternative timbre vectors of at least two alternative singer audios in the sound effect pool, obtain the timbre-sound effect vector of the audio of the singer to be matched and the at least two alternative timbre-sound effect vectors of at least two alternative singer audios. The process of how to obtain the timbre vector of the audio of the singer to be matched and the multiple timbre vectors of multiple alternative singer audios in the sound effect pool will be described in the following embodiments and will not be elaborated here.

[0099] Further, after obtaining the timbre vector of the audio of the singer to be matched and the at least two alternative timbre vectors of at least two alternative singer audios, the timbre vector of the audio of the singer to be matched and the at least two alternative timbre vectors of at least two alternative singer audios are input into a timbre-sound effect conversion model to obtain the timbre-sound effect vector of the audio of the singer to be matched output by the timbre-sound effect conversion model, and the at least two alternative timbre-sound effect vectors of at least two alternative singer audios. Here, the timbre-sound effect conversion model may be a neural convolutional network model or a recurrent neural network model, etc. The type of the timbre-sound effect conversion model is not specifically limited here.

[0100] Further, the timbre-sound effect conversion model in the embodiment of the present application is trained according to contrastive learning and a triplet loss function. The training process of the timbre-sound effect conversion model will be described in the following embodiments and will not be elaborated here.

[0101] 502. Determine the similarity between the timbre-sound effect vector and each alternative timbre-sound effect vector respectively;

[0102] After obtaining the timbre and sound effect vector of the singer audio to be matched and at least two alternative timbre and sound effect vectors of at least two alternative singer audios, the similarity between the timbre and sound effect vector of the singer to be matched and each alternative timbre and sound effect vector is determined respectively. Here, the method for calculating the similarity can be the cosine of the included angle method or the vector distance method, etc. There is no specific limitation on the method for calculating the similarity here.

[0103] 503. Determine a target timbre and sound effect vector that matches the timbre and sound effect vector of the singer to be matched from at least two alternative timbre and sound effect vectors according to the similarity.

[0104] After obtaining at least two timbre and sound effect vectors of at least two alternative singer audios, a target timbre and sound effect vector that matches the timbre and sound effect vector of the singer audio to be matched is determined from at least two alternative timbre and sound effect vectors of at least two alternative singer audios.

[0105] Specifically, in the process of screening the target timbre and sound effect vector, it can be screened based on the similarity and the first threshold, or the similarity and the second threshold, and the size of the threshold can be customized according to actual needs. There is no specific limitation on the size of the threshold here.

[0106] 504. Determine the sound effect parameters of the audio corresponding to the target timbre and sound effect vector as the sound effect parameters of the singer audio to be matched.

[0107] After obtaining a target timbre and sound effect vector that matches the timbre and sound effect vector of the singer audio to be matched, the sound effect parameters of the audio corresponding to the target timbre and sound effect vector are further determined, and the sound effect parameters of the audio corresponding to the target timbre and sound effect vector are determined as the sound effect parameters of the singer audio to be matched.

[0108] Specifically, when obtaining the sound effect parameters of the audio corresponding to the target timbre and sound effect vector, the audio of the target alternative singer corresponding to the target timbre and sound effect vector stored in advance in the sound effect pool can be obtained, and then the sound effect parameters set for the target singer audio are determined as the sound effect parameters of the singer audio to be matched.

[0109] In the embodiment of the present application, it is not based on the timbre vector to determine the target timbre vector that is most similar to the timbre of the singer to be matched, but based on the intermediate parameter timbre and sound effect vector of the timbre vector and the sound effect vector to determine the target timbre and sound effect vector that is most similar to the timbre and sound effect vector of the singer audio to be matched, and further obtain the audio of the target alternative singer corresponding to the target timbre and sound effect vector, and then determine the sound effect parameters of the target alternative singer audio as the sound effect parameters of the singer audio to be matched, thereby further improving the accuracy of the sound effect parameters set for the singer audio to be matched.

[0110] Based on Figure 5In the described embodiments, before performing step 501, it is also necessary to train the initialized timbre and sound effect conversion model. The following describes the training process of the initialized timbre and sound effect conversion model. Please refer to Figure 6 :

[0111] 601. Obtain the training samples of the initialized timbre and sound effect conversion model. The training samples include the first anchor sample, the first positive sample, and the first negative sample;

[0112] When training the initialized timbre and sound effect conversion model, it is necessary to first obtain the training samples of the initialized timbre and sound effect conversion model. Since the timbre and sound effect conversion model here is trained using contrastive learning and the triplet loss function, the training samples that need to be obtained in the embodiments of the present application include the first anchor sample, the first positive sample, and the first negative sample. Among them, the first anchor sample here is the timbre vector of the anchor singer (the anchor singer here is any singer in the sound effect pool), and the first positive sample is the timbre vector similar to the anchor sample, and the first negative sample is the timbre vector not similar to the anchor sample.

[0113] Regarding the process of how to obtain the first anchor sample, the first positive sample, and the first negative sample, it will be described in the following embodiments and will not be elaborated here.

[0114] 602. Input the first anchor sample into the initialized timbre and sound effect conversion model to obtain the timbre and sound effect vector of the first anchor sample output by the initialized timbre and sound effect conversion model;

[0115] After obtaining the training samples, according to the training process of contrastive learning, input the first anchor sample into the initialized timbre and sound effect conversion model (the initialized timbre and sound effect conversion model here is the timbre and sound effect conversion model before training) to obtain the timbre and sound effect vector corresponding to the first anchor sample output by the initialized timbre and sound effect conversion model.

[0116] 603. Input the first positive sample into the initialized timbre and sound effect conversion model to obtain the timbre and sound effect vector of the first positive sample output by the initialized timbre and sound effect conversion model;

[0117] Similar to step 602, input the first positive sample into the initialized timbre and sound effect conversion model to obtain the timbre and sound effect vector of the first positive sample output by the initialized timbre and sound effect conversion model.

[0118] 604. Input the first negative sample into the initialized timbre and sound effect conversion model to obtain the timbre and sound effect vector of the first negative sample output by the initialized timbre and sound effect conversion model;

[0119] Similar to step 602, input the first negative sample into the initialized timbre and sound effect conversion model to obtain the timbre and sound effect vector of the first negative sample output by the initialized timbre and sound effect conversion model.

[0120] 605. Calculate the loss value according to the timbre and sound effect vector of the first anchor sample, the timbre and sound effect vector of the first positive sample, the timbre and sound effect vector of the first negative sample, and the triplet loss function;

[0121] After obtaining the timbre and sound effect vector of the first anchor sample, the timbre and sound effect vector of the first positive sample, and the timbre and sound effect vector of the first negative sample, calculate the loss value according to the triplet loss function. Specifically, the triplet loss function is as follows:

[0122] L(a,p,n)=max(d(a,p)-d(a,n)+margin,0);

[0123] Among them, d(a,p) is the distance between the timbre and sound effect vector of the first anchor sample and the timbre and sound effect vector of the first positive sample, d(a,n) is the distance between the timbre and sound effect vector of the first anchor sample and the timbre and sound effect vector of the first negative sample, and margin is a hyperparameter used to adjust the threshold of the distance between the first positive sample and the first negative sample.

[0124] 606. Train the initialized timbre and sound effect conversion model according to the loss value and the backpropagation algorithm to obtain a pre-trained timbre and sound effect conversion model.

[0125] After obtaining the loss value and the backpropagation algorithm, train the initialized timbre and sound effect conversion model according to the loss value and the backpropagation algorithm to obtain a pre-trained timbre and sound effect conversion model. As for the specific training process, it is similar to the description in the prior art and will not be elaborated here.

[0126] In the embodiments of the present application, the training process of the timbre and sound effect conversion model using contrastive learning and the triplet loss function is described in detail. The process of training the timbre and sound effect conversion model using contrastive learning has at least the following advantages:

[0127] 1. Efficient use of data: Because contrastive learning does not require a large number of labels, only a small amount of paired data is needed to train the initialized timbre and sound effect conversion model, that is, the number of labels is reduced;

[0128] 2. Strong robustness: Because contrastive learning focuses on the relative relationship between samples, it has a high tolerance for small noise and small errors, improving the generalization ability of the model;

[0129] 3. Good generalization: Because the model does not need to directly predict the target value, but learns how to distinguish samples, it helps the model understand and capture the essence of the data.

[0130] Based on Figure 6 the above-described embodiments, the process of obtaining training samples for the timbre and sound effect conversion model will be described below. Please refer to Figure 7 ;

[0131] 701. Calculate multiple similarities between the sound effect vectors of the anchor singer's audio and multiple alternative sound effect vectors of multiple alternative singer audios;

[0132] Specifically, when determining the first anchor sample, the first positive sample, and the first negative sample, one singer can be randomly determined from the sound effect pool, and this singer is determined as the anchor singer. Then, calculate the sound effect vector of the anchor singer's audio, and calculate multiple similarities between the sound effect vector of the anchor singer's audio and multiple sound effect vectors of multiple different singer audios in the sound effect pool.

[0133] Specifically, when determining the sound effect vector of any singer in the sound effect pool, the sound effect parameters set for the audio of this anchor singer and the audios of each different singer can be obtained from the sound effect pool. The sound effect parameters here include at least one of the equalizer adjustment hyperparameters, reverb hyperparameters, and dynamic compression hyperparameters, and a sound effect vector is generated corresponding to the sound effect parameters. For example, after obtaining the equalizer adjustment hyperparameters, reverb hyperparameters, and dynamic compression hyperparameters, the above parameters can be mapped to corresponding vectors according to a certain mapping rule or a specific mapping function.

[0134] After obtaining the sound effect vector of the anchor singer's audio and multiple sound effect vectors of multiple different singer audios in the sound effect pool, the similarity between the vectors can be calculated according to the cosine similarity or the distance between the vectors. The process and method of calculating the vector similarity are not specifically limited here.

[0135] 702. Determine multiple first alternative sound effect vectors similar to the sound effect vector of the anchor singer's audio according to multiple similarities and a preset first threshold, where the preset first threshold is used to characterize the similarity between the sound effect vector and the alternative sound effect vector;

[0136] After obtaining multiple similarities between the sound effect vector of the anchor singer's audio and multiple sound effect vectors of multiple different singer audios in the sound effect pool, multiple similar sound effect vectors of multiple different singer audios similar to the sound effect vector of the anchor singer's audio can be determined from the sound effect pool according to the similarity and a preset first threshold. Here, the first threshold is used to characterize the similarity between the sound effect vector of the anchor singer's audio and the sound effect vectors of other singer audios.

[0137] Since the first threshold here is used to determine multiple similar sound effect vectors similar to the sound effect vector of the anchor singer from the sound effect pool, the number of similar sound effect vectors can be selected by changing the value of the first threshold.

[0138] 703. Determine multiple second alternative sound effect vectors that are dissimilar to the sound effect vector of the anchor singer audio according to multiple similarity degrees and a preset second threshold, where the preset second threshold is used to represent the dissimilarity between the sound effect vector and the alternative sound effect vector;

[0139] After obtaining the multiple similarity degrees between the sound effect vector of the anchor singer audio and the multiple sound effect vectors of the audio of each different singer in the sound effect pool, multiple second alternative sound effect vectors that are dissimilar to the sound effect vector of the anchor singer audio can also be determined from the sound effect pool according to the multiple similarity degrees and the preset second threshold. Here, the preset second threshold is used to represent the dissimilarity, and the number of dissimilar sound effect vectors can also be selected by changing the size of the second threshold.

[0140] 704. Determine the timbre vector of the anchor singer audio as the first anchor sample;

[0141] After respectively determining multiple first alternative sound effect vectors that are similar to the sound effect vector of the anchor singer audio and multiple second alternative sound effect vectors that are dissimilar to the sound effect vector of the anchor singer audio in steps 702 - 703, the timbre vector of the anchor singer audio is determined as the first anchor sample.

[0142] 705. Determine the timbre vectors of the first alternative singer audios corresponding to the multiple first alternative sound effect vectors as the first positive samples;

[0143] After obtaining the anchor singer, the timbre vectors of the first alternative singer audios corresponding to the multiple first alternative sound effect vectors are determined as the first positive samples.

[0144] 706. Determine the timbre vectors of the second alternative singer audios corresponding to the multiple second alternative sound effect vectors as the first negative samples.

[0145] After obtaining the anchor singer, the timbre vectors of the second alternative singer audios corresponding to the multiple second alternative sound effect vectors are determined as the first negative samples.

[0146] In the embodiments of the present application, the process of determining the first anchor sample, the first positive sample, and the first negative sample for training the timbre and sound effect conversion model is described in detail. In the embodiments of the present application, the similarity between the sound effect vector of the anchor singer and the sound effect vectors of multiple different singers in the sound effect pool is used to determine the first alternative sound effect vector similar to the sound effect vector of the anchor singer and the second alternative sound effect vector not similar to the sound effect vector of the anchor singer. The timbre vector of the first alternative singer audio corresponding to the first alternative sound effect vector is regarded as the first positive sample, the timbre vector of the second alternative singer audio corresponding to the second alternative sound effect vector is determined as the first negative sample, and the timbre and sound effect conversion model is trained using the first anchor sample, the first positive sample, and the first negative sample. Therefore, the timbre and sound effect conversion model trained in the embodiments of the present application can learn the conversion logic from the timbre vector to the timbre and sound effect vector, and correspondingly improve the accuracy of converting the timbre vector into the timbre and sound effect vector.

[0147] Based on Figure 5 the above-described embodiments, the process of obtaining the timbre vector of the singer audio to be matched and the timbre vectors of multiple alternative singer audios in the sound effect pool in step 501 will be described in detail below. Please refer to Figure 8 , Figure 8 which is a refinement step of step 501:

[0148] 801. Obtain the spectral feature vector of the singer audio to be matched and at least two alternative spectral feature vectors of at least two alternative singer audios;

[0149] Specifically, after the embodiments of the present application obtain the audio of the singer to be matched and the audios of at least two alternative singers, in order to improve the accuracy of audio data processing and the convenience of audio data processing, the spectral feature vector of the singer audio to be matched and at least two alternative spectral feature vectors of at least two alternative singer audios can be further extracted from the audio of the singer to be matched and the audios of at least two alternative singers, so as to reduce the amount of data input to the pre-trained timbre prediction model and improve the training speed and training efficiency of the model.

[0150] Further, when extracting the spectral feature vectors of each audio, the feature matrix of multiple songs can be extracted by using the Mel spectrum or through audio vector technology.

[0151] As an alternative embodiment, when using the Mel spectrum to extract the feature matrix of the audio, the waveform of the audio can be transformed through short-time Fourier transform and Mel spectrum to obtain the Mel spectrum of the audio.

[0152] 802. Input the spectral feature vector of the singer audio to be matched and at least two alternative spectral feature vectors of at least two alternative singer audios into the pre-trained timbre prediction model respectively, so as to obtain the timbre vector of the singer audio to be matched output by the pre-trained timbre prediction model and at least two alternative timbre vectors of at least two alternative singer audios.

[0153] After obtaining the spectral feature vector of the singer audio to be matched and at least two alternative spectral feature vectors of at least two alternative singer audios, the spectral feature vector of the singer audio to be matched and at least two alternative spectral feature vectors of at least two alternative singer audios can be input into the pre-trained timbre prediction model respectively, so as to obtain the timbre vector of the singer audio to be matched output by the pre-trained timbre prediction model and at least two alternative timbre vectors of at least two alternative singer audios. Among them, the process of obtaining the timbre vector of the singer audio to be matched and at least two alternative timbre vectors of at least two alternative singer audios by using the pre-trained timbre prediction model is similar to Figure 2 the description in step 201 in [reference], and they can be referred to each other, so it will not be elaborated here.

[0154] In the embodiment of the present application, the process of obtaining the timbre vector of the singer audio to be matched and the timbre vectors of multiple alternative singer audios is described in detail. And in the embodiment of the present application, the timbre vectors of the singer to be matched and the timbre vectors of multiple alternative singers are directly obtained through the timbre prediction model, thereby improving the convenience of the process of obtaining the timbre vector.

[0155] Furthermore, before using the timbre prediction model, it is also necessary to train the initialized timbre prediction model. The training process of the initialized timbre model is described below. Please refer to Figure 9 :

[0156] 901. Obtain the second anchor sample, the second positive sample and the second negative sample. Among them, the second anchor sample includes the spectral feature vector of the singer to be matched singing the target song audio, the second positive sample includes the spectral feature vector of the singer to be matched singing other songs different from the target song, and the second negative sample includes the spectral feature vector of other singers different from the singer to be matched singing any song;

[0157] 902. Input the second anchor sample into the initialized timbre prediction model to obtain the timbre vector of the second anchor sample output by the initialized timbre prediction model;

[0158] 903. Input the second positive sample into the initialized timbre prediction model to obtain the timbre vector of the second positive sample output by the initialized timbre prediction model;

[0159] 904. Input the second negative sample into the initialized timbre prediction model to obtain the timbre vector of the second negative sample output by the initialized timbre prediction model;

[0160] 905. Calculate the loss value according to the timbre vector of the second anchor sample, the timbre vector of the second positive sample, the timbre vector of the second negative sample, and the triplet loss function;

[0161] 906. Train the initialized timbre prediction model according to the loss value and the backpropagation algorithm to obtain a pre-trained timbre prediction model.

[0162] Among them, the description process of steps 901 to 906 is similar to the description of steps 301 to 306 in Figure 3 the embodiment and can be referred to each other, and will not be elaborated here.

[0163] In the embodiment of the present application, the process of training the timbre prediction model is described in detail, and this process uses contrastive learning and the triplet loss function for training, thereby improving the convenience of the process of training the timbre prediction model.

[0164] Based on Figure 9 the above-mentioned embodiment, in order to better extract the overall timbre characteristics of the singer to be matched, the following steps can also be performed on the audio of the singer to be matched. Please refer to Figure 10 :

[0165] 1001. Obtain the audio of the singer to be matched and the audio of multiple candidate singers;

[0166] It is easy to understand that before performing preprocessing on the audio of the singer to be matched and the audio of multiple candidate singers, it is necessary to first obtain the audio of the singer to be matched and the audio of multiple candidate singers. Among them, when obtaining the audio of the singer to be matched and the audio of multiple candidate singers, it can be obtained from the sound effect pool pre-stored in the terminal or server, or obtained from the local terminal. The method of obtaining the audio of the singer to be matched and the audio of multiple candidate singers is not specifically limited here.

[0167] 1002. Perform preprocessing on the audio of the singer to be matched and the audio of multiple candidate singers to obtain the preprocessed audio of the singer to be matched and the preprocessed audio of multiple candidate singers. Among them, the preprocessing includes extracting the human voice audio from the audio, removing the silent segments in the human voice audio, and splicing the human voice audio after removing the silent segments.

[0168] 1003. Obtain multiple spectral feature vectors of multiple audio pieces of the singer to be matched after preprocessing;

[0169] 1004. Input the multiple spectral feature vectors of multiple pieces of audio of the singer to be matched after preprocessing into the pre-trained timbre prediction model respectively, so as to obtain multiple timbre vectors of the singer to be matched output by the pre-trained timbre prediction model;

[0170] 1005. Screen out the first number of outlier timbre vectors from the multiple timbre vectors of the singer to be matched, so as to obtain the second number of timbre vectors;

[0171] 1006. Calculate the timbre vector of the audio of the singer to be matched according to the second number of timbre vectors.

[0172] Steps 1002 to 1006 are similar to the descriptions of steps 401 to 405 in Figure 4 the embodiment, and will not be elaborated here.

[0173] In the embodiment of the present application, in order to improve the integrity and accuracy of obtaining the timbre of the singer to be matched, the human voice audio can be extracted from the audio of the singer to be matched first, and then the timbre vectors of multiple songs sung by the singer to be matched are obtained, and the first number of discrete points are screened out from the timbre vectors of multiple songs, and the average value of the second number of timbre vectors after screening out the discrete points is determined as the timbre vector of the audio of the singer to be matched, thereby improving the integrity and accuracy of obtaining the timbre vector of the audio of the singer to be matched.

[0174] It can be understood that in various embodiments of the present invention, the magnitudes of the serial numbers of the above steps do not mean the sequence of execution, and the execution sequence of each step should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0175] The embodiment of the present application also provides a computer program product, storing a computer program, which when executed by a processor is used to implement each step in the above method embodiment.

[0176] The computer device in the embodiment of the present invention will be described from the perspective of hardware processing below. Please refer to Figure 11 :

[0177] One embodiment of the computer device in the embodiment of the present invention includes:

[0178] One or more central processing units (CPUs) 1101 and a memory 1105, and one or more application programs or data are stored in the memory 1105.

[0179] Among them, the memory 1105 can be volatile storage or persistent storage. The program stored in the memory 1105 can include one or more modules, and each module can include a series of instruction operations on the server. Further, the central processing unit 1101 can be configured to communicate with the memory 1105 and execute a series of instruction operations in the memory 1105 on the computer device.

[0180] The computer device may further include one or more power supplies 1102, one or more wired or wireless network interfaces 1103, one or more input / output interfaces 1104, and / or one or more operating systems, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0181] The central processing unit 1101 can execute each step in the above method embodiments.

[0182] When the processor in the computer device described above executes the computer program, it can also implement the functions of each unit in the corresponding device embodiments described above, which will not be elaborated here. Exemplarily, the computer program can be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the computer device. For example, the computer program can be divided into the respective units in the above computer device, and each unit can implement the specific functions as described in the corresponding computer device description above.

[0183] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the processor and the memory are only examples of the computer device and do not constitute a limitation on the computer device. It may include more or fewer components, or combine certain components, or different components. For example, the computer device may further include input / output devices, network access devices, a bus, etc.

[0184] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the computer device, and connects various parts of the entire computer device through various interfaces and lines.

[0185] The memory can be used to store the computer program and / or modules. The processor realizes various functions of the computer device by running or executing the computer program and / or modules stored in the memory, and by calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.

[0186] The present invention also provides a computer-readable storage medium for implementing the functions of a computer device, on which a computer program is stored. When the computer program is executed by a processor, the processor can be used to execute each step in the above method embodiments.

[0187] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, indirect couplings or communication connections of devices or units, and can be in electrical, mechanical or other forms.

[0188] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0189] In addition, each functional unit in various embodiments of the present invention may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0190] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0191] As described above, the above embodiments are only used to illustrate the technical solution of the present invention and are not intended to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of the present invention.

Claims

1. A sound effect matching method, characterized in that: include: The timbre features of the to-be-matched singer audio and at least two candidate singer audios are extracted by a pre-trained timbre prediction model to obtain a timbre vector of the to-be-matched singer corresponding to the to-be-matched singer audio and at least two candidate timbre vectors of the candidate singers corresponding to the at least two candidate singer audios, wherein the pre-trained timbre prediction model is trained by contrastive learning and a ternary loss function; Respectively determining the similarity between the timbre vector and each of the candidate timbre vectors; Determining a target timbre vector matching the timbre vector from the at least two candidate timbre vectors according to the similarity; The sound effect parameters of the audio corresponding to the target timbre vector are determined as the sound effect parameters of the singer audio to be matched.

2. The method according to claim 1, characterized in that Before extracting timbre features from the to-be-matched singer audio and the at least two candidate singer audios using the pre-trained timbre prediction model, the method further comprises: The initialized timbre prediction model is trained, and the training process includes: Obtain anchor samples, positive samples and negative samples, wherein the anchor samples include frequency spectrum feature vectors of the target song audio sung by the singer to be matched, the positive samples include frequency spectrum feature vectors of the other song audios sung by the singer to be matched that are different from the target song, and the negative samples include frequency spectrum feature vectors of any song sung by other singers that are different from the singer to be matched; Inputting the anchor point sample into the initialized timbre prediction model to obtain a timbre vector of the anchor point sample output by the initialized timbre prediction model; Inputting the positive sample into the initialized timbre prediction model to obtain a timbre vector of the positive sample output by the initialized timbre prediction model; Inputting the negative sample into the initialized timbre prediction model to obtain a timbre vector of the negative sample output by the initialized timbre prediction model; Calculating a loss amount according to the timbre vector of the anchor sample, the timbre vector of the positive sample, the timbre vector of the negative sample and a ternary loss function; The initialized timbre prediction model is trained according to the loss amount and the back propagation algorithm to obtain the pre-trained timbre prediction model.

3. The method according to claim 1, characterized in that Before extracting timbre features from the to-be-matched singer audio and the at least two candidate singer audios using the pre-trained timbre prediction model, the method further includes: Obtain multiple frequency spectrum feature vectors of multiple audios of the singer to be matched; Obtaining at least two candidate frequency spectrum feature vectors corresponding to the at least two candidate singer audios; The method of extracting timbre features from the audio of the singer to be matched and the audio of at least two candidate singers by using the pre-trained timbre prediction model to obtain the timbre vector of the singer to be matched corresponding to the audio of the singer to be matched and at least two candidate timbre vectors of the candidate singers corresponding to the audio of the at least two candidate singers includes: Inputting the plurality of spectrum feature vectors into the pre-trained timbre prediction model respectively to obtain a plurality of timbre vectors output by the pre-trained timbre prediction model; Inputting the at least two candidate spectrum feature vectors into the pre-trained timbre prediction model to obtain at least two candidate timbre vectors output by the pre-trained timbre prediction model; Before respectively determining the similarity between the timbre vector and each of the candidate timbre vectors, the method further includes: Screening out a first number of outlier timbre vectors from the plurality of timbre vectors to obtain a second number of timbre vectors; The timbre vector is calculated based on the second number of timbre vectors.

4. The method according to any one of claims 1 to 3, characterized in that Before extracting timbre features from the to-be-matched singer audio and the at least two candidate singer audios using the pre-trained timbre prediction model, the method further includes: Preprocessing the to-be-matched singer audio and the at least two candidate singer audios to obtain preprocessed to-be-matched singer audio and at least two candidate singer audios, wherein the preprocessing includes extracting vocal audio from the audio, removing silent sections from the vocal audio, and splicing the vocal audio after removing the silent sections; The method of extracting timbre features from the audio of the singer to be matched and the audio of at least two candidate singers by using the pre-trained timbre prediction model includes: The timbre features of the preprocessed singer audio to be matched and the preprocessed at least two candidate singer audios are extracted through a pretrained timbre prediction model.

5. A sound effect matching method, characterized in that: include: The timbre and sound effect conversion is performed on the timbre vector of the to-be-matched singer audio and the timbre vectors of at least two candidate singer audios by a pre-trained timbre and sound effect conversion model to obtain the timbre and sound effect vector of the to-be-matched singer audio and at least two candidate timbre and sound effect vectors of the at least two candidate singer audios, wherein the pre-trained timbre and sound effect conversion model is trained by contrastive learning and a ternary loss function; Respectively determining the similarity between the timbre sound effect vector and each of the candidate timbre sound effect vectors; According to the similarity, determining a target timbre sound effect vector matching the timbre sound effect vector from the at least two candidate timbre sound effect vectors; The sound effect parameters of the audio corresponding to the target timbre sound effect vector are determined as the sound effect parameters of the singer audio to be matched.

6. The method according to claim 5, characterized in that Before performing timbre and sound effect conversion on the timbre vector of the matching singer audio and the timbre vectors of at least two candidate singer audios through the pre-trained timbre and sound effect conversion model, the method further includes: The initialized timbre and sound effect conversion model is trained, and the training process includes: Acquire training samples of the initialized timbre and sound effect conversion model, wherein the training samples include a first anchor point sample, a first positive sample, and a first negative sample; Inputting the first anchor point sample into the initialized timbre and sound effect conversion model to obtain a timbre and sound effect vector of the first anchor point sample output by the initialized timbre and sound effect conversion model; Inputting the first positive sample into the initialized timbre-sound effect conversion model to obtain a timbre-sound effect vector of the first positive sample output by the initialized timbre-sound effect conversion model; Inputting the first negative sample into the initialized timbre and sound effect conversion model to obtain a timbre and sound effect vector of the first negative sample output by the initialized timbre and sound effect conversion model; Calculating a loss value according to the timbre and sound effect vector of the first anchor point sample, the timbre and sound effect vector of the first positive sample, the timbre and sound effect vector of the first negative sample, and a ternary loss function; The initialized timbre and sound effect conversion model is trained according to the loss value and the back propagation algorithm to obtain the pre-trained timbre and sound effect conversion model.

7. The method according to claim 6, characterized in that The step of obtaining the training sample of the timbre and sound effect conversion model comprises: Determine multiple similarities between a sound effect vector of an anchor singer's audio and multiple candidate sound effect vectors of multiple candidate singers' audio, wherein the anchor singer is any singer; Determine a plurality of first candidate sound effect vectors similar to the sound effect vector of the anchor singer audio according to the plurality of similarities and a preset first threshold, wherein the preset first threshold is used to characterize the similarity between the sound effect vector and the candidate sound effect vector; Determine a plurality of second candidate sound effect vectors that are dissimilar to the sound effect vector of the anchor singer audio according to the plurality of similarities and a preset second threshold, wherein the preset second threshold is used to characterize the dissimilarity between the sound effect vector and the candidate sound effect vectors; Determine the timbre vector of the anchor singer audio as the first anchor sample; Determine the timbre vectors of the first candidate singer audio corresponding to the plurality of first candidate sound effect vectors as the first positive samples; The timbre vectors of the second candidate singer audio corresponding to each of the plurality of second candidate sound effect vectors are determined as the first negative samples.

8. The method according to claim 7, characterized in that Before determining multiple similarities between the sound effect vector of the anchor singer audio and multiple candidate sound effect vectors of multiple candidate singer audios, the method further includes: Acquire sound effect parameters set for the anchor singer audio and each candidate singer audio, wherein the sound effect parameters include at least one of an equalizer adjustment super parameter, a reverberation super parameter, and a dynamic compression super parameter; The sound effect parameters set for the anchor singer audio and for each alternative singer audio are respectively converted into sound effect vectors to obtain the sound effect vector of the anchor singer audio and multiple sound effect vectors of each alternative singer audio.

9. The method according to claim 5, characterized in that Before performing timbre and sound effect feature conversion on the timbre vector of the matching singer audio and the timbre vectors of at least two candidate singer audios through the pre-trained timbre and sound effect conversion model, the method further includes: Obtaining a frequency spectrum feature vector of the to-be-matched singer's audio and at least two candidate frequency spectrum feature vectors of at least two candidate singers' audio; The spectral feature vector of the singer's audio to be matched and the at least two alternative spectral feature vectors of the at least two alternative singers' audios are respectively input into the pre-trained timbre prediction model to obtain the timbre vector of the singer's audio to be matched and the at least two alternative timbre vectors of the at least two alternative singers' audios output by the pre-trained timbre prediction model.

10. The method according to claim 9, characterized in that Before respectively inputting the frequency spectrum feature vector of the to-be-matched singer audio and the multiple candidate frequency spectrum feature vectors of the multiple candidate singer audios into the pre-trained timbre prediction model, the method further includes: The initialized timbre prediction model is trained, and the training process includes: Obtain a second anchor point sample, a second positive sample, and a second negative sample, wherein the second anchor point sample includes a frequency spectrum feature vector of the audio of the target song sung by the singer to be matched, the second positive sample includes a frequency spectrum feature vector of the audio of other songs sung by the singer to be matched that are different from the target song, and the second negative sample includes a frequency spectrum feature vector of any song sung by other singers that are different from the singer to be matched; Inputting the second anchor point sample into the initialized timbre prediction model to obtain a timbre vector of the second anchor point sample output by the initialized timbre prediction model; Inputting the second positive sample into the initialized timbre prediction model to obtain a timbre vector of the second positive sample output by the initialized timbre prediction model; Inputting the second negative sample into the initialized timbre prediction model to obtain a timbre vector of the second negative sample output by the initialized timbre prediction model; Calculate a loss value according to the timbre vector of the second anchor point sample, the timbre vector of the second positive sample, the timbre vector of the second negative sample and a ternary loss function; The initialized timbre prediction model is trained according to the loss value and the back propagation algorithm to obtain the pre-trained timbre prediction model.

11. The method according to claim 9, characterized in that The step of obtaining the frequency spectrum feature vector of the singer audio to be matched includes: Obtain multiple frequency spectrum feature vectors of multiple audios of the singer to be matched; Inputting the frequency spectrum feature vector of the to-be-matched singer audio into a pre-trained timbre prediction model to obtain the timbre vector of the to-be-matched singer output by the timbre prediction model, including: Inputting multiple frequency spectrum feature vectors of multiple audios of the singer to be matched into the pre-trained timbre prediction model respectively to obtain multiple timbre vectors of the singer to be matched output by the pre-trained timbre prediction model; Before performing timbre and sound effect feature conversion on the timbre vector of the to-be-matched singer audio through the pre-trained timbre and sound effect conversion model, the method further comprises: Screening out a first number of outlier timbre vectors from a plurality of timbre vectors of the singer to be matched, so as to obtain a second number of timbre vectors; According to the second number of timbre vectors, the timbre vector of the singer audio to be matched is calculated.

12. The method according to any one of claims 9 to 11, characterized in that Before obtaining the frequency spectrum feature vector of the to-be-matched singer audio and at least two candidate frequency spectrum feature vectors of at least two candidate singer audios, the method further includes: Obtain the audio of the singer to be matched and at least two alternative singers; Preprocessing the to-be-matched singer audio and the at least two candidate singer audios to obtain preprocessed to-be-matched singer audio and at least two candidate singer audios, wherein the preprocessing includes extracting vocal audio from the audio, removing silent sections from the vocal audio, and splicing the vocal audio after removing the silent sections; The step of obtaining the frequency spectrum feature vector of the to-be-matched singer audio and at least two candidate frequency spectrum feature vectors of at least two candidate singer audios includes: Obtain the preprocessed frequency spectrum feature vector of the to-be-matched singer audio and at least two candidate frequency spectrum feature vectors of at least two candidate singer audios after preprocessing.

13. A computer device comprising a processor, characterized in that: When executing the computer program stored in the memory, the processor is used to implement the sound effect matching method as described in any one of claims 1 to 4, or the sound effect matching method as described in any one of claims 5 to 12.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it is used to implement the sound effect matching method according to any one of claims 1 to 4, or the sound effect matching method according to any one of claims 5 to 12.

15. A computer program product having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it is used to implement the sound effect matching method according to any one of claims 1 to 4, or the sound effect matching method according to any one of claims 5 to 12.