Audio style migration method, computer equipment, storage medium and program product

By iteratively processing the audio and updating the sound effect configuration parameters in real time, the problem of low efficiency in traditional audio style transfer is solved, achieving efficient and real-time audio style transfer and reducing dependence on data volume and computing resources.

CN120954362APending Publication Date: 2025-11-14TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511085627.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Traditional audio style transfer methods rely on end-to-end neural network models, which require a large amount of training data, resulting in low transfer efficiency.

Method used

By acquiring multiple sets of current sound effect configuration parameters, the initial audio is iteratively processed, candidate audios that meet the similarity condition with the target audio are selected, and the sound effect configuration parameter set is updated based on the candidate audios until the iteration stops, thus generating the target sound effect configuration parameter set to achieve audio style transfer.

Benefits of technology

It achieves real-time and efficient transfer of audio styles, reduces dependence on data volume and computing resources, significantly improves transfer efficiency, and avoids the offline training step in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954362A_ABST
    Figure CN120954362A_ABST
Patent Text Reader

Abstract

The invention relates to an audio style migration method, computer equipment, a storage medium and a program product. The method comprises the following steps: acquiring a plurality of current sound effect configuration parameter sets; respectively applying the plurality of current sound effect configuration parameter sets to the initial audio to perform iterative processing on the initial audio to obtain a plurality of processed audios; the processed audios meeting the similarity condition with the target audio are selected from the multiple processed audios to serve as candidate audios; generating a plurality of updated current sound effect configuration parameter sets according to the current sound effect configuration parameter set corresponding to the candidate audio under the condition that the iteration stop condition is not met, and continuing to perform iteration processing on the initial audio by utilizing the plurality of updated current sound effect configuration parameter sets; and determining a target sound effect configuration parameter set of the initial audio according to the current sound effect configuration parameter set corresponding to the candidate audio under the condition that an iteration stop condition is satisfied. By adopting the method, the audio style migration efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to an audio style transfer method, computer device, computer-readable storage medium, and computer program product. Background Technology

[0002] Audio style transfer technology, as an important branch of digital audio processing, has seen a surge in demand in recent years for applications such as music production, film and television post-production, and game sound effects. The core objective of audio style transfer technology is to transfer a target style (such as classic records or specific scene sound effects) to input audio (such as vocals or instruments), thereby enabling personalized creation or industrialized content production. With the explosive growth of multimedia content, users' demand for fast, high-quality audio style conversion is becoming increasingly urgent. However, traditional audio style transfer methods rely on end-to-end neural network models, requiring large amounts of training data for offline training, resulting in low efficiency. Summary of the Invention

[0003] Therefore, it is necessary to provide an audio style transfer method, computer device, computer-readable storage medium, and computer program product that can improve the efficiency of audio style transfer in response to the above-mentioned technical problems.

[0004] Firstly, this application provides an audio style transfer method, including:

[0005] Obtain multiple sets of current sound effect configuration parameters, wherein the current sound effect configuration parameter set includes multiple sound effect configuration parameters used for sound effect configuration of audio;

[0006] The multiple sets of current sound effect configuration parameters are applied to the initial audio to iteratively process the initial audio and obtain multiple processed audios; the processed audios that meet the similarity condition with the target audio are selected from the multiple processed audios as candidate audios.

[0007] If the iteration stopping condition is not met, multiple updated sets of current sound effect configuration parameters are generated based on the current sound effect configuration parameter set corresponding to the candidate audio, and the initial audio is further iteratively processed using the multiple updated sets of current sound effect configuration parameters.

[0008] If the iteration stopping condition is met, the target sound effect configuration parameter set of the initial audio is determined according to the current sound effect configuration parameter set corresponding to the candidate audio. The target sound effect configuration parameter set is used to transfer the audio style of the target audio to the initial audio.

[0009] In one embodiment, obtaining a set of multiple current sound effect configuration parameters includes:

[0010] Get the initial sound effect configuration parameter set;

[0011] Based on the various random adjustment results of the initial sound effect configuration parameter set, the multiple current sound effect configuration parameter sets are generated.

[0012] In one embodiment, the plurality of sound effect configuration parameters include effect parameters corresponding to each of the plurality of sound effect processors; applying the plurality of current sound effect configuration parameter sets to the initial audio respectively includes:

[0013] Adjust the parameters of each effect unit in each current sound effect configuration parameter set to obtain a set of sound effect configuration parameters corresponding to each current sound effect configuration parameter set;

[0014] The sound effect configuration parameters of each group are applied to the initial audio respectively.

[0015] In one embodiment, adjusting each of the effect parameters in each of the current sound effect configuration parameter sets includes:

[0016] Obtain the adjustment range of each effect parameter in each of the current sound effect configuration parameter sets;

[0017] The effect parameters are adjusted according to the adjustment range of the effect parameters and the corresponding parameter adjustment ratio.

[0018] In one embodiment, generating updated sets of multiple current sound effect configuration parameters based on the current sound effect configuration parameter set corresponding to the candidate audio includes:

[0019] Generate a random array; the random array includes multiple random numbers, each of which corresponds to a sound effect configuration parameter in the current sound effect configuration parameter set.

[0020] Based on multiple random arrays, the set of current sound effect configuration parameters corresponding to the candidate audio is updated to obtain the updated set of multiple current sound effect configuration parameters.

[0021] In one embodiment, generating the random array includes:

[0022] Random vectors are obtained by sampling from a Gaussian distribution;

[0023] The noise vector is determined based on the random vector and the changing step size;

[0024] The random array is generated based on the noise vector.

[0025] In one embodiment, the sound effect configuration parameters are normalized parameters; updating the current sound effect configuration parameter set corresponding to the candidate audio to obtain the updated multiple current sound effect configuration parameter sets includes:

[0026] The current sound effect configuration parameter set corresponding to the candidate audio is updated to obtain multiple new sound effect configuration parameter sets;

[0027] The sound effect configuration parameters that exceed the normalized parameter range in the multiple new sound effect configuration parameter sets are clipped to obtain multiple clipped sound effect configuration parameter sets.

[0028] The multiple sets of trimmed sound effect configuration parameters are used as the updated sets of current sound effect configuration parameters.

[0029] In one embodiment, selecting processed audio that meets the similarity condition with the target audio from the plurality of processed audios as candidate audio includes:

[0030] The processed audio and the target audio are input into a pre-trained sound effect feature extraction model to obtain the sound effect features of the processed audio and the sound effect features of the target audio.

[0031] Calculate the similarity between the sound effect features of the processed audio and the sound effect features of the target audio;

[0032] Based on the calculated similarity, processed audio that meets the similarity condition is selected from the plurality of processed audios as candidate audio.

[0033] In one embodiment, the method further includes:

[0034] The target sound effect configuration parameter set is applied to the initial audio to generate audio with the audio style of the target audio.

[0035] Secondly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.

[0036] Thirdly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.

[0037] Fourthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.

[0038] The aforementioned audio style transfer method, apparatus, computer device, computer-readable storage medium, and computer program product, by acquiring multiple sets of current sound effect configuration parameters, can simultaneously apply multiple sound effect configurations to the initial audio and generate multiple processed audios. It can test multiple sound effect configurations at once, quickly filtering out candidate audios that are closer to the target audio style. Furthermore, based on the current sound effect configuration parameter sets corresponding to the candidate audios, iteratively generating multiple new sets of current sound effect configuration parameters ensures that each iteration is adjusted based on the current optimal sound effect configuration. This allows the sound effect configuration parameters to consistently converge towards a direction closer to the target audio style, shortening the path to finding the target sound effect configuration parameter set. This achieves real-time and efficient transfer of the target audio style, and the entire process does not require large-scale model training beforehand. Instead, it directly approximates the target style by iteratively adjusting parameters in real time, eliminating the offline training required by traditional end-to-end models. This significantly reduces the dependence on data volume and computing resources, thereby significantly improving the efficiency of audio style transfer. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0040] Figure 1 This is an application environment diagram of an audio style transfer method in one embodiment of this application;

[0041] Figure 2 This is a flowchart illustrating an audio style transfer method in one embodiment of this application;

[0042] Figure 3 This is a flowchart of an audio style transfer method in one embodiment of this application;

[0043] Figure 4 This is a flowchart illustrating an audio style transfer method in another embodiment of this application;

[0044] Figure 5 This is a structural block diagram of an audio style transfer device according to one embodiment of this application;

[0045] Figure 6 This is an internal structural diagram of a computer device in one embodiment of this application. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0047] The audio style transfer method provided in this application can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located in the cloud or on another network server. Terminal 102 acquires multiple sets of current sound effect configuration parameters, which include multiple sound effect configuration parameters used for sound effect configuration of audio. Terminal 102 applies the multiple sets of current sound effect configuration parameters to the initial audio to iteratively process the initial audio, resulting in multiple processed audios. Terminal 102 selects processed audios that meet the similarity condition with the target audio from the multiple processed audios as candidate audios. If the iteration stopping condition is not met, terminal 102 generates updated sets of current sound effect configuration parameters based on the current sound effect configuration parameter sets corresponding to the candidate audios, and continues to iteratively process the initial audio using the updated sets of current sound effect configuration parameters. If the iteration stopping condition is met, terminal 102 determines the target sound effect configuration parameter set for the initial audio based on the current sound effect configuration parameter set corresponding to the candidate audios. The target sound effect configuration parameter set is used to transfer the audio style of the target audio to the initial audio. The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle systems, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. The server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0048] In one exemplary embodiment, such as Figure 2 As shown, this application provides an audio style transfer method, which is applied to... Figure 1 Taking terminal 102 as an example, the explanation includes:

[0049] Step S202: Obtain a set of multiple current sound effect configuration parameters.

[0050] The current sound effect configuration parameter set includes multiple sound effect configuration parameters used to configure the audio. During audio style transfer, the current sound effect configuration parameter set can refer to the set of sound effect configuration parameters used to configure the initial audio in the current iteration.

[0051] For example, the multiple sound effect configuration parameters in the current sound effect configuration parameter set may include the effect parameters corresponding to each of the multiple sound effect units. In practical applications, audio style can refer to a style label subjectively perceived by the human ear after audio is processed through a sound effect chain composed of sound effect units, such as "retro" or "ethereal". In the sound effect chain, each sound effect unit can process the audio sequentially, and each sound effect unit corresponds to adjustable effect parameters.

[0052] For example, sound effects may include equalizers, compressors, reverbs, and delays. The equalizer's parameters may include, but are not limited to: low frequency gain, low frequency center, mid frequency gain, mid frequency center, high frequency gain, and high frequency center. The compressor's parameters may include, but are not limited to: threshold, compression ratio, attack time, and release time. The reverb's parameters may include, but are not limited to: reverb time, early reflections intensity, and diffusivity. The delay's parameters may include, but are not limited to: delay time and feedback ratio.

[0053] Step S204: Apply multiple sets of current sound effect configuration parameters to the initial audio to iteratively process the initial audio and obtain multiple processed audios; select the processed audio that meets the similarity condition with the target audio from the multiple processed audios as candidate audios.

[0054] The initial audio is the audio of the audio style to be migrated.

[0055] In practice, applying multiple sets of current sound effect configuration parameters to the initial audio can refer to processing the initial audio through the audio processing functions of the sound effect chain. Assume the initial audio is... The i-th current sound effect configuration parameter set is Set the current sound effect configuration parameters Apply to initial audio The processed audio is above. , can be represented as:

[0056] ;

[0057] Where A represents the audio processing function of the sound effect link, which can perform various different sound effect configuration processing flows on the initial audio based on multiple current sound effect configuration parameter sets to obtain multiple processed audio.

[0058] In the specific implementation, multiple sets of current sound effect configuration parameters are applied to the initial audio to perform iterative processing on the initial audio, resulting in multiple processed audios. The iterative processing may include the step of applying multiple sets of current sound effect configuration parameters to the initial audio in each iteration. Thus, in the process of audio style transfer, the similarity of audio style between the processed audio and the target audio is gradually optimized by repeatedly adjusting the sets of sound effect configuration parameters and applying them to the initial audio.

[0059] For example, the multiple sound effect configuration parameters in the current sound effect configuration parameter set may include the effect parameters corresponding to each of the multiple sound effect processors. In practical applications, applying the current sound effect configuration parameter set to the initial audio can be done by processing the initial audio based on the effect parameters in the current sound effect configuration parameter set. For instance, if the current sound effect configuration parameter set includes the effect parameters corresponding to the equalizer, compressor, and reverb, then frequency band gain adjustment, dynamic range compression, and spatial reverb overlay processing can be performed on the initial audio based on these effect parameters to obtain the processed audio. Therefore, applying multiple current sound effect configuration parameter sets to the initial audio can be done by processing the same initial audio based on multiple different sound effect configuration processes, thereby obtaining multiple processed audios corresponding to different audio styles. For example, applying five current sound effect configuration parameter sets to the initial audio can be done by processing the same initial audio based on five different sound effect configuration processes to obtain five processed audios with different audio styles.

[0060] In practical applications, in the initial iteration round, the multiple sets of current sound effect configuration parameters obtained can be multiple sets of sound effect configuration parameters derived from a basic set of sound effect configuration parameters; or, the multiple sets of current sound effect configuration parameters can be multiple sets of preset sound effect configuration parameters corresponding to different audio styles.

[0061] In practice, candidate audio is selected from multiple processed audio files that meet a similarity condition with the target audio. Meeting the similarity condition may include a degree of audio style similarity between the processed audio and the target audio exceeding a certain threshold.

[0062] In practice, sound effect features of multiple processed audio files and the target audio file can be extracted. Based on the similarity between the sound effect features of the processed audio files and the target audio files, the processed audio files that meet the similarity criteria with the target audio files are selected from the multiple processed audio files. The sound effect features are those extracted from the audio files that can be used to quantify the sound effect configuration results.

[0063] The target audio is audio with the target audio style. The sound effect features of the target audio can be used as reference sound effect features to select the processed audio with a high similarity to the reference sound effect features from multiple processed audios as candidate audio.

[0064] Here, similarity can refer to the degree of similarity between two sound effect features. Optionally, similarity can be determined by any calculation method such as Euclidean distance, cosine similarity, correlation coefficient, and Jaccard similarity coefficient. In specific implementation, similarity conditions can include any of the following conditions: maximum similarity, top K similarity rankings (K>=1), or similarity greater than a certain threshold.

[0065] Step S206: If the iteration stop condition is not met, generate multiple updated current sound effect configuration parameter sets based on the current sound effect configuration parameter sets corresponding to the candidate audio, and continue iterative processing of the initial audio using the updated multiple current sound effect configuration parameter sets.

[0066] The iteration stopping condition may include at least one of the following conditions: reaching the maximum number of iterations, or having at least one similarity between the sound effect features of the processed audio and the sound effect features of the target audio that is greater than a similarity threshold.

[0067] The current sound effect configuration parameter set corresponding to the candidate audio is the set of current sound effect configuration parameters applied to the initial audio to obtain the candidate audio. For example, if the current sound effect configuration parameter set 'a' is applied to the initial audio to obtain the processed audio A, then the current sound effect configuration parameter set corresponding to the processed audio A is the current sound effect configuration parameter set 'a'. Therefore, if the processed audio A is determined to be a candidate audio, then the current sound effect configuration parameter set corresponding to the candidate audio is the current sound effect configuration parameter set 'a'.

[0068] In practice, based on the current sound effect configuration parameter set corresponding to the candidate audio, multiple updated current sound effect configuration parameter sets are generated. Alternatively, multiple new sound effect configuration parameter sets can be derived from the current sound effect configuration parameter set corresponding to the candidate audio, serving as the updated sets. For example, the current sound effect configuration parameter set corresponding to the candidate audio can be adjusted according to various parameter adjustment strategies to obtain multiple new derived sound effect configuration parameter sets, which are then used as the updated sets. In these various parameter adjustment strategies, the types of effect parameters adjusted differ, or the adjustment range for the same effect parameter may also vary.

[0069] In one embodiment, multiple random arrays can be obtained. Then, based on these random arrays, the current sound effect configuration parameter set corresponding to the candidate audio is adjusted to obtain updated sets of current sound effect configuration parameters. Each random array includes multiple random numbers, and the number of random numbers in each array is the same as the number of sound effect configuration parameters in the current sound effect configuration parameter set. For example, suppose there are three different random arrays a, b, and c, each containing 10 random numbers. The current sound effect configuration parameter set d corresponding to the candidate audio contains 10 sound effect configuration parameters. The 10 random numbers from random array a can be superimposed on the 10 sound effect configuration parameters in the current sound effect configuration parameter set d to obtain the updated sets of current sound effect configuration parameters. Similarly, the 10 random numbers in the random array b are superimposed on the 10 sound effect configuration parameters in the current sound effect configuration parameter set d to obtain the updated current sound effect configuration parameter set. The 10 random numbers from the random array c are then superimposed onto the 10 sound effect configuration parameters in the current sound effect configuration parameter set d, resulting in the updated current sound effect configuration parameter set. Therefore, the current sound effect configuration parameter set d can be used to generate... , and This serves as a set of updated current sound effect configuration parameters.

[0070] In one embodiment, the Covariance Matrix Adaptation Evolution Strategy (CMA-ES) algorithm can be used to process the current sound effect configuration parameter set corresponding to the candidate audio, resulting in multiple updated current sound effect configuration parameter sets. Specifically, the CMA-ES algorithm treats the current sound effect configuration parameter set corresponding to the candidate audio as an initial population. Based on this initial population, a small random perturbation is applied to each parameter vector in the current sound effect configuration parameter set corresponding to the candidate audio to generate a new population. Applying multiple different random perturbations to the same candidate audio's current sound effect configuration parameter set yields multiple new populations, which are the updated current sound effect configuration parameter sets. The random perturbation can be normally distributed, with a mean of 0 and a standard deviation determined based on the current covariance matrix and the variable asynchronous length. The variable asynchronous length of the covariance matrix adaptive evolution strategy algorithm is used to determine the magnitude of random perturbation in each iteration. In each round of iteration, the covariance matrix and the variable asynchronous length can be updated according to the performance of the new population. For example, the variable asynchronous length can be increased or decreased according to the performance of the new population to improve search efficiency.

[0071] After generating multiple updated sets of current sound effect configuration parameters, the initial audio is iteratively processed using these updated sets. This means that we can return to the step of applying each set of current sound effect configuration parameters to the initial audio to iteratively process it and obtain multiple processed audio files, thus entering the next iteration. It is evident that the sets of current sound effect configuration parameters applied to the initial audio are different in each iteration.

[0072] Step S208: If the iteration stopping condition is met, determine the target sound effect configuration parameter set of the initial audio based on the current sound effect configuration parameter set corresponding to the candidate audio.

[0073] As an example, assuming there is only one candidate audio, the current sound effect configuration parameter set corresponding to the candidate audio can be used as the target sound effect configuration parameter set for the initial audio.

[0074] As another example, assuming there are multiple candidate audios, the target audio configuration parameter set for the initial audio can be determined based on the current audio configuration parameter set corresponding to each of the multiple candidate audios. For example, for the current audio configuration parameter sets corresponding to each of the multiple candidate audios, the average value of the same audio configuration parameter in each current audio configuration parameter set can be calculated, and the target audio configuration parameter set can be generated based on the average value of various audio configuration parameters.

[0075] The target sound effect configuration parameter set is used to transfer the audio style of the target audio to the initial audio.

[0076] In practice, the target audio effect configuration parameter set is applied to the initial audio, which can transfer the audio style of the target audio to the initial audio.

[0077] In the above audio style transfer method, in each iteration, multiple sets of current sound effect configuration parameters updated in real time are directly applied to the initial audio to generate multiple processed audios. Then, candidate audios that meet the similarity condition are selected from the multiple processed audios. If the iteration stopping condition is not met, the set of current sound effect configuration parameters corresponding to the candidate audios is updated to enter the next iteration process. If the iteration stopping condition is met, the target sound effect configuration parameter set is generated based on the set of current sound effect configuration parameters corresponding to the candidate audios selected in the latest iteration, thus realizing the real-time and efficient transfer of the audio style of the target audio.

[0078] In another embodiment, obtaining multiple sets of current sound effect configuration parameters includes: obtaining an initial set of sound effect configuration parameters; and generating multiple sets of current sound effect configuration parameters based on various random adjustment results of the initial set of sound effect configuration parameters.

[0079] The initial sound effect configuration parameter set includes multiple sound effect configuration parameters used to configure the sound effects of the audio.

[0080] In one embodiment, the initial sound effect configuration parameter set can be adjusted according to multiple random arrays to obtain various random adjustment results for the initial sound effect configuration parameter set, which serve as multiple current sound effect configuration parameter sets. Each random adjustment result corresponds to an adjustment result of the initial sound effect configuration parameter set by a random array.

[0081] In practical implementation, the initial sound effect configuration parameter set can include multiple sound effect parameters corresponding to each of multiple sound effect units. The initial sound effect configuration parameter set can be a set of parameter vectors corresponding to multiple sound effect units. In other words, in the initial sound effect configuration parameter set, each sound effect parameter can be represented as a parameter vector corresponding to any sound effect unit, and the parameters in this parameter vector are the effect parameters of that sound effect unit. For example, the initial sound effect configuration parameter set can be represented as... ,in, Configure parameters for the i-th sound effect, i.e., the parameter vector corresponding to sound effect i; parameter vector Parameters in This can represent the parameters of the j-th effect contained in the parameter vector of the i-th sound effect. For example, suppose the parameter vector corresponding to the equalizer is... The parameter vector Parameters in This can include the effects parameters corresponding to the equalizer (such as low-frequency gain, low-frequency center frequency, mid-frequency gain, etc.).

[0082] As an example, a random array could be in the range of Random vectors obtained by sampling from a random distribution. Multiple random vectors obtained by sampling from a random distribution. For the initial sound effect configuration parameter set By making adjustments, multiple sets of current sound effect configuration parameters can be obtained. The specific process can be represented as follows:

[0083] .

[0084] Among them, the i-th random vector can be used Initial sound effect configuration parameter set Adjustments are made to obtain the i-th current sound effect configuration parameter set. .

[0085] As another example, a random array can be a random vector sampled from a Gaussian distribution. Multiple random vectors sampled from a Gaussian distribution... By adjusting the initial set of sound effect configuration parameters, multiple sets of current sound effect configuration parameters can be obtained. The specific process can be represented as follows:

[0086] .

[0087] In one embodiment, the Covariance Matrix Adaptation Evolution Strategy (CMA-ES) algorithm can be used to process the initial sound effect configuration parameter set, resulting in multiple random adjustments to the initial sound effect configuration parameter set, which serve as multiple current sound effect configuration parameter sets. Specifically, the Covariance Matrix Adaptation Evolution Strategy algorithm treats the initial sound effect configuration parameter set as an initial population. Based on this initial population, a small random perturbation is applied to each parameter vector in the initial sound effect configuration parameter set to generate a new population. Applying multiple different random perturbations to the same initial sound effect configuration parameter set yields multiple new populations, which constitute multiple current sound effect configuration parameter sets.

[0088] The technical solution of this embodiment generates multiple current sound effect configuration parameter sets based on various random adjustment results of the initial sound effect configuration parameter set, providing more possibilities for sound effect configuration and improving the flexibility, comprehensiveness and efficiency of parameter generation during audio style transfer.

[0089] In another implementation, multiple sound effect configuration parameters include the effect parameters corresponding to each of multiple sound effect processors; applying multiple sets of current sound effect configuration parameters to the initial audio includes: adjusting each effect parameter in each set of current sound effect configuration parameters to obtain a set of sound effect configuration parameters corresponding to each set of current sound effect configuration parameters; and applying each set of sound effect configuration parameters to the initial audio.

[0090] In practice, the parameters of each effect in each current sound effect configuration parameter set can be adjusted according to the parameter adjustment information to obtain a set of sound effect configuration parameters corresponding to each current sound effect configuration parameter set. The parameter adjustment information may include predefined parameter mapping rules. For example, parameter mapping rules can be used to define how to map effect parameters to a numerical range suitable for audio style transfer processing; for instance, parameter mapping rules can define how to map normalized effect parameters to real physical values.

[0091] In practice, the effect parameters of each sound effect processor included in each current sound effect configuration parameter set can be mapped to a reasonable range according to predefined parameter mapping rules, thereby enabling the adjustment of each effect parameter included in each current sound effect configuration parameter set.

[0092] Specifically, for each current sound effect configuration parameter set, the corresponding set of sound effect configuration parameters includes the effect parameters corresponding to each sound effect effect after adjustment. Applying each set of sound effect configuration parameters to the initial audio can process the same initial audio based on multiple different sound effect configuration processes, thereby obtaining multiple processed audios corresponding to different audio styles.

[0093] The technical solution of the above embodiments, by constructing a set of sound effect configuration parameters through interpretable effect parameters, can solve the problem of poor interpretability caused by the inability of black box neural network models to guide parameter adjustment strategies in traditional technologies, and enhance the flexibility, accuracy and universality of audio style transfer; furthermore, by adjusting the effect parameters, the values ​​of sound effect configuration parameters can be optimized flexibly and efficiently, reducing the risk of numerical errors and improving the efficiency and reliability of sound effect audio style transfer.

[0094] In another embodiment, adjusting each effect parameter in each current sound effect configuration parameter set includes: obtaining the adjustment range of each effect parameter in each current sound effect configuration parameter set; and adjusting the effect parameter according to the adjustment range of the effect parameter and the parameter adjustment ratio corresponding to the effect parameter.

[0095] In practical applications, each sound effect processor's parameters have their own valid physical range. A valid physical range refers to a range determined based on the working principles of physical hardware or the objective laws governing physical acoustic phenomena. To avoid invalid parameters during audio style transfer, effect processor parameters can be mapped to their corresponding valid physical ranges, which can then serve as the adjustment range for the effect processor parameters.

[0096] As an example, the parameter adjustment ratio can be a coefficient that controls the range of change of the effector parameter. The adjusted effector parameter can be obtained by adding the product of the adjustment range and the parameter adjustment ratio to the original effector parameter.

[0097] As another example, suppose the effects parameters are normalized parameters, and the normalized parameter range for each effects parameter is [0, 1]. The adjustment range is Then, based on the effects pedal parameters Adjustment range and effects parameters The normalized parameter range [0, 1] determines the adjustment ratio of the effect pedal parameters, thus allowing for adjustments to the effect pedal parameters based on the adjustment range and the parameter adjustment ratio. The formula for making the adjustment can be expressed as:

[0098] ;

[0099] in, For the parameters of this effect The parameter values ​​mapped to the valid physical range are the adjusted effect parameters. Based on the adjusted effect parameters in each current sound effect configuration parameter set, a set of sound effect configuration parameters corresponding to each current sound effect configuration parameter set can be generated.

[0100] In the specific implementation, the sound effect configuration parameters of each group are applied to the initial audio. Let's assume the initial audio is... Each set of sound effect configuration parameters is as follows: Set a set of sound effect configuration parameters Apply to initial audio The processed audio is above. , can be represented as:

[0101] ;

[0102] Where A represents the audio processing function of the sound effect link, which can simulate the processing flow of the sound effect processor according to the configuration parameters of each group of sound effects, process the initial audio, and obtain multiple processed audio.

[0103] The technical solution of the above embodiments can adjust the effects parameters according to the parameter adjustment range and parameter adjustment ratio, thereby improving the flexibility, reliability and efficiency of audio style transfer.

[0104] In another embodiment, generating multiple updated sets of current sound effect configuration parameters based on the current sound effect configuration parameter set corresponding to the candidate audio includes: generating a random array; updating the current sound effect configuration parameter set corresponding to the candidate audio based on the multiple random arrays to obtain multiple updated sets of current sound effect configuration parameters.

[0105] The random array contains multiple random numbers, each corresponding to a sound effect configuration parameter in the current sound effect configuration parameter set. For example, if the random array contains [p, q, r] and the sound effect configuration parameters are [e1, e2, e3], then p is used to adjust e1, q is used to adjust e2, and r is used to adjust e3. Adjusting the sound effect configuration parameters using random numbers can include multiplying, subtracting, or adding the random numbers to the sound effect configuration parameters.

[0106] In the specific implementation, the random numbers contained in the multiple random arrays are different. The current sound effect configuration parameter set corresponding to the candidate audio is updated through multiple random arrays. The update results of the current sound effect configuration parameter set by multiple random arrays can be used as the updated multiple current sound effect configuration parameter sets.

[0107] As an example, a random array could be in the range of Random vectors obtained by sampling from a random distribution. Multiple random vectors obtained by sampling from a random distribution. The set of current sound effect configuration parameters corresponding to the candidate audio. By making adjustments, you can obtain multiple sets of adjusted sound effect configuration parameters. The specific process can be represented as follows:

[0108] .

[0109] As another example, a random array can be a random vector sampled from a Gaussian distribution. Multiple random vectors sampled from a Gaussian distribution... The set of current sound effect configuration parameters corresponding to the candidate audio. By making adjustments, you can obtain multiple sets of adjusted sound effect configuration parameters. The specific process can be represented as follows:

[0110] .

[0111] The technical solution of the above embodiment updates the current sound effect configuration parameter set corresponding to the candidate audio through multiple random arrays to obtain multiple updated current sound effect configuration parameter sets. By introducing uncertainty, it breaks out of the current local optimum and explores more possible combinations of sound effect parameters. During iterative processing, processed audio with different audio style tendencies can be obtained, thereby improving the fineness of audio style transfer.

[0112] In another embodiment, generating a random array includes: sampling a random vector from a Gaussian distribution; determining a noise vector based on the random vector and a changing step size; and generating a random array based on the noise vector.

[0113] The definition of a random vector can be found in the explanation of the above embodiments. For example, a random vector can be represented as [0.12, -0.85, ..., 0.37], where the number of elements in the random vector is the same as the number of sound effect configuration parameters in the current sound effect configuration parameter set corresponding to the candidate audio. Multiple random vectors can be used to represent multiple different random perturbations.

[0114] The step size can be used to control the search range of the parameters. This step size can be the variable step size in the covariance matrix adaptive evolution strategy algorithm described above, used to determine the magnitude of random perturbation in each iteration. For example, in the early stage of iteration, the step size can be increased to perform a wide-range search and discover potential solutions; in the later stage of iteration, the step size can be decreased to accurately approximate the optimal solution.

[0115] In the specific implementation, the noise vector can be determined based on the random variable and the change step size, and each noise vector is used to generate a random array. Based on multiple random arrays, the set of current sound effect configuration parameters corresponding to the candidate audio is updated to obtain multiple updated sets of current sound effect configuration parameters.

[0116] For example, suppose the change step size is , For the i-th random vector, the set of current sound effect configuration parameters corresponding to the candidate audio. By making adjustments, we can obtain the i-th updated set of current sound effect configuration parameters. The specific process can be represented as follows:

[0117] .

[0118] The technical solution of the above embodiments ensures that the direction of random disturbance conforms to acoustic laws through Gaussian distribution sampling, thus guaranteeing the real-time performance and reliability of the current sound effect configuration parameter set update and improving the efficiency and accuracy of audio style transfer.

[0119] In another embodiment, the sound effect configuration parameters are normalized parameters; updating the current sound effect configuration parameter set corresponding to the candidate audio to obtain multiple updated current sound effect configuration parameter sets includes: updating the current sound effect configuration parameter set corresponding to the candidate audio to obtain multiple new sound effect configuration parameter sets; pruning the sound effect configuration parameters in the multiple new sound effect configuration parameter sets that exceed the range of their respective normalized parameters to obtain multiple pruned sound effect configuration parameter sets; and using the multiple pruned sound effect configuration parameter sets as the updated multiple current sound effect configuration parameter sets.

[0120] In the specific implementation, if there are effect parameters in the new sound effect configuration parameter set that exceed the normalized parameter range, such as exceeding the range of [0, 1], then the sound effect configuration parameter is clipped to limit it to the normalized parameter range, and multiple clipped sound effect configuration parameter sets are used as multiple updated current sound effect configuration parameter sets.

[0121] For example, suppose that among the multiple sound effect configuration parameters included in the new sound effect configuration parameter set, there exists an effect parameter. If the value exceeds the range of the normalized parameters, then the effect parameters in the new sound effect configuration parameter set will be adjusted. The process of cutting can be represented as:

[0122] .

[0123] The technical solution of the above embodiments constrains out-of-range parameters within the normalized range through the trimming operation, avoiding errors or abnormal processing caused by illegal parameters, and ensuring the reliability and effectiveness of audio style transfer.

[0124] In another embodiment, selecting processed audio that satisfies a similarity condition with the target audio from multiple processed audios as candidate audio includes: inputting the processed audio and the target audio into a pre-trained sound effect feature extraction model to obtain sound effect features of the processed audio and the target audio; calculating the similarity between the sound effect features of the processed audio and the sound effect features of the target audio; and selecting processed audio that satisfies the similarity condition from multiple processed audios based on the calculated similarity as candidate audio.

[0125] Among them, sound effect features are features extracted from audio that can be used to quantify the sound effect configuration results.

[0126] In specific implementations, the pre-trained audio feature extraction model can include a mixing encoder model. This model extracts feature embeddings from each processed audio file, serving as the audio effect features for each processed audio file. It also extracts feature embeddings from the target audio file, serving as the audio effect features for the target audio file. The mixing encoder model can be a pre-trained multimodal neural network that maps the original audio signal into low-dimensional embedding vectors through end-to-end learning, serving as audio effect features. These features can simultaneously encode the audio's temporal dynamic characteristics (such as transient response), frequency domain structural characteristics (such as harmonic distribution), and audio effects processing traces (such as reverberation spatiality).

[0127] For example, the similarity between the sound effect features of the processed audio and the sound effect features of the target audio can be measured by cosine distance, which can also be called cosine similarity loss.

[0128] For example, the calculation of cosine distance can be expressed as:

[0129] ;

[0130] in, Cosine distance Let i be the sound effect feature of the i-th processed audio. The sound effects characteristics of the target audio.

[0131] In a specific implementation, the cosine distance between the sound effect features of any processed audio and the sound effect features of the target audio can be calculated as the similarity. If the cosine distance is less than a preset threshold, the processed audio is determined to meet the similarity condition, and the processed audio that meets the similarity condition is taken as the candidate audio.

[0132] The technical solution of the above embodiment improves the accuracy of audio style representation and comparison by extracting the sound effect features of the audio to determine the similarity of audio styles. This allows for the precise selection of processed audio with a similar audio style to the target audio as candidate audio, which is then used to generate multiple updated sets of current sound effect configuration parameters. This achieves real-time optimization and updating of audio parameters without the need to retrain the model, enabling rapid migration of the target audio style.

[0133] In another embodiment, after determining the target sound effect configuration parameter set for the initial audio, the target sound effect configuration parameter set can be applied to the initial audio to generate audio with the audio style of the target audio.

[0134] In practice, applying the target audio effect configuration parameter set to the initial audio can mean processing the initial audio through the audio processing function of the audio effect link to obtain the final audio, which is audio with the audio style of the target audio, thereby transferring the audio style of the target audio to the initial audio.

[0135] The technical solution of the above embodiments applies the target sound effect configuration parameter set obtained by iterative optimization to the initial audio, realizing the entire process from sound effect parameter optimization to audio style transfer result output, and improving the automation and real-time performance of audio style transfer.

[0136] For the convenience of those skilled in the art, Figure 3 An exemplary flowchart of an audio style transfer method is provided. Figure 3 As shown, the initial audio first passes through the audio effect link, loading multiple sets of current audio effect configuration parameters to generate multiple processed audio files. Then, the target audio and processed audio files are synchronously input into the mixing encoder model to generate their respective feature embeddings. The feature embeddings of the two are compared, and the cosine similarity is used to measure the style similarity. Candidate audio files are then selected from the multiple processed audio files. Next, an adaptive evolutionary strategy based on the covariance matrix is ​​used to generate updated sets of current audio effect configuration parameters corresponding to the candidate audio files. These updated sets are then re-input into the audio effect link to process the initial audio files again, entering the next iteration process. This process continues until the iteration stopping condition is met (e.g., reaching the required number of iterations or achieving a required cosine similarity). When the iteration stopping condition is met, the current audio effect configuration parameter set of the candidate audio files becomes the target audio effect configuration parameter set. The audio effect link configuration can then be output to process the initial audio files, thereby transferring the audio style of the target audio to the initial audio files, completing the entire audio style transfer process.

[0137] In another embodiment, such as Figure 4 As shown, an audio style transfer method is provided, which is applied to... Figure 1 Taking terminal 102 as an example, the explanation includes the following steps:

[0138] Step S402: Obtain the initial sound effect configuration parameter set.

[0139] Step S404: Based on the various random adjustment results of the initial sound effect configuration parameter set, generate multiple current sound effect configuration parameter sets.

[0140] Step S406: Adjust the parameters of each effect in each current sound effect configuration parameter set to obtain a set of sound effect configuration parameters corresponding to each current sound effect configuration parameter set.

[0141] In one embodiment, the parameter adjustment information includes a parameter adjustment range; adjusting each effect parameter included in each current sound effect configuration parameter set according to the parameter adjustment information of each effect parameter included in each current sound effect configuration parameter set includes: obtaining the parameter adjustment range of each effect parameter for each effect parameter included in each current sound effect configuration parameter set; and adjusting each effect parameter according to the parameter adjustment range of each effect parameter and the normalized parameter range to which each effect parameter belongs.

[0142] Step S408: Apply each set of sound effect configuration parameters to the initial audio to iteratively process the initial audio and obtain multiple processed audio files.

[0143] Step S410: Input the processed audio and the target audio into the pre-trained sound effect feature extraction model to obtain the sound effect features of the processed audio and the target audio; calculate the similarity between the sound effect features of the processed audio and the sound effect features of the target audio; select the processed audio that meets the similarity condition from multiple processed audios based on the calculated similarity as candidate audio.

[0144] Step S412: Determine whether the iteration stopping condition is met; if not, proceed to step S414; if met, proceed to step S416.

[0145] Step S414: Generate multiple updated sets of current sound effect configuration parameters based on the set of current sound effect configuration parameters corresponding to the candidate audio, and then return to step S406.

[0146] Step S416: Determine the target sound effect configuration parameter set for the initial audio based on the current sound effect configuration parameter set corresponding to the candidate audio.

[0147] Step S418: Apply the target sound effect configuration parameter set to the initial audio to generate audio with the audio style of the target audio.

[0148] It should be noted that the specific limitations of the above steps can be found in the specific limitations of an audio style transfer method described above.

[0149] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0150] Based on the same inventive concept, this application also provides an audio style transfer apparatus for implementing the audio style transfer method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more audio style transfer apparatus embodiments provided below can be found in the limitations of the audio style transfer method described above, and will not be repeated here.

[0151] In one exemplary embodiment, such as Figure 5 As shown, an audio style transfer apparatus is provided, comprising:

[0152] The acquisition module 510 is used to acquire multiple sets of current sound effect configuration parameters, which include multiple sound effect configuration parameters used to configure audio sound effects.

[0153] The processing module 520 is used to apply the multiple sets of current sound effect configuration parameters to the initial audio respectively to perform iterative processing on the initial audio to obtain multiple processed audio; and to select the processed audio that meets the similarity condition with the target audio from the multiple processed audio as candidate audio.

[0154] The update module 530 is used to generate multiple updated current sound effect configuration parameter sets based on the current sound effect configuration parameter set corresponding to the candidate audio if the iteration stop condition is not met, and to continue iterative processing of the initial audio using the multiple updated current sound effect configuration parameter sets.

[0155] The migration module 540 is used to determine the target sound effect configuration parameter set of the initial audio based on the current sound effect configuration parameter set corresponding to the candidate audio when the iteration stopping condition is met. The target sound effect configuration parameter set is used to migrate the audio style of the target audio to the initial audio.

[0156] In one embodiment, the acquisition module 510 can acquire an initial sound effect configuration parameter set; and generate the plurality of current sound effect configuration parameter sets based on the various random adjustment results of the initial sound effect configuration parameter set.

[0157] In one embodiment, the multiple sound effect configuration parameters include the effect parameters corresponding to each of the multiple sound effect processors; the processing module 520 can adjust each of the effect parameters in each current sound effect configuration parameter set to obtain a set of sound effect configuration parameters corresponding to each current sound effect configuration parameter set; and apply each set of sound effect configuration parameters to the initial audio.

[0158] In one embodiment, the processing module 520 can obtain the adjustment range of each effect parameter in each current sound effect configuration parameter set; and adjust the effect parameter according to the adjustment range of the effect parameter and the parameter adjustment ratio corresponding to the effect parameter.

[0159] In one embodiment, the update module 530 can generate a random array; the random array includes multiple random numbers, which correspond to each of the sound effect configuration parameters in the current sound effect configuration parameter set; based on the multiple random arrays, the current sound effect configuration parameter set corresponding to the candidate audio is updated to obtain the updated multiple current sound effect configuration parameter sets.

[0160] In one embodiment, the update module 530 can sample a random vector from a Gaussian distribution; determine a noise vector based on the random vector and the change step size; and generate the random array based on the noise vector.

[0161] In one embodiment, the update module 530 can update the current sound effect configuration parameter set corresponding to the candidate audio to obtain multiple new sound effect configuration parameter sets; trim the sound effect configuration parameters in the multiple new sound effect configuration parameter sets that exceed the normalized parameter range to obtain multiple trimmed sound effect configuration parameter sets; and use the multiple trimmed sound effect configuration parameter sets as the updated multiple current sound effect configuration parameter sets.

[0162] In one embodiment, the processing module 520 can input the processed audio and the target audio into a pre-trained sound effect feature extraction model to obtain the sound effect features of the processed audio and the sound effect features of the target audio; calculate the similarity between the sound effect features of the processed audio and the sound effect features of the target audio; and select processed audio that meets the similarity condition from the plurality of processed audios based on the calculated similarity as candidate audio.

[0163] In one embodiment, the migration module 540 can apply the target sound effect configuration parameter set to the initial audio to generate audio with the audio style of the target audio.

[0164] Each module in the aforementioned audio style transfer device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0165] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 6 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When executed by the processor, the computer program implements an audio style transfer method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0166] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0167] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0168] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0169] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0170] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0171] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0172] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0173] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An audio style transfer method, characterized in that, The method includes: Obtain multiple sets of current sound effect configuration parameters, wherein the current sound effect configuration parameter sets include multiple sound effect configuration parameters used for sound effect configuration of audio; The multiple sets of current sound effect configuration parameters are applied to the initial audio to iteratively process the initial audio and obtain multiple processed audios; the processed audios that meet the similarity condition with the target audio are selected from the multiple processed audios as candidate audios. If the iteration stopping condition is not met, multiple updated sets of current sound effect configuration parameters are generated based on the current sound effect configuration parameter set corresponding to the candidate audio, and the initial audio is further iteratively processed using the multiple updated sets of current sound effect configuration parameters. If the iteration stopping condition is met, the target sound effect configuration parameter set of the initial audio is determined according to the current sound effect configuration parameter set corresponding to the candidate audio. The target sound effect configuration parameter set is used to transfer the audio style of the target audio to the initial audio.

2. The method according to claim 1, characterized in that, The process of obtaining multiple sets of current sound effect configuration parameters includes: Get the initial sound effect configuration parameter set; Based on the various random adjustment results of the initial sound effect configuration parameter set, the multiple current sound effect configuration parameter sets are generated.

3. The method according to claim 1, characterized in that, The plurality of sound effect configuration parameters include the effect parameters corresponding to each of the plurality of sound effect processors; the step of applying the plurality of current sound effect configuration parameter sets to the initial audio includes: Adjust the parameters of each effect unit in each current sound effect configuration parameter set to obtain a set of sound effect configuration parameters corresponding to each current sound effect configuration parameter set; The sound effect configuration parameters of each group are applied to the initial audio respectively.

4. The method according to claim 3, characterized in that, The adjustment of each effect parameter in each of the current sound effect configuration parameter sets includes: Obtain the adjustment range of each effect parameter in each of the current sound effect configuration parameter sets; The effect parameters are adjusted according to the adjustment range of the effect parameters and the corresponding parameter adjustment ratio.

5. The method according to claim 1, characterized in that, The step of generating multiple updated sets of current sound effect configuration parameters based on the set of current sound effect configuration parameters corresponding to the candidate audio includes: Generate a random array; the random array includes multiple random numbers, each of which corresponds to a sound effect configuration parameter in the current sound effect configuration parameter set. Based on multiple random arrays, the set of current sound effect configuration parameters corresponding to the candidate audio is updated to obtain the updated set of multiple current sound effect configuration parameters.

6. The method according to claim 5, characterized in that, The generation of the random array includes: Random vectors are obtained by sampling from a Gaussian distribution; The noise vector is determined based on the random vector and the changing step size; The random array is generated based on the noise vector.

7. The method according to claim 5, characterized in that, The sound effect configuration parameters are normalized parameters; updating the current sound effect configuration parameter set corresponding to the candidate audio to obtain the updated multiple current sound effect configuration parameter sets includes: The current sound effect configuration parameter set corresponding to the candidate audio is updated to obtain multiple new sound effect configuration parameter sets; The sound effect configuration parameters that exceed the normalized parameter range in the multiple new sound effect configuration parameter sets are clipped to obtain multiple clipped sound effect configuration parameter sets. The multiple sets of trimmed sound effect configuration parameters are used as the updated sets of current sound effect configuration parameters.

8. The method according to claim 1, characterized in that, The step of selecting candidate audio from the plurality of processed audios that meets the similarity condition with the target audio includes: The processed audio and the target audio are input into a pre-trained sound effect feature extraction model to obtain the sound effect features of the processed audio and the sound effect features of the target audio. Calculate the similarity between the sound effect features of the processed audio and the sound effect features of the target audio; Based on the calculated similarity, processed audio that meets the similarity condition is selected from the plurality of processed audios as candidate audio.

9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: The target sound effect configuration parameter set is applied to the initial audio to generate audio with the audio style of the target audio.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the audio style transfer method according to any one of claims 1 to 9.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the audio style transfer method according to any one of claims 1 to 9.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the audio style transfer method according to any one of claims 1 to 9.