Audio processing method, device, computing equipment and medium

By automatically determining the tone similarity of the audio group and applying preset tuning parameters, the problem of low tuning efficiency in the prior art is solved, and the automation and efficiency improvement of audio processing is achieved.

CN115019814BActive Publication Date: 2025-05-23HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210524704.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-13
Publication Date
2025-05-23
Estimated Expiration
2042-05-13

AI Technical Summary

Technical Problem

In the prior art, the tuning process is less efficient and relies on manual operations of professional tuners, resulting in wasted time and energy.

Method used

By acquiring the to-process audio, determining the tone similarity between each candidate audio and the to-process audio in the candidate audio group, selecting the candidate audio group with the smallest sum of the to-color similarity is the target, and automatically tuning the to-process audio based on the preset tuning parameters of the target audio.

Benefits of technology

It improves the tuning efficiency of the audio processing process, reduces the need for manual operations, realizes automatic tuning of the audio to be processed, and enhances the accuracy of the processing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115019814B_ABST
    Figure CN115019814B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide an audio processing method, apparatus, computing device, and medium. After obtaining the audio to be processed, the present disclosure determines the timbre similarity between each candidate audio in the candidate audio group and the audio to be processed. Among them, the sum of the similarities between the audios included in the candidate audio group is the smallest among the sums of the similarities between the audios included in each audio group, so that the timbres of the candidate audios included in the candidate audio group are more diverse, so that the target audio can be determined from more diverse candidate audios to improve the accuracy of the determined target audio. Furthermore, based on the preset tuning parameters of the target audio, the audio to be processed is tuned to realize the automatic tuning process of the audio to be processed, without manual operation by relevant technicians, thereby improving the tuning efficiency of the audio processing process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the technical field of audio processing. More specifically, the embodiments of the present disclosure relate to an audio processing method, apparatus, computing device and medium. Background Art

[0002] This section is intended to provide a background or context to the embodiments of the present disclosure. No description herein is admitted to be prior art by virtue of its inclusion in this section.

[0003] Tuning is a form of music production that involves making post-production adjustments to make the audio sound more natural and more pleasing to the eye.

[0004] In the related art, audio is mainly tuned by professional tuners through a mixing console, which places extremely high professional requirements on the tuners and requires manual operation by the tuners, resulting in low audio tuning efficiency. Summary of the invention

[0005] In this context, embodiments of the present disclosure are intended to provide an audio processing method, apparatus, computing device, and medium to improve the tuning efficiency of the audio processing process.

[0006] In a first aspect of the embodiments of the present disclosure, there is provided an audio processing method, the method comprising:

[0007] Get the audio to be processed;

[0008] Determine the timbre similarity between each candidate audio in the candidate audio group and the audio to be processed, the candidate audio group is an audio group determined from multiple audio groups, the sum of the timbre similarities corresponding to the candidate audio group is the smallest among the sums of the timbre similarities corresponding to the various audio groups, the sum of the timbre similarities is the sum of the similarities between the audios included in the audio group, and each candidate audio corresponds to a preset tuning parameter;

[0009] Determine a target audio from multiple candidate audios based on the timbre similarity between the audio to be processed and each candidate audio;

[0010] The audio to be processed is tuned based on the preset tuning parameters of the target audio.

[0011] In one embodiment of the present disclosure, determining the timbre similarity between each candidate audio in the candidate audio group and the audio to be processed includes:

[0012] Obtaining the timbre characteristics of each candidate audio and the timbre characteristics of the audio to be processed;

[0013] For any candidate audio, based on the timbre features of the candidate audio and the timbre features of the audio to be processed, the timbre similarity between the candidate audio and the audio to be processed is determined.

[0014] In one embodiment of the present disclosure, obtaining the audio timbre of each candidate audio and the timbre feature of the audio to be processed includes:

[0015] For any audio, determine the vocal part of the audio;

[0016] Extract a target number of audio frames from the vocal part of the audio;

[0017] Based on the timbre characteristics of the target number of audio frames, a timbre characteristic of the audio is determined.

[0018] In one embodiment of the present disclosure, extracting a target number of audio frames from a vocal part of an audio comprises:

[0019] Determine the sampling frequency based on the duration of the vocal part of the audio;

[0020] According to the sampling frequency, audio frames are extracted from the vocal part of the audio to obtain the target number of audio frames.

[0021] In one embodiment of the present disclosure, determining the timbre features of the audio based on the timbre features of the target number of audio frames includes:

[0022] Acquire the timbre features of each audio frame to obtain the target number of timbre features;

[0023] The timbre features of the target number are averaged to obtain the timbre features of the audio.

[0024] In one embodiment of the present disclosure, for any candidate audio, determining the timbre similarity between the candidate audio and the audio to be processed based on the timbre features of the candidate audio and the timbre features of the audio to be processed includes:

[0025] For any candidate audio, a cosine distance between the timbre features of the candidate audio and the timbre features of the audio to be processed is determined, and the determined cosine distance is used as the timbre similarity between the candidate audio and the audio to be processed.

[0026] In one embodiment of the present disclosure, determining a target audio from a plurality of candidate audios based on the timbre similarity between the audio to be processed and each candidate audio includes:

[0027] The candidate audio with the greatest timbre similarity to the audio to be processed among the multiple candidate audios is determined as the target audio.

[0028] In one embodiment of the present disclosure, the process of determining the candidate audio group includes:

[0029] Get multiple audio samples;

[0030] Obtain the timbre characteristics of each sample audio;

[0031] Based on the timbre characteristics of the plurality of sample audios, determining the timbre similarity between every two sample audios;

[0032] Taking a set number of sample audios as an audio group, determining the sum of timbre similarities between two of the set number of sample audios included in each audio group;

[0033] The audio group with the smallest sum of timbre similarities is taken as a candidate audio group.

[0034] In one embodiment of the present disclosure, after acquiring a plurality of sample audios, the method further includes:

[0035] Preprocessing is performed on the plurality of sample audios, where the preprocessing includes at least one of noise reduction processing, de-essing processing, and volume normalization processing.

[0036] In one embodiment of the present disclosure, after taking the audio group with the smallest sum of timbre similarities as the candidate audio group, the method further includes:

[0037] Dynamic EQ tuning parameters and static EQ tuning parameters of the audio included in the candidate audio group are obtained as preset tuning parameters.

[0038] In one embodiment of the present disclosure, the audio is dry audio, and the timbre features are composed of features of a target dimension in Mel-frequency cepstral coefficient (MFCC) features.

[0039] In a second aspect of the embodiments of the present disclosure, there is provided an audio processing device, the device comprising:

[0040] An acquisition module, used to acquire audio to be processed;

[0041] A similarity determination module is used to determine the timbre similarity between each candidate audio in a candidate audio group and the audio to be processed, wherein the candidate audio group is an audio group determined from multiple audio groups, the sum of the timbre similarities corresponding to the candidate audio group is the smallest among the sums of the timbre similarities corresponding to the various audio groups, the sum of the timbre similarities is the sum of the similarities between the audios included in the audio group, and each candidate audio corresponds to a preset tuning parameter;

[0042] An audio determination module, used to determine a target audio from a plurality of candidate audios based on the timbre similarity between the audio to be processed and each candidate audio;

[0043] The processing module is used to perform tuning processing on the audio to be processed based on the preset tuning parameters of the target audio.

[0044] In one embodiment of the present disclosure, the similarity determination module, when used to determine the timbre similarity between each candidate audio in the candidate audio group and the audio to be processed, includes:

[0045] An acquisition submodule, used to acquire the timbre characteristics of each candidate audio and the timbre characteristics of the audio to be processed;

[0046] The determination submodule is used to determine the timbre similarity between any candidate audio and the audio to be processed based on the timbre features of the candidate audio and the timbre features of the audio to be processed.

[0047] In one embodiment of the present disclosure, the acquisition submodule, when used to acquire the audio timbre of each candidate audio and the timbre feature of the audio to be processed, includes:

[0048] A determination unit, used for determining a vocal part of any audio;

[0049] An extraction unit, used for extracting a target number of audio frames from the vocal part of the audio;

[0050] The determination unit is further used to determine the timbre features of the audio based on the timbre features of the target number of audio frames.

[0051] In one embodiment of the present disclosure, the extraction unit, when used to extract a target number of audio frames from a vocal part of an audio, is used to:

[0052] Determine the sampling frequency based on the duration of the vocal part of the audio;

[0053] According to the sampling frequency, audio frames are extracted from the vocal part of the audio to obtain the target number of audio frames.

[0054] In one embodiment of the present disclosure, the determining unit, when determining the timbre features of the audio based on the timbre features of the target number of audio frames, is configured to:

[0055] Acquire the timbre features of each audio frame to obtain the target number of timbre features;

[0056] The timbre features of the target number are averaged to obtain the timbre features of the audio.

[0057] In one embodiment of the present disclosure, the determination submodule, when used to determine the timbre similarity between the candidate audio and the audio to be processed based on the timbre features of the candidate audio and the timbre features of the audio to be processed for any candidate audio, is used to:

[0058] For any candidate audio, a cosine distance between the timbre features of the candidate audio and the timbre features of the audio to be processed is determined, and the determined cosine distance is used as the timbre similarity between the candidate audio and the audio to be processed.

[0059] In one embodiment of the present disclosure, the audio determination module, when used to determine the target audio from multiple candidate audios based on the timbre similarity between the audio to be processed and each candidate audio, is used to:

[0060] The candidate audio with the greatest timbre similarity to the audio to be processed among the multiple candidate audios is determined as the target audio.

[0061] In one embodiment of the present disclosure, the audio processing device further includes an audio group determination module, which is used to determine a candidate audio group from a plurality of audio groups; the audio group determination module includes:

[0062] The audio acquisition submodule is used to acquire multiple sample audios;

[0063] A feature acquisition submodule is used to obtain the timbre features of each sample audio;

[0064] A similarity determination submodule, used for determining the timbre similarity between every two sample audios based on the timbre features of the multiple sample audios;

[0065] The similarity determination submodule is further used to take a set number of sample audios as an audio group and determine the sum of the timbre similarities between the set number of sample audios included in each audio group;

[0066] The audio group determination submodule is used to select the audio group with the smallest sum of timbre similarities as a candidate audio group.

[0067] In one embodiment of the present disclosure, the audio group determination module further includes:

[0068] The processing submodule is used to pre-process the multiple sample audios, where the pre-processing includes at least one of noise reduction processing, de-essing processing and volume normalization processing.

[0069] In one embodiment of the present disclosure, the audio group determination module further includes:

[0070] The parameter acquisition submodule is used to acquire dynamic EQ tuning parameters and static EQ tuning parameters of the audio included in the candidate audio group as preset tuning parameters.

[0071] In one embodiment of the present disclosure, the audio is dry audio, and the timbre features are composed of features of a target dimension in Mel-frequency cepstral coefficient (MFCC) features.

[0072] In a third aspect of the embodiments of the present disclosure, a computing device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the operations performed by the audio processing method provided in the first aspect and any embodiment of the first aspect are implemented.

[0073] In a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which a program is stored, and the program is used by a processor to execute the operations performed by the audio processing method provided in the first aspect and any embodiment of the first aspect.

[0074] In a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, including a computer program, which, when executed by a processor, implements the operations performed by the audio processing method provided in the first aspect and any embodiment of the first aspect.

[0075] According to the audio processing method, apparatus, computing device and medium provided in the embodiments of the present disclosure, after obtaining the audio to be processed, the timbre similarity between each candidate audio in the candidate audio group and the audio to be processed is determined, wherein the sum of the similarities between the audios included in the candidate audio group is the smallest of the sums of the similarities between the audios included in each audio group, thereby making the timbre of the candidate audios included in the candidate audio group more diverse, thereby making it possible to determine the target audio from more diverse candidate audios, so as to improve the accuracy of the determined target audio, and then based on the preset tuning parameters of the target audio, perform tuning processing on the audio to be processed, thereby realizing an automatic tuning process for the audio to be processed, without the need for manual operation by relevant technical personnel, thereby improving the tuning efficiency of the audio processing process. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, in which:

[0077] Figure 1 is a flowchart of an audio processing method according to an exemplary embodiment of the present disclosure;

[0078] Figure 2 is a flowchart of a process for determining a candidate audio group according to an exemplary embodiment of the present disclosure;

[0079] Figure 3 is a schematic diagram of a static EQ tuning parameter according to an exemplary embodiment of the present disclosure;

[0080] Figure 4 is a schematic diagram of a dynamic EQ tuning parameter according to an exemplary embodiment of the present disclosure;

[0081] Figure 5 is a schematic diagram of another dynamic EQ tuning parameter according to an exemplary embodiment of the present disclosure;

[0082] Figure 6 is a block diagram of an audio processing device according to an exemplary embodiment of the present disclosure;

[0083] Figure 7 is a schematic diagram of a computer-readable storage medium according to an exemplary embodiment of the present disclosure;

[0084] Figure 8 is a structural schematic diagram of a computing device according to an exemplary embodiment of the present disclosure;

[0085] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts. DETAILED DESCRIPTION

[0086] The principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present disclosure, and are not intended to limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.

[0087] Those skilled in the art will appreciate that the embodiments of the present disclosure may be implemented as a system, device, apparatus, method or computer program product. Therefore, the present disclosure may be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0088] It should be understood herein that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction rather than having any limiting meaning.

[0089] According to the embodiments of the present disclosure, an audio processing method, apparatus, computing device and medium are proposed. The above method can be executed by a computing device, and is used to tune the acquired audio to be processed after the computing device acquires the audio to be processed. The computing device can be a server, such as a server, multiple servers, a server cluster, a cloud computing platform, etc. Optionally, the computing device can also be a terminal device, such as a smart phone, a tablet computer, a desktop computer, a portable computer, a smart speaker, etc. The present disclosure does not limit the device type and number of computing devices.

[0090] For example, the singer's vocals can be recorded through an audio recording component, and the recorded vocals can be used as audio to be processed, so that the computing device can tune the audio to be processed through the audio processing method provided by the present disclosure.

[0091] The audio recording component may be a separate audio recording device, such as a recorder, etc. Optionally, the audio recording component may also be a component built into other devices, such as a microphone, etc. The present disclosure does not limit the specific type of the audio recording component.

[0092] The above is an introduction to the application scenarios of the present disclosure. It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principle of the present disclosure, and the embodiments of the present disclosure are not limited in this respect. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.

[0093] After introducing the application scenarios of the present disclosure, the principle and spirit of the present disclosure are explained in detail below with reference to several representative embodiments of the present disclosure.

[0094] See also Figure 1 , Figure 1 is a flowchart of an audio processing method according to an exemplary embodiment of the present disclosure, the method comprising:

[0095] Step 101: Obtain audio to be processed.

[0096] Step 102: determine the timbre similarity between each candidate audio in the candidate audio group and the audio to be processed, where the candidate audio group is an audio group determined from multiple audio groups, the sum of the timbre similarities corresponding to the candidate audio group is the smallest among the sums of the timbre similarities corresponding to the various audio groups, the sum of the timbre similarities is the sum of the similarities between the audios included in the audio group, and each candidate audio corresponds to a preset tuning parameter.

[0097] Step 103: Determine a target audio from multiple candidate audios based on the timbre similarity between the audio to be processed and each candidate audio.

[0098] It should be noted that, since each candidate audio in the candidate audio group corresponds to a preset tuning parameter, the preset tuning parameter corresponding to the target audio determined from the multiple candidate audios is also determined.

[0099] Step 104: Perform tuning processing on the audio to be processed based on the preset tuning parameters of the target audio.

[0100] The present invention determines the timbre similarity between each candidate audio in the candidate audio group and the audio to be processed after acquiring the audio to be processed, wherein the sum of the similarities between the audios included in the candidate audio group is the smallest of the sums of the similarities between the audios included in each audio group, thereby making the timbre of the candidate audios included in the candidate audio group more diverse, thereby making it possible to determine the target audio from more diverse candidate audios, so as to improve the accuracy of the determined target audio, and then perform tuning processing on the audio to be processed based on the preset tuning parameters of the target audio, thereby realizing an automatic tuning process for the audio to be processed, without the need for manual operation by relevant technical personnel, thereby improving the tuning efficiency of the audio processing process.

[0101] After introducing the basic principles of the present disclosure, various non-limiting embodiments of the present disclosure are described in detail below.

[0102] In some embodiments, the computing device may have a built-in or external audio recording component. Accordingly, for step 101, when obtaining the audio to be processed, there may be the following two methods:

[0103] When the audio recording component is built in the computing device, the computing device can directly obtain the audio to be processed recorded by the audio recording component.

[0104] In the case where the audio recording component is externally connected to a computing device, after recording the audio to be processed, the audio recording component can send the recorded audio to be processed to the computing device so that the computing device can obtain the audio to be processed.

[0105] The audio to be processed may be dry audio, that is, recorded audio without any audio post-processing, for example, pure human voice without music.

[0106] After the audio to be processed is obtained, the obtained audio to be processed can be processed.

[0107] In some embodiments, for step 102, when determining the timbre similarity between each candidate audio in the candidate audio group and the audio to be processed, the following steps may be included:

[0108] Step 1021: Acquire the timbre features of each candidate audio and the timbre features of the audio to be processed.

[0109] For any audio among multiple candidate audios and the audio to be processed, the timbre characteristics of the audio can be obtained by the following steps:

[0110] Step 1021 - 1 : For any audio, determine the vocal part of the audio.

[0111] In a possible implementation, a voice activity detection (VAD) algorithm may be used to filter out all time periods containing human voices, that is, the human voice part of the audio.

[0112] Optionally, a VAD algorithm based on a signal-to-noise ratio, a VAD algorithm based on a deep neural network (DNN), an energy-based VAD algorithm, a decoder-based VAD algorithm, etc. may be used. The present disclosure does not limit which specific method is used.

[0113] Step 1021 - 2: Extract a target number of audio frames from the vocal part of the audio.

[0114] In a possible implementation, the sampling frequency may be determined based on the duration of the vocal portion of the audio; and then, audio frames may be extracted from the vocal portion of the audio according to the sampling frequency to obtain a target number of audio frames.

[0115] Optionally, when the sampling frequency is determined based on the duration of the vocal part, the ratio of the target number to the duration of the vocal part may be determined as the sampling frequency.

[0116] Among them, the target number can be any value. In order to calculate the timbre features of a complete audio, it is generally necessary to obtain an audio clip of sufficient length to ensure that the determined timbre features can represent the audio features of the entire audio. In one possible implementation, a sufficient number of audio frames need to be extracted to ensure that an audio clip of sufficient length can be obtained for determining the timbre features. For example, the audio clip used to determine the timbre features can be 3 seconds or more to ensure that the timbre features determined based on the audio clip of 3 seconds or more can represent the audio features of the entire audio. In the case where the audio clip is 3 seconds long, combined with an audio clip of about 0.023 seconds per audio frame, it can be determined that the target number can be 130. Optionally, the target number can also be other values ​​greater than 130. The present disclosure does not limit the specific value of the target number.

[0117] Taking the target number as 130 as an example, when the vocal part of the audio is a seconds, 130 / a can be used as the sampling frequency for extracting the target number of audio frames, that is, an audio frame can be extracted every a / 130 seconds in the vocal part of a seconds, so that 130 audio frames can be extracted from a seconds of audio (each audio frame is about 0.023 seconds of audio, and 130 audio frames are about 3 seconds of audio), so as to determine the timbre characteristics based on the extracted target number of audio frames, that is, 130 audio frames.

[0118] Through the above process, multiple audio frames located at different time points can be extracted, and the timbre features can be determined based on the extracted audio frames. This can take into account the timbre differences caused by the pronunciation of different parts of the entire audio (including the main chorus, true and false voices, and different lyrics), making the subsequently determined timbre features more accurate.

[0119] Step 1021 - 3: Determine the timbre features of the audio based on the timbre features of the target number of audio frames.

[0120] In a possible implementation, the timbre features of each audio frame may be acquired to obtain a target number of timbre features; then, an average process is performed based on the target number of timbre features to obtain the timbre features of the audio.

[0121] The timbre feature is composed of the features of the target dimension in the Mel Frequency Cepstrum Coefficient (MFCC) feature. Optionally, the timbre feature can also be other types of features, and the present disclosure does not limit the specific types of the timbre feature.

[0122] Taking the timbre feature composed of the features of the target dimension in the MFCC feature as an example, the timbre feature can be obtained in the following way:

[0123] For any audio frame, the MFCC features of the audio frame are determined, so as to obtain the features of the target dimension from the determined MFCC features as the timbre features of the audio frame.

[0124] The process of determining the MFCC features of the audio frame may be:

[0125] The audio frame is windowed, and the windowed audio frame is fast Fourier transformed to obtain the spectrum of the audio frame; the obtained spectrum is passed through a set of Mel-scale triangular filter groups to obtain the energy value of the audio frame in the frequency band corresponding to each filter (for example, when the filter group includes 22 filters, 22 energy values ​​can be obtained); the obtained energy value is taken logarithm, so as to bring the logarithmic energy corresponding to each filter into discrete cosine transform to obtain Mel-scale coefficient (Mel-scale Cepstrum) parameters; the first-order differential and second-order differential of the obtained Mel-scale Cepstrum parameters are calculated, so as to use the combination of the Mel-scale Cepstrum parameters, the first-order differential of the Mel-scale Cepstrum parameters and the second-order differential of the Mel-scale Cepstrum parameters as the MFCC features of the audio frame.

[0126] Through the above process, 39-dimensional MFCC features can be obtained, so that timbre features can be obtained based on the obtained MFCC features.

[0127] Optionally, the target dimension may be any dimension, for example, the target dimension may be the 2nd to 14th dimension (a total of 13 dimensions). Taking the target dimension as the 2nd to 14th dimension as an example, after obtaining the 39-dimensional MFCC features through the above process, the 2nd to 14th dimension features may be obtained from the 39-dimensional MFCC features as the timbre features of the audio frame.

[0128] The above only takes the process of obtaining the timbre features of one audio frame as an example for explanation. The process of obtaining the timbre features of other audio frames is similar to it and will not be described in detail here.

[0129] It should be noted that the present disclosure uses the 2nd to 14th dimensions (a total of 13 dimensions) of the MFCC feature as the timbre feature, so that the number of dimensions included in the timbre feature is relatively small, that is, a relatively shallow feature is used as the timbre feature, so that the subsequent processing based on the timbre feature is simpler and more efficient. In addition, the first dimension of the MFCC feature represents the volume, and the volume cannot be used as a timbre feature, while the first dozen or so dimensions of the general MFCC feature are sufficient to represent the timbre feature. Therefore, the timbre of the audio frame can be expressed while ensuring that the number of dimensions of the timbre feature is relatively small, while reducing the computing pressure of the computing device and ensuring the accuracy of the acquired timbre feature. Moreover, the MFCC feature can better represent the original timbre of the human voice and the timbre decoration produced by the recording device, and is more suitable for tuning, thereby improving the accuracy of the subsequent tuning process.

[0130] Through the above process, the timbre features corresponding to the target number of audio frames can be obtained, that is, the target number of timbre features. In the case where the timbre features are the 2nd to 14th dimensions of the MFCC features, each timbre feature is a 13-dimensional vector, so the timbre features of the target number can be arithmetic averaged, and the result of the arithmetic average is used as the timbre feature of the audio.

[0131] Step 1022: For any candidate audio, determine the timbre similarity between the candidate audio and the audio to be processed based on the timbre features of the candidate audio and the timbre features of the audio to be processed.

[0132] In a possible implementation, for any candidate audio, a cosine distance between the timbre features of the candidate audio and the timbre features of the audio to be processed may be determined, and the determined cosine distance is used as the timbre similarity between the candidate audio and the audio to be processed.

[0133] Among them, when the cosine distance is used as the timbre similarity, the value range of the timbre similarity is [-1,1]. The timbre similarity of two completely identical audios is 1, and the timbre similarity of two audios with completely different timbres is close to -1. The closer the timbre similarity between two audios is to 1, the smaller the timbre distance between the two audios is, that is, the more similar the timbres of the two audios are; conversely, the closer the timbre similarity between two audios is to -1, the greater the timbre distance between the two audios is, that is, the less similar the timbres of the two audios are.

[0134] It should be noted that the above is only an exemplary method for determining the timbre similarity between the candidate audio and the audio to be processed. In more possible implementations, other methods can also be used to determine the timbre similarity. The present disclosure does not limit the specific method for determining the timbre similarity.

[0135] Through the above process, the timbre similarity between each candidate audio and the audio to be processed can be determined, and thus the target audio can be determined based on the determined timbre similarity.

[0136] In some embodiments, for step 103, when determining the target audio from multiple candidate audios based on the timbre similarity between the audio to be processed and each candidate audio, it can be implemented in the following manner:

[0137] In a possible implementation, the candidate audio with the greatest timbre similarity to the audio to be processed among multiple candidate audios is determined as the target audio.

[0138] Since the preset tuning parameters of each candidate audio included in the candidate audio group are determined, the preset tuning parameters of the determined target audio are also determined. Therefore, after the target audio is determined in step 103, step 104 can be used to perform tuning processing on the audio to be processed based on the preset tuning parameters of the target audio.

[0139] It should be noted that the above process is about how to determine the target audio that is most similar to the timbre of the audio to be processed from the candidate audio group after obtaining the audio to be processed, so as to perform tuning processing on the audio to be processed based on the target audio, wherein the candidate audio group can be predetermined, see Figure 2 , Figure 2 : is a flowchart of a process for determining a candidate audio group according to an exemplary embodiment of the present disclosure, and the process for determining a candidate audio group includes:

[0140] Step 201: Acquire multiple sample audios.

[0141] The sample audio may be a dry audio, that is, n dry audios may be obtained as n sample audios.

[0142] It should be noted that, since the timbre similarity between the n sample audios needs to be calculated later, in order to avoid memory overflow of the computing device, the value of n is generally taken to be less than 80.

[0143] Optionally, after acquiring the plurality of sample audios, the plurality of sample audios may be preprocessed, wherein the preprocessing may include at least one of noise reduction processing, de-essing processing and volume normalization processing.

[0144] For any sample audio, when performing noise reduction processing on the sample audio, a noise reduction algorithm based on a linear filter, a noise reduction algorithm based on a nonlinear filter, or a noise reduction algorithm with a neural network algorithm as the core can be used. The present disclosure does not limit the specific noise reduction method to be used.

[0145] By performing noise reduction processing on the sample audio, the ambient noise and recording background noise in the sample audio can be removed, thereby improving the audio quality of the sample audio, and making the subsequent timbre features obtained based on the sample audio more accurate.

[0146] De-essing the sample audio may be performed through a de-essing plug-in or a dynamic equalizer of a computing device. In one possible implementation, the de-essing plug-in or the dynamic equalizer may automatically identify the sibilance band (generally a frequency band range of 2 to 10 kHz) in the sample audio. When the level of the sibilance band exceeds a set threshold, the de-essing plug-in or the dynamic equalizer may automatically attenuate the sibilance band to achieve de-essing of the sample audio. The set threshold is an arbitrary value, and the present disclosure does not limit the specific value of the set threshold.

[0147] By de-essing the sample audio, the dry singing defect in the sample audio can be reduced to improve the audio quality of the sample audio, thereby making the subsequent timbre characteristics obtained based on the sample audio more accurate.

[0148] When the volume of the sample audio is normalized, the volume of the sample audio can be normalized to 18LUFS. For example, the peak value normalization method can be used to adjust the position with the maximum volume of the sample audio to a specific size, that is, the maximum volume of the sample audio is adjusted to 18FLUS, and the volume of other positions is increased / decreased accordingly, thereby achieving the volume normalization of the sample audio.

[0149] By normalizing the sample audio, the volume of the dry singing voice of multiple sample audios can be unified, which is convenient for subsequent operations.

[0150] Step 202: Acquire the timbre characteristics of each audio sample.

[0151] It should be noted that the process of obtaining the timbre features of each sample audio may refer to the process of obtaining the audio timbre features in step 1021, which will not be repeated here.

[0152] Step 203: Determine the timbre similarity between every two sample audios based on the timbre features of the multiple sample audios.

[0153] In a possible implementation, for any two sample audios among the multiple sample audios, a cosine distance between the timbre features of the two sample audios may be determined, and the determined cosine distance is used as the timbre similarity between the two sample audios.

[0154] It should be noted that, taking the number of sample audios as n as an example, when calculating the timbre similarity between the n sample audios, a total of Calculation times.

[0155] Optionally, the n audio samples may be numbered, for example, the n audio samples are numbered from 0 to n, and when the timbre similarity of any two audio samples is calculated, the numbers corresponding to the two audio samples may be recorded, so that the two audio samples to which the timbre similarity corresponds may be determined based on the recorded numbers. For example, (2,27) represents the timbre similarity between the 2nd audio sample and the 27th audio sample.

[0156] Step 204: Take the set number of sample audios as an audio group, and determine the sum of the timbre similarities between the set number of sample audios included in each audio group.

[0157] The set number can be any value. Let the set number be k, then n audio samples can form audio groups.

[0158] Generally speaking, in order to avoid memory overflow of the computing device, the set number is generally set to 6. Optionally, the set number can also be a value smaller than 6. The present disclosure does not limit the specific value of the set number.

[0159] Optionally, after determining each audio group, the numbers of the sample audios included in each audio group may be recorded. For example, when the number is set to 6, (2, 11, 17, 32, 35, 48) represents an audio group consisting of the 2nd, 11th, 17th, 32nd, 35th, and 48th sample audios.

[0160] It should be noted that, when the number is set to k, each audio group corresponds to In one possible implementation, for any audio group, this timbre similarity can be calculated. The sum of the timbre similarities is taken as the sum of the timbre similarities corresponding to the audio group.

[0161] For example, if the number is set to 6, each audio group has If there are 15 timbre similarities, then for any audio group, the sum of these 15 timbre similarities can be calculated as the sum of the timbre similarities corresponding to the audio group.

[0162] Step 205: The audio group with the smallest sum of timbre similarities is taken as a candidate audio group.

[0163] It should be noted that the smaller the sum of the timbre similarities corresponding to the audio groups, the less similar the timbre of the sample audios included in the audio groups. Therefore, determining the audio group with the smallest sum of timbre similarities as the candidate audio group can ensure that the timbre difference of the sample audios included in the candidate audio group is the greatest, so that the candidate audio group can cover audios of various different timbres, making the candidate audio group more representative.

[0164] Optionally, after the candidate audio group is determined, the sample audios (ie, the candidate audios) included in the determined candidate audio group may be tuned to obtain preset tuning parameters of the candidate audios.

[0165] In a possible implementation, dynamic EQ tuning parameters and static EQ tuning parameters of the audio included in the candidate audio group may be obtained as preset tuning parameters.

[0166] Among them, the static equalizer (EQ) tuning parameter refers to the gain or attenuation value of a certain frequency band of audio that does not change dynamically with the signal level, and the dynamic EQ tuning parameter refers to the gain or attenuation value of a certain frequency band of audio that changes dynamically according to the signal level.

[0167] Optionally, a professional (such as a tuner) may perform EQ tuning on the candidate audios included in the candidate audio group, including static EQ tuning and dynamic EQ tuning. Figure 3 , Figure 3 is a schematic diagram of a static EQ tuning parameter according to an exemplary embodiment of the present disclosure. Figure 3 The audio shown has a static EQ tuning parameter of 400Hz as the center, a Q value of 1, and an attenuation of 5dB. Figure 4 , Figure 4 is a schematic diagram of a dynamic EQ tuning parameter according to an exemplary embodiment of the present disclosure. Figure 4The audio shown in the figure has a dynamic EQ tuning parameter of 500 Hz as the center, a Q value of 1, and an actual attenuation amount (curve 401) that varies according to the signal level in the frequency band. The upper limit of the attenuation amount is curve 402, and the lower limit of the attenuation amount is curve 403; see Figure 5 , Figure 5 is a schematic diagram of another dynamic EQ tuning parameter according to an exemplary embodiment of the present disclosure. Figure 5 The audio shown in the figure has a dynamic EQ tuning parameter of 500 Hz as the center, a Q value of 1, and an actual attenuation (curve 501) that varies according to the signal level in the frequency band. The upper limit of the attenuation is curve 502, and the lower limit of the attenuation is curve 503. The difference is that Figure 4 The signal level in is high, Figure 5 The signal level in is low, so Figure 4 The attenuation in is large. Figure 5 The attenuation in is small.

[0168] Due to the complexity and diversity of human voice timbre, timbre similarity is only a value, and the ability to distinguish timbre differences in different frequency bands is limited. The dynamic EQ tuning parameter is a tuning parameter suitable for processing defects caused by changes in frequency band levels, and can achieve tuning processing for audio with similar overall timbre but different timbre in different frequency bands. By combining the static EQ tuning parameter and the dynamic EQ tuning parameter as the preset tuning parameter, the static EQ tuning parameter is suitable for processing the inherent frequency band defects of the timbre, thereby increasing the universality of the tuning parameters disclosed in the present invention.

[0169] It should be noted that the above-mentioned process of determining the candidate audio group and obtaining the preset tuning parameters of the candidate audio can be used as a preparatory stage of the audio processing process, that is, the candidate audio group can be determined in advance, and the preset tuning parameters of the candidate audio included in the candidate audio group can be obtained, so that when the audio to be processed is obtained, it can be directly processed as follows. Figure 1 The process shown is used to tune the audio to be processed.

[0170] For example, if the number is set to 6, Figure 1 The processing process shown can determine the timbre similarities between the audio to be processed and the six candidate audios respectively, and obtain six timbre similarities. Then, based on the six timbre similarities, the target audio with the greatest timbre similarity to the audio to be processed is found from the six candidate audios, and the preset tuning parameters of the target audio are used as the preset tuning parameters of the audio to be processed to tune the audio to be processed, thereby realizing automatic tuning of the audio to be processed.

[0171] In addition, it should be noted that since the process of determining the candidate audio group and obtaining the preset tuning parameters of the candidate audio belongs to the preparation stage of the audio processing process, the process of determining the candidate audio group and obtaining the preset tuning parameters of the candidate audio will not affect the Figure 1 The audio processing efficiency shown is achieved.

[0172] Based on the above embodiments, the audio processing method provided by the present disclosure can at least bring the following effects:

[0173] The present invention utilizes the features of the target dimension in the MFCC features as the timbre features, which are compatible with the timbre differences caused by different human voices and different recording devices. At the same time, the timbre differences caused by the main and chorus, true and false voices, and different pronunciations of lyrics during the singing process are taken into account, making the calculation of the timbre features more accurate.

[0174] In addition, adding dynamic EQ tuning parameters to the tuning parameters used can accommodate the situation where the timbre of the audio to be processed and the target audio are similar but there are slight differences in different frequency bands.

[0175] In addition, the scheme for determining candidate audio groups has been optimized, and the sum of the timbre similarities corresponding to each audio group has been compared, so that the audio group with the smallest sum of timbre similarities is selected as the candidate audio group. The timbre similarity between each pair of candidate audios in the candidate audio group is the smallest, so that the candidate audio group can cover audios of various different timbres, making the candidate audio group more representative.

[0176] After introducing the audio processing method according to the exemplary embodiment of the present disclosure, the structure of the audio processing apparatus according to the exemplary embodiment of the present disclosure and the computing device for implementing the audio processing method will be described next.

[0177] See also Figure 6 , Figure 6 is a block diagram of an audio processing device according to an exemplary embodiment of the present disclosure, the device comprising:

[0178] An acquisition module 601 is used to acquire audio to be processed;

[0179] A similarity determination module 602 is used to determine the timbre similarity between each candidate audio in the candidate audio group and the audio to be processed, wherein the candidate audio group is an audio group determined from multiple audio groups, the sum of the timbre similarities corresponding to the candidate audio group is the smallest among the sums of the timbre similarities corresponding to the various audio groups, the sum of the timbre similarities is the sum of the similarities between the audios included in the audio group, and each candidate audio corresponds to a preset tuning parameter;

[0180] The audio determination module 603 is used to determine the target audio from multiple candidate audios based on the timbre similarity between the audio to be processed and each candidate audio;

[0181] The processing module 604 is used to perform tuning processing on the audio to be processed based on the preset tuning parameters of the target audio.

[0182] In one embodiment of the present disclosure, the similarity determination module 602, when used to determine the timbre similarity between each candidate audio in the candidate audio group and the audio to be processed, includes:

[0183] An acquisition submodule, used to acquire the timbre characteristics of each candidate audio and the timbre characteristics of the audio to be processed;

[0184] The determination submodule is used to determine the timbre similarity between any candidate audio and the audio to be processed based on the timbre features of the candidate audio and the timbre features of the audio to be processed.

[0185] In one embodiment of the present disclosure, the acquisition submodule, when used to acquire the audio timbre of each candidate audio and the timbre feature of the audio to be processed, includes:

[0186] A determination unit, used for determining a vocal part of any audio;

[0187] An extraction unit, used for extracting a target number of audio frames from the vocal part of the audio;

[0188] The determination unit is further used to determine the timbre features of the audio based on the timbre features of the target number of audio frames.

[0189] In one embodiment of the present disclosure, the extraction unit, when used to extract a target number of audio frames from a vocal part of an audio, is used to:

[0190] Determine the sampling frequency based on the duration of the vocal part of the audio;

[0191] According to the sampling frequency, audio frames are extracted from the vocal part of the audio to obtain the target number of audio frames.

[0192] In one embodiment of the present disclosure, the determining unit, when determining the timbre features of the audio based on the timbre features of the target number of audio frames, is configured to:

[0193] Acquire the timbre features of each audio frame to obtain the target number of timbre features;

[0194] The timbre features of the target number are averaged to obtain the timbre features of the audio.

[0195] In one embodiment of the present disclosure, the determination submodule, when used to determine the timbre similarity between the candidate audio and the audio to be processed based on the timbre features of the candidate audio and the timbre features of the audio to be processed for any candidate audio, is used to:

[0196] For any candidate audio, a cosine distance between the timbre features of the candidate audio and the timbre features of the audio to be processed is determined, and the determined cosine distance is used as the timbre similarity between the candidate audio and the audio to be processed.

[0197] In one embodiment of the present disclosure, the audio determination module 603, when used to determine the target audio from multiple candidate audios based on the timbre similarity between the audio to be processed and each candidate audio, is used to:

[0198] The candidate audio with the greatest timbre similarity to the audio to be processed among the multiple candidate audios is determined as the target audio.

[0199] In one embodiment of the present disclosure, the audio processing device further includes an audio group determination module, which is used to determine a candidate audio group from a plurality of audio groups; the audio group determination module includes:

[0200] The audio acquisition submodule is used to acquire multiple sample audios;

[0201] A feature acquisition submodule is used to obtain the timbre features of each sample audio;

[0202] A similarity determination submodule, used for determining the timbre similarity between every two sample audios based on the timbre features of the multiple sample audios;

[0203] The similarity determination submodule is further used to take a set number of sample audios as an audio group and determine the sum of the timbre similarities between the set number of sample audios included in each audio group;

[0204] The audio group determination submodule is used to select the audio group with the smallest sum of timbre similarities as a candidate audio group.

[0205] In one embodiment of the present disclosure, the audio group determination module further includes:

[0206] The processing submodule is used to pre-process the multiple sample audios, where the pre-processing includes at least one of noise reduction processing, de-essing processing and volume normalization processing.

[0207] In one embodiment of the present disclosure, the audio group determination module further includes:

[0208] The parameter acquisition submodule is used to acquire dynamic EQ tuning parameters and static EQ tuning parameters of the audio included in the candidate audio group as preset tuning parameters.

[0209] In one embodiment of the present disclosure, the audio is dry audio, and the timbre features are composed of features of a target dimension in Mel-frequency cepstral coefficient (MFCC) features.

[0210] It should be noted that although several modules / units of the audio processing device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to an embodiment of the present disclosure, the features and functions of two or more modules / units described above may be embodied in one module / unit. Conversely, the features and functions of one module / unit described above may be further divided to be embodied by multiple modules / units.

[0211] The embodiment of the present disclosure also provides a computer-readable storage medium. Figure 7 is a schematic diagram of a computer-readable storage medium according to an exemplary embodiment of the present disclosure, such as Figure 7 As shown, the storage medium stores a computer program 701, and when the computer program 701 is executed by a processor, the audio processing method provided by any embodiment of the present disclosure can be executed.

[0212] The present disclosure also provides a computing device, which may include a memory and a processor, wherein the memory is used to store computer instructions that can be run on the processor, and the processor is used to implement the audio processing method provided by any embodiment of the present disclosure when executing the computer instructions. Figure 8 , Figure 8 It is a structural diagram of a computing device according to an exemplary embodiment of the present disclosure. The computing device 800 may include but is not limited to: a processor 810, a memory 820, and a bus 830 connecting different system components (including the memory 820 and the processor 810).

[0213] The memory 820 stores computer instructions, which can be executed by the processor 810, so that the processor 810 can execute the audio processing method provided by any embodiment of the present disclosure. The memory 820 may include a random access memory unit RAM 821, a cache memory unit 822 and / or a read-only memory unit ROM 823. The memory 820 may also include: a program tool 825 having a set of program modules 824, and the program modules 824 include but are not limited to: an operating system, one or more application programs, other program modules and program data, and one or more combinations of these program modules may include the implementation of a network environment.

[0214] The bus 830 may include, for example, a data bus, an address bus, and a control bus. The computing device 800 may also communicate with an external device 850 through an I / O interface 840. The external device 850 may be, for example, a keyboard, a Bluetooth device, etc. The computing device 800 may also communicate with one or more networks through a network adapter 860. For example, the network may be a local area network, a wide area network, a public network, etc. Figure 8 As shown, the network adapter 860 can also communicate with other modules of the computing device 800 via the bus 830 .

[0215] The embodiments of the present disclosure further provide a computer program product, which includes a computer program. When the program is executed by the processor 810 of the computing device 800, the audio processing method provided by any embodiment of the present disclosure can be implemented.

[0216] In addition, although the operations of the disclosed method are described in a specific order in the drawings, this does not require or imply that the operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0217] Although the spirit and principle of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the disclosed specific embodiments, and the division of various aspects does not mean that the features in these aspects cannot be combined to benefit, and such division is only for the convenience of expression. The present disclosure is intended to cover various modifications and equivalent arrangements included in the spirit and scope of the attached claims.

Claims

1. An audio processing method, It is characterized in that The method comprises: Get the audio to be processed; Based on the timbre feature of each candidate audio in the candidate audio group and the timbre feature of the audio to be processed, determining the timbre similarity between each candidate audio in the candidate audio group and the audio to be processed, wherein the candidate audio group is an audio group determined from multiple audio groups, the sum of the timbre similarities corresponding to the candidate audio group is the smallest among the sums of the timbre similarities corresponding to the various audio groups, the sum of the timbre similarities is the sum of the similarities between the audios included in the audio group, and each of the candidate audios corresponds to a preset tuning parameter; Determining a target audio from a plurality of candidate audios based on the timbre similarity between the audio to be processed and each candidate audio; Based on the preset tuning parameters of the target audio, the audio to be processed is tuned; For each candidate audio and any audio among the audios to be processed, the process of determining the timbre characteristics of the audio includes: determining a vocal portion of the audio; Extracting a target number of audio frames from the vocal portion of the audio; Acquire the timbre features of each audio frame to obtain the target number of timbre features; Performing an averaging process based on the target number of timbre features to obtain the timbre features of the audio; The timbre features are composed of shallow features of the target dimension in the Mel-frequency cepstral coefficient (MFCC) features.

2. The method according to claim 1, It is characterized in that The step of extracting a target number of audio frames from the vocal part of the audio comprises: Determining a sampling frequency based on the duration of the vocal portion of the audio; According to the sampling frequency, audio frames are extracted from the vocal part of the audio to obtain a target number of audio frames.

3. The method according to claim 2, It is characterized in that The determining, based on the timbre feature of each candidate audio in the candidate audio group and the timbre feature of the audio to be processed, the timbre similarity between each candidate audio in the candidate audio group and the audio to be processed comprises: For any candidate audio, a cosine distance between the timbre feature of the candidate audio and the timbre feature of the audio to be processed is determined, and the determined cosine distance is used as the timbre similarity between the candidate audio and the audio to be processed.

4. The method according to claim 1, It is characterized in that The step of determining the target audio from a plurality of candidate audios based on the timbre similarity between the audio to be processed and each candidate audio comprises: The candidate audio with the greatest timbre similarity to the audio to be processed among the multiple candidate audios is determined as the target audio.

5. The method according to claim 1, It is characterized in that The process of determining the candidate audio group includes: Get multiple audio samples; Obtain the timbre characteristics of each sample audio; Determining the timbre similarity between every two sample audios based on the timbre features of the multiple sample audios; Taking a set number of sample audios as an audio group, determining the sum of timbre similarities between two of the set number of sample audios included in each audio group; The audio group with the smallest sum of timbre similarities is taken as the candidate audio group.

6. The method according to claim 5, It is characterized in that After obtaining the plurality of sample audios, the method further comprises: The plurality of sample audios are preprocessed, wherein the preprocessing includes at least one of noise reduction processing, de-essing processing, and volume normalization processing.

7. The method according to claim 5, It is characterized in that After taking the audio group with the smallest sum of timbre similarities as the candidate audio group, the method further comprises: Dynamic EQ tuning parameters and static EQ tuning parameters of the audio included in the candidate audio group are obtained as preset tuning parameters.

8. The method according to any one of claims 1 to 7, It is characterized in that The audio is dry audio.

9. An audio processing device, It is characterized in that The device comprises: An acquisition module, used to acquire audio to be processed; A similarity determination module, for determining the timbre similarity between each candidate audio in the candidate audio group and the audio to be processed based on the timbre feature of each candidate audio in the candidate audio group and the timbre feature of the audio to be processed, wherein the candidate audio group is an audio group determined from multiple audio groups, the sum of the timbre similarities corresponding to the candidate audio group is the smallest among the sums of the timbre similarities corresponding to the various audio groups, the sum of the timbre similarities is the sum of the similarities between the audios included in the audio group, and each of the candidate audios corresponds to a preset tuning parameter; An audio determination module, configured to determine a target audio from a plurality of candidate audios based on the timbre similarity between the audio to be processed and each candidate audio; A processing module, configured to perform tuning processing on the audio to be processed based on preset tuning parameters of the target audio; The similarity determination module includes an acquisition submodule, and for each candidate audio and any audio in the audio to be processed, the acquisition submodule is used to acquire the timbre characteristics of the audio; The acquisition submodule includes: A determination unit, which determines a vocal part of the audio; An extraction unit extracts a target number of audio frames from the vocal part of the audio; The determining unit is further used to obtain the timbre features of each audio frame to obtain a target number of timbre features; The determining unit is further configured to perform an averaging process based on the target number of timbre features to obtain the timbre features of the audio; The timbre features are composed of shallow features of the target dimension in the Mel-frequency cepstral coefficient (MFCC) features.

10. The device according to claim 9, It is characterized in that The extraction unit, when used to extract a target number of audio frames from the vocal part of the audio, is used to: Determining a sampling frequency based on the duration of the vocal portion of the audio; According to the sampling frequency, audio frames are extracted from the vocal part of the audio to obtain a target number of audio frames.

11. The device according to claim 10, It is characterized in that The similarity determination module includes a determination submodule, and for any candidate audio, the determination submodule is used to determine the timbre similarity between the candidate audio and the audio to be processed based on the timbre characteristics of the candidate audio and the timbre characteristics of the audio to be processed; Wherein, the determination submodule is specifically used for: For any candidate audio, a cosine distance between the timbre feature of the candidate audio and the timbre feature of the audio to be processed is determined, and the determined cosine distance is used as the timbre similarity between the candidate audio and the audio to be processed.

12. The device according to claim 9, It is characterized in that The audio determination module, when used to determine the target audio from a plurality of candidate audios based on the timbre similarity between the audio to be processed and each candidate audio, is used to: The candidate audio with the greatest timbre similarity to the audio to be processed among the multiple candidate audios is determined as the target audio.

13. The device according to claim 9, It is characterized in that The apparatus further comprises an audio group determination module, configured to determine a candidate audio group from a plurality of audio groups; The audio group determination module includes: The audio acquisition submodule is used to acquire multiple sample audios; A feature acquisition submodule is used to obtain the timbre features of each sample audio; A similarity determination submodule, used to determine the timbre similarity between every two sample audios based on the timbre features of the multiple sample audios; The similarity determination submodule is further used to take a set number of sample audios as an audio group and determine the sum of the timbre similarities between the set number of sample audios included in each audio group; The audio group determination submodule is used to select the audio group with the smallest sum of timbre similarities as the candidate audio group.

14. The device according to claim 13, It is characterized in that The audio group determination module further includes: The processing submodule is used to pre-process the multiple sample audios, wherein the pre-processing includes at least one of noise reduction processing, de-essing processing and volume normalization processing.

15. The device according to claim 13, It is characterized in that The audio group determination module further includes: The parameter acquisition submodule is used to acquire the dynamic EQ tuning parameters and the static EQ tuning parameters of the audio included in the candidate audio group as preset tuning parameters.

16. The device according to any one of claims 9 to 15, It is characterized in that The audio is dry audio.

17. A computing device, It is characterized in that The computing device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the operations performed by the audio processing method according to any one of claims 1 to 8 when executing the program.

18. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a program, and the program is used by a processor to execute the operations performed by the audio processing method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Diversity-based sample screening method, system and device and medium

    CN112308143A

  • Intelligent tone tuning method and device based on timbre, medium and computing equipment

    CN113870873A

  • Audio signal dynamic equalization processing control

    US20120063614A1