Audio processing method, computer device and computer program product

By filtering and transposing the dry audio of a song, and combining it with the accompaniment audio to generate choral sound effects, the problems of long time consumption and high cost in traditional technology have been solved, and the acquisition of choral accompaniment audio has been achieved efficiently.

CN115171632BActive Publication Date: 2025-10-28TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210675793.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-15
Publication Date
2025-10-28
Estimated Expiration
2042-06-15

AI Technical Summary

Technical Problem

Traditional techniques require a dedicated choir and recording process to obtain the choral accompaniment audio for a song, which is time-consuming and costly, resulting in low efficiency in obtaining accompaniment audio.

Method used

By acquiring the set of dry audio files corresponding to the target song, filtering out the dry audio files that meet the preset conditions, and performing pitch shifting using chord progressions that match the song's key, the files are then blended with the accompaniment audio to generate an accompaniment audio with chorus effects.

Benefits of technology

It can quickly generate accompaniment audio with highly realistic choral sound effects without the need for a specially configured choir, thus improving the efficiency of obtaining song accompaniment audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115171632B_ABST
    Figure CN115171632B_ABST
Patent Text Reader

Abstract

This application relates to an audio processing method, computer device, and computer program product. The method includes: acquiring a set of dry audio files corresponding to a target song, and selecting at least one target dry audio file from the dry audio file that meets preset conditions; the dry audio file is obtained by acquiring sound signals generated by at least one target object singing the target song; acquiring the key of the target song, and using chord progressions matching the key of the song to perform pitch shifting on each of the target dry audio files, obtaining pitch-shifted dry audio files; performing a fusion process on each of the pitch-shifted dry audio files to obtain target harmonic audio files; and mixing the target harmonic audio files with the accompaniment audio files corresponding to the target song to output target accompaniment audio files with chorus effects. This method can improve the efficiency of audio acquisition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to an audio processing method, computer device, and computer program product. Background Technology

[0002] With the development of internet technology, more and more people are singing songs through the song singing function provided by their devices. When users sing songs using this function, the device can often play the accompaniment audio, such as instrumental or choral accompaniment, to help users sing the songs better.

[0003] However, traditional techniques for obtaining the choral accompaniment audio for a song often require a specially configured choir and the recording of the choir singing the target song's accompaniment, which is time-consuming and costly, making it difficult to improve the efficiency of obtaining the song's accompaniment audio. Summary of the Invention

[0004] Therefore, it is necessary to provide an audio processing method, computer device, and computer program product that can improve the efficiency of acquiring accompaniment audio, in response to the above-mentioned technical problems.

[0005] In a first aspect, this application provides an audio processing method, the method comprising:

[0006] Obtain the dry audio set corresponding to the target song, and filter out at least one target dry audio that meets preset conditions from the dry audio set; the dry audio set is obtained by collecting the sound signal generated by at least one target object singing the target song;

[0007] Obtain the key of the target song, and use the chord progression that matches the key of the song to perform pitch shifting on each of the target dry audio files to obtain the pitch-shifted dry audio files.

[0008] The modified dry audio files are fused together to obtain the target harmonic audio.

[0009] The target harmonic audio is mixed with the accompaniment audio corresponding to the target song to output the target accompaniment audio with chorus sound effects.

[0010] In one embodiment, obtaining the key of the target song and using chord progressions that match the key of the song to perform pitch shifting on each of the target dry audio files to obtain pitch-shifted dry audio files includes:

[0011] Obtain the key offset that matches the key of the song; the key offset is used to characterize the chord progression that matches the key of the song;

[0012] Adjust the original pitch corresponding to the key of the song according to the key offset to determine the pitch adjustment target for the key shifting process;

[0013] Using the aforementioned pitch adjustment targets, the dry audio of each target is subjected to pitch shifting processing to obtain the pitch-shifted dry audio.

[0014] In one embodiment, before the step of performing pitch shifting processing on each of the target dry audio samples to obtain pitch-shifted dry audio samples, the method further includes:

[0015] The loudness of each target dry audio audio is adjusted to obtain the target dry audio audio with adjusted loudness; the loudness of each target dry audio audio with adjusted loudness is within a preset loudness range.

[0016] In one embodiment, the fusion processing of each of the pitch-shifted dry audio samples to obtain the target acoustic audio includes:

[0017] Obtain the harmonic weights corresponding to each of the pitch-shifted dry audio samples;

[0018] According to the harmonic weights corresponding to each of the pitch-shifted dry audio signals, the audio signals corresponding to each of the pitch-shifted dry audio signals are weighted and summed to output the target harmonic audio signal.

[0019] In one embodiment, after the step of fusing the pitch-shifted dry audio samples to obtain the target and acoustic audio samples, the method further includes:

[0020] The audio signal of the target and acoustic audio is processed with sound effects adjustment, and the sound effects adjusted target and acoustic audio is output.

[0021] The sound effect adjustment processing includes at least one of audio equalization processing, dynamic range control processing, and reverb effect addition processing.

[0022] In one embodiment, the step of mixing the target harmonic audio with the accompaniment audio corresponding to the target song to output the target accompaniment audio with chorus effects includes:

[0023] The loudness of the target and acoustic audio frequencies is adjusted so that the loudness of the adjusted target and acoustic audio frequencies is less than the loudness of the accompaniment audio frequencies;

[0024] The adjusted target harmony audio is mixed with the accompaniment audio to output the target accompaniment audio.

[0025] The loudness of the target accompaniment audio is adjusted so that the loudness of the adjusted target accompaniment audio is equal to a preset loudness threshold.

[0026] In one embodiment, obtaining the key of the target song includes:

[0027] The total cumulative duration of each note in the target song is identified; the total cumulative duration of each note is used to characterize the pitch distribution information of the target song.

[0028] Based on the pitch distribution information and the tonality weights corresponding to each note, calculate the Pearson coefficient corresponding to at least one candidate song tonality.

[0029] The key of the candidate song with the highest Pearson coefficient is taken as the key of the target song.

[0030] In one embodiment, the step of selecting at least one target dry audio file that meets preset conditions from the dry audio file set includes:

[0031] Based on the sound quality of each dry audio file in the dry audio file set, a first audio file set is selected from the dry audio file set; the sound quality of each dry audio file in the first audio file set meets the preset sound quality conditions.

[0032] Based on the pitch of each dry audio file in the first audio set, a second audio set is selected from the first audio set; the pitch of each dry audio file in the second audio set meets the preset pitch conditions.

[0033] A preset number of dry audio files are determined from the second audio set as the at least one target dry audio file.

[0034] In one embodiment, the step of filtering out a first audio set from the dry audio set based on the sound quality of each dry audio file in the dry audio set includes:

[0035] Remove dry audio with abnormal sound quality from the dry audio set to obtain the removed dry audio; the abnormal sound quality includes at least one of the following: noise, popping sound, backslide of accompaniment, audio length less than a preset length threshold, and audio energy less than a preset energy threshold.

[0036] The removed dry audio files are input into the sound quality scoring model to obtain the sound quality scores corresponding to each of the removed dry audio files.

[0037] After removing the dry audio files whose sound quality scores are greater than or equal to a preset score threshold, add them to the first audio set.

[0038] In one embodiment, the step of filtering out a second audio set from the first audio set based on the pitch of each dry audio file in the first audio set includes:

[0039] Obtain the rhythm sequence of each dry audio in the first audio set, and determine the rhythm difference between the rhythm sequence of each dry audio in the first audio set and the reference rhythm sequence corresponding to the target song;

[0040] Obtain the fundamental frequency of each dry audio file in the first audio set, and determine the fundamental frequency difference between the fundamental frequency of each dry audio file in the first audio set and the reference fundamental frequency corresponding to the target song;

[0041] At least one target pitch audio is determined in the first audio set, and each target pitch audio is added to the second audio set; the rhythm difference corresponding to the target pitch audio is less than a preset rhythm difference level, and the fundamental frequency difference corresponding to the target pitch audio is less than a preset fundamental frequency difference level.

[0042] Secondly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program that, when executed by the processor, implements the steps of the method described above.

[0043] Thirdly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the above-described method.

[0044] The aforementioned audio processing method, apparatus, computer equipment, storage medium, and computer program product acquire a set of dry audio files corresponding to a target song, and then select at least one target dry audio file that meets preset conditions from the dry audio file set. The dry audio file set is obtained by collecting sound signals generated by at least one target object singing the target song. Next, the tonality of the target song is acquired, and each target dry audio file is pitch-shifted using chord progressions that match the tonality, resulting in pitch-shifted dry audio files. Finally, the pitch-shifted dry audio files are fused to obtain the target harmony. Finally, by mixing the target harmony audio with the accompaniment audio corresponding to the target song, the target accompaniment audio with chorus sound effects is output. In this way, by using the dry audio generated by at least one target object singing the target song as the sound source for synthesizing the chorus audio of the target song, and by adaptively transposing each dry audio using the key of the song, it is possible to quickly achieve accompaniment audio with a high degree of realism in chorus sound effects, without having to specially configure a chorus for the target song and record the corresponding chorus audio, which greatly improves the efficiency of obtaining song accompaniment audio. Attached Figure Description

[0045] Figure 1 This is a diagram illustrating the application environment of an audio processing method in one embodiment.

[0046] Figure 2 This is a flowchart illustrating an audio processing method in one embodiment;

[0047] Figure 3 This is a flowchart illustrating another audio processing method in one embodiment;

[0048] Figure 4 This is a flowchart illustrating an audio processing method in another embodiment;

[0049] Figure 5 This is a schematic diagram of the architecture of an audio processing method in one embodiment;

[0050] Figure 6 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0051] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0052] The audio processing method provided in this application embodiment can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or located in the cloud or on another network server. When a user sings a target song using terminal 102 with a karaoke application (e.g., a karaoke software) installed, terminal 102 can collect the sound signal generated by the user singing the target song, thus obtaining dry audio. Then, terminal 102 can upload the collected dry audio to server 104, which then uses this dry audio to build a dry audio library (i.e., a dry audio set) for each song. Server 104 selects at least one target dry audio file that meets preset conditions from the dry audio set; Server 104 obtains the key of the target song and uses chord progressions matching the key to perform pitch shifting on each target dry audio file to obtain pitch-shifted dry audio; Server 104 performs fusion processing on each pitch-shifted dry audio file to obtain target harmony audio; Server 104 mixes the target harmony audio with the accompaniment audio corresponding to the target song to output target accompaniment audio with chorus effects. In specific implementation, terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc. Server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0053] In one embodiment, such as Figure 2 As shown, an audio processing method is provided. Taking the application of this method to a server as an example, it includes the following steps:

[0054] Step S202: Obtain the set of dry audio files corresponding to the target song, and select at least one target dry audio file that meets the preset conditions from the set of dry audio files.

[0055] Dry audio can refer to pure human voice audio without music.

[0056] In practical applications, the dry audio set is obtained by collecting the sound signal generated by at least one target object singing the target song.

[0057] For example, when a user sings a target song using a terminal with a karaoke application (such as a karaoke software) installed, the terminal can collect the sound signal generated by the user singing the target song, thus obtaining dry audio. The terminal can then upload the collected dry audio to a server, which can then use the server to build a dry audio library (i.e., a dry audio collection) for each song.

[0058] In practical applications, the dry audio set corresponding to the target song can include dry audio generated from the same user singing the target song multiple times; it can also include dry audio generated from different users singing the target song.

[0059] In practice, the server can respond to a request to generate a chorus accompaniment for a target song by obtaining a set of dry audio files corresponding to that target song. Then, the server selects at least one target dry audio file from the dry audio file that meets preset conditions. In practical applications, the target dry audio file can be named high-quality dry audio file.

[0060] Specifically, the server can select N target dry audio files from the dry audio set that meet preset requirements for sound quality, rhythm, and pitch, using these as the original sound sources for generating choir accompaniment audio. In practical applications, N is a positive integer greater than 0. For example, N can be equal to 10.

[0061] Step S204: Obtain the key of the target song and use the chord progression that matches the key of the song to perform pitch shifting on each target dry audio file to obtain the pitch-shifted dry audio file.

[0062] In practice, after the server filters out the target dry audio, it can obtain the key of the target song. Then, the server uses the chord progression that matches the key of the song to perform pitch shifting on each target dry audio, thus obtaining the pitch-shifted dry audio.

[0063] Specifically, the server can input the audio of the target song into the tonality detector to determine the key of the song. Then, the server determines the chord tones corresponding to the key of the song, and determines the harmonic positions based on the chord tones. Using the harmonic tones indicated by the harmonic positions, the server performs pitch shifting on each target dry audio file to obtain the pitch-shifted dry audio.

[0064] In another embodiment, before the step of performing pitch shifting processing on each target dry audio to obtain pitch-shifted dry audio, the method further includes: adjusting the loudness of each target dry audio to obtain loudness-adjusted target dry audio; the loudness corresponding to each loudness-adjusted target dry audio is within a preset loudness range.

[0065] In practice, before the server performs pitch shifting on each target dry audio and obtains the pitch-shifted dry audio, the server can also use volume equalization to adjust the loudness of each target dry audio so that the loudness of each target dry audio after loudness adjustment is within a preset loudness range.

[0066] Specifically, the server uses Dynamic Range Control (DRC) to effectively compress and control the volume of each target dry audio signal, thereby obtaining the target dry audio signal with adjusted loudness. The preset loudness range can be set to be controlled at around -14dB.

[0067] The technical solution of this embodiment adjusts the loudness of each target dry audio by using volume equalization, so that the effective dry audio after screening is controlled within a similar loudness range, preventing the sound volume from being disharmonious after synthesis due to individual sounds being too loud or too soft.

[0068] Step S206: Perform fusion processing on each pitch-shifted dry audio to obtain the target harmonic audio.

[0069] In practice, after obtaining the dry audio files after each pitch shift, the server performs a fusion process on these dry audio files to obtain the target harmonic audio. Specifically, the server can obtain the weights corresponding to each harmonic pitch and use these weights to adjust and fuse the loudness of each dry audio file after pitch shift to obtain the target harmonic audio (i.e., the choir's wet audio).

[0070] In another embodiment, after the step of fusing the pitch-shifted dry audio signals to obtain the target harmony audio, the method further includes: performing sound effect adjustment processing on the audio signal of the target harmony audio, and outputting the sound effect adjusted target harmony audio.

[0071] The sound effect adjustment processing includes at least one of the following: audio equalization processing, dynamic range control processing, and reverb effect addition processing.

[0072] In practice, after the server performs fusion processing on the dry audio after each pitch shift to obtain the target harmony audio, the server can also perform post-processing on the target harmony audio (i.e., wet audio). Specifically, the server can perform audio equalization (EQ), dynamic range control (DRC), and add preset reverb effects on the audio signal of the target harmony audio, and output the target harmony audio after sound effect adjustment to broaden the sound field spatial sense of the target harmony audio.

[0073] Step S208: Mix the target harmony audio with the accompaniment audio corresponding to the target song, and output the target accompaniment audio with chorus sound effects.

[0074] In practice, the server can mix the target harmonic audio with the corresponding accompaniment audio of the target song to output a target accompaniment audio with chorus effects. Specifically, the server can adjust the energy ratio between the target harmonic audio and the corresponding accompaniment audio of the target song and perform the mixing process to obtain a target accompaniment audio with chorus effects.

[0075] In the above audio processing method, a set of dry audio samples corresponding to the target song is obtained, and at least one target dry audio sample that meets preset conditions is selected from the dry audio sample set. The dry audio sample set is obtained by collecting sound signals generated by at least one target object singing the target song. Then, the key of the target song is obtained, and each target dry audio sample is pitch-shifted using chord progressions that match the key of the song to obtain pitch-shifted dry audio. The pitch-shifted dry audio samples are then fused to obtain the target harmony audio. Finally, the target harmony audio is blended with the accompaniment audio corresponding to the target song. The audio is mixed to output a target accompaniment audio with choral effects. In this way, by using the dry audio generated by at least one target object singing the target song as the sound source for synthesizing the chorus audio of the target song, and by adaptively transposing each dry audio using the key of the song, it is possible to quickly generate accompaniment audio with a high degree of realism in choral effects. This achieves the effect of simulating a chorus singing the target song and the corresponding audio track, without the need to specially configure a chorus for the target song and record the chorus audio corresponding to the target song, which greatly improves the efficiency of obtaining the song accompaniment audio.

[0076] In another embodiment, the key of the target song is obtained, and the target dry audio is pitch-shifted using chord progressions that match the key of the song to obtain pitch-shifted dry audio. This includes: obtaining a key offset that matches the key of the song; adjusting the original pitch corresponding to the key of the song according to the key offset to determine the pitch adjustment target for the pitch shifting process; and using the pitch adjustment target to perform pitch shifting on each target dry audio to obtain pitch-shifted dry audio.

[0077] The tonality offset is used to characterize the chord progression that matches the key of the song.

[0078] In the specific implementation, the server obtains the key of the target song and performs pitch shifting on each target dry audio file using chord progressions that match the key. During this process, the server can acquire a key shift value representing the chord progressions that match the song's key. Then, the server adjusts the original pitch corresponding to the song's key according to the key shift value, determining the pitch adjustment target for the pitch shifting process. Finally, the server uses the pitch adjustment target to perform pitch shifting on each target dry audio file, obtaining the pitch-shifted dry audio file.

[0079] Specifically, after determining the key of a song, the server uses the chords within the key to determine the harmonic positions. For simplicity and efficiency, third harmony can be used directly. Taking a song in the key of "C major" as an example, the server can use the third / fifth intervals within the key as chord tones. Taking the fifth interval as an example, the key offsets are determined to be +4, +7, and +11 respectively. The server adjusts the original key corresponding to the song's key according to the key offsets to determine the target tone adjustment (i.e., harmonic tone) for the key-shifting process. For ease of understanding by those skilled in the art, the harmonic tone can be represented as:

[0080]

[0081] Where, N base This refers to the original key corresponding to the song's key. In practical applications, to maintain the stability of the harmonic pitch as much as possible, the server can select a key triad (i.e.,...). () as the final harmonic tone.

[0082] The server determines the pitch adjustment target for pitch shifting. Using this target, the server performs pitch shifting on each target's dry audio, obtaining the pitch-shifted dry audio. Specifically, the server can calculate the constant pitch shift coefficient for the entire song according to the pitch adjustment target, resulting in the pitch-shifted harmonious dry audio. The server can use preset signal processing methods (e.g., TSM, Worldvocoder) or pre-trained neural networks (e.g., WaveNet, LPCNet models) to perform pitch shifting on each target's dry audio, obtaining the pitch-shifted dry audio. The pitch-shifted dry audio can be represented as:

[0083]

[0084] Where, ζ k (·) indicates the number of semitones in the tone shift process when the pitch shift is k, where k = 4, 7, which is the number of semitones of the tone shift; This represents the transposed dry voice after k semitones of the u-th dry voice, and i represents the audio sample index.

[0085] The technical solution of this embodiment obtains a tone offset that matches the tone of the song; adjusts the original tone corresponding to the tone of the song according to the tone offset to determine the tone adjustment target for tone shifting; and uses the tone adjustment target to perform tone shifting on each target dry audio to obtain tone-shifted dry audio. Thus, the dry audio of the target song can be adaptively tone-shifted based on the tone of the target song, so that the synthesized sound effect synthesized based on the tone-shifted dry audio can match the target song well.

[0086] In another embodiment, the target harmonic audio is obtained by fusing the dry audio signals after each pitch shifting, including: obtaining the harmonic weights corresponding to each dry audio signal after each pitch shifting; and performing a weighted summation of the audio signals corresponding to each dry audio signal after each pitch shifting according to the harmonic weights, and outputting the target harmonic audio.

[0087] In the specific implementation, during the process of the server fusing the dry audio after each pitch shift to obtain the target harmonic audio, the server can obtain the harmonic weights corresponding to each dry audio after each pitch shift; according to the harmonic weights corresponding to each dry audio after each pitch shift, the server performs a weighted summation of the audio signals corresponding to each dry audio after each pitch shift, and outputs the target harmonic audio.

[0088] In practical applications, the target and the audio frequency can be represented as:

[0089]

[0090] Where α0, α1, α2, and α3 represent the corresponding harmonic weights: s syn (i) represents the final synthesized wet sound signal, i.e., the target sound and audio frequency; This represents the transposed dry voice after k semitones of the u-th dry voice, and i represents the audio sample index.

[0091] In practical applications, α0 = 0.1 can be set as the backing tone for the main track, and α1 = 0.3 and α2 = 0.3 can be set as the relative weights on the harmonic intervals.

[0092] The technical solution of this embodiment obtains the harmonic weights corresponding to each pitch-shifted dry audio; according to the harmonic weights corresponding to each pitch-shifted dry audio, the audio signals corresponding to each pitch-shifted dry audio are weighted and summed, so that each pitch-shifted dry audio can be reasonably and effectively fused to output a target harmonic audio with good listening experience.

[0093] In another embodiment, mixing the target harmony audio with the accompaniment audio corresponding to the target song to output a target accompaniment audio with chorus effects includes: adjusting the loudness of the target harmony audio so that the loudness of the adjusted target harmony audio is less than the loudness of the accompaniment audio; mixing the adjusted target harmony audio with the accompaniment audio to output a target accompaniment audio; and adjusting the loudness of the target accompaniment audio so that the loudness of the adjusted target accompaniment audio is equal to a preset loudness threshold.

[0094] In practice, during the mixing process of the target harmony audio and the corresponding accompaniment audio of the target song, and the output of the target accompaniment audio with chorus effects, the server can adjust the energy ratio between the target harmony audio and the corresponding accompaniment audio to prevent the dry vocals from overpowering the accompaniment. Specifically, the server can adjust the loudness of the target harmony audio so that the loudness of the adjusted target harmony audio is less than the loudness of the accompaniment audio. For example, the loudness of the adjusted target harmony audio can be 3dB lower than the loudness of the accompaniment audio.

[0095] Then, the server mixes the adjusted target harmonic audio with the accompaniment audio, outputs the target accompaniment audio, and adjusts the loudness of the output target accompaniment audio to the same loudness as the original accompaniment. That is, the loudness of the output target accompaniment audio is adjusted to be equal to a preset loudness threshold, which can be -14dB or the loudness of the accompaniment audio corresponding to the target song.

[0096] The technical solution of this embodiment adjusts the loudness of the target harmony audio to make the loudness of the adjusted target harmony audio less than the loudness of the accompaniment audio; then mixes the adjusted target harmony audio with the accompaniment audio to output the target accompaniment audio; and then adjusts the loudness of the target accompaniment audio to make the loudness of the adjusted target accompaniment audio equal to a preset loudness threshold. In this way, by adjusting the energy ratio between the target harmony audio and the accompaniment audio corresponding to the target song, it is possible to prevent the energy of the target harmony audio from being too large and overpowering the original accompaniment audio in the target song. This allows for the rapid acquisition of accompaniment audio with a high degree of realism in choral sound effects, greatly improving the efficiency of acquiring song accompaniment audio.

[0097] In another embodiment, obtaining the key of the target song includes: identifying the total cumulative duration of each note in the target song; the total cumulative duration of each note is used to characterize the pitch distribution information of the target song; calculating the Pearson coefficient corresponding to at least one candidate song key according to the pitch distribution information and the key weight corresponding to each note; and taking the candidate song key with the highest Pearson coefficient as the key of the target song.

[0098] In the specific implementation, during the process of obtaining the key of the target song, the server can identify the total cumulative duration of each note in the target song that represents the pitch distribution information of the target song; then, according to the pitch distribution information and the key weight corresponding to each note, the server calculates the Pearson coefficient corresponding to at least one candidate song key; finally, the server takes the candidate song key with the highest Pearson coefficient as the key of the target song.

[0099] For example, the process of a server obtaining the key of a target song includes the following steps:

[0100] 1. Fundamental frequency extraction:

[0101] The server can extract the fundamental frequency from the original audio of the target song and convert it into the corresponding pitch. The conversion formula can be:

[0102]

[0103] Where f represents the detected fundamental frequency, and N represents the pitch corresponding to the fundamental frequency.

[0104] 2. Calculate the total duration of different notes based on the converted pitch information (in practical applications, the octave distinction can be ignored).

[0105] 3. Select the tonal weight corresponding to each note.

[0106] The server can select the tone weights in Simple mode, and the weight values ​​are shown in Table 1 below:

[0107] tone C C# D D# E F F# G G# A A# B c c# d d# e f f# g g# a a# b Weight 2 0 1 0 1 1 0 2 0 1 0 1 2 0 1 1 0 1 0 2 1 0 0.5 0.5

[0108] Table 1 shows the tonal weights corresponding to each note.

[0109] Of course, if the server stores the sheet music information corresponding to the target song (e.g., in MIDI information), the server can read the sheet music information corresponding to the target song and determine the key of the target song based on the sheet music information.

[0110] 4. Adjust the total duration of the notes. Calculate the Pearson coefficient (PCC) for the first note in the tonic list, resulting in 24 values.

[0111] 5. Select the key corresponding to the maximum value of the Pearson coefficient as the final key of the song. The Pearson coefficient parameters are defined as follows:

[0112]

[0113] The final estimated key of the song is as follows:

[0114]

[0115] Among them, key k This represents the weights of the 24 tonal values, and N represents the current tone distribution of the song.

[0116] The technical solution of this embodiment identifies the total cumulative duration of each note in the target song that represents the pitch distribution information of the target song; and calculates the Pearson coefficient corresponding to at least one candidate song key according to the pitch distribution information and the key weight corresponding to each note; and takes the candidate song key with the highest Pearson coefficient as the key of the target song, thereby enabling the key of the target song to be identified quickly and accurately.

[0117] In another embodiment, such as Figure 3 As shown, an audio processing method is provided, which is applied to... Figure 1 Taking the server in the example, the following steps are included:

[0118] Step S302: Obtain the dry audio set corresponding to the target song, and filter out at least one target dry audio that meets the preset conditions from the dry audio set; the dry audio set is obtained by collecting the sound signal generated by at least one target object singing the target song.

[0119] Step S304: Adjust the loudness of each target dry audio audio to obtain the target dry audio audio with loudness adjustment; the loudness of each target dry audio audio with loudness adjustment is within the preset loudness range.

[0120] Step S306: Obtain the key offset that matches the key of the target song; the key offset is used to represent the chord progression that matches the key of the song.

[0121] Step S308: Adjust the original pitch corresponding to the key of the song according to the key offset to determine the pitch adjustment target for the key shifting process.

[0122] Step S310: Using pitch adjustment targets, the dry audio of each target is pitch-shifted to obtain the pitch-shifted dry audio.

[0123] Step S312: According to the harmonic weights corresponding to each pitch-shifted dry audio, perform weighted summation on the audio signals corresponding to each pitch-shifted dry audio, and output the target harmonic audio.

[0124] Step S314: Perform sound effect adjustment processing on the audio signal of the target harmony audio and output the sound effect adjusted target harmony audio; wherein, the sound effect adjustment processing includes at least one of audio equalization processing, dynamic range control processing, and reverb effect addition processing.

[0125] Step S316: Adjust the loudness of the target harmony audio to make the loudness of the adjusted target harmony audio less than the loudness of the accompaniment audio.

[0126] Step S318: Mix the adjusted target harmony audio with the accompaniment audio and output the target accompaniment audio.

[0127] Step S320: Adjust the loudness of the target accompaniment audio so that the loudness of the adjusted target accompaniment audio is equal to the preset loudness threshold.

[0128] It should be noted that the specific limitations of the above steps can be found in the specific limitations of an audio processing method described above.

[0129] In another embodiment, selecting at least one target dry audio that meets preset conditions from the dry audio set includes: selecting a first audio set from the dry audio set based on the sound quality of each dry audio in the dry audio set; selecting a second audio set from the first audio set based on the pitch of each dry audio in the first audio set; and determining a preset number of dry audio in the second audio set as at least one target dry audio.

[0130] In the first audio set, the sound quality of each dry audio file meets the preset sound quality conditions.

[0131] In the second audio set, the pitch of each dry audio file meets the preset pitch conditions.

[0132] In the specific implementation, during the process of the server selecting at least one target dry audio that meets the preset conditions from the dry audio set, the server can perform sound quality filtering on each dry audio in the dry audio set: specifically, the server can select the dry audio that meets the preset sound quality conditions as the target dry audio based on the sound quality of each dry audio in the dry audio set.

[0133] For example, dry audio that meets the preset sound quality conditions can be audio that is free of noise, has no accompaniment, has an audio length that meets the preset length threshold, has an audio energy that meets the preset energy threshold, and has no popping sounds.

[0134] The server can perform pitch accuracy filtering on each dry audio file in the first audio set: Specifically, the server can filter out dry audio files in the first audio set whose pitch accuracy meets preset pitch conditions, and use them as the second audio set.

[0135] For example, dry audio that meets preset pitch conditions can be dry audio whose rhythm and / or pitch differ from the rhythm and / or pitch of the reference audio corresponding to the target song by less than a preset difference threshold.

[0136] After filtering out dry audio files that meet preset conditions in terms of sound quality and pitch from the dry audio set, the server can obtain a preset number of dry audio files as target dry audio files. In practical applications, the server can select 10 high-quality dry audio files, i.e., U=10 (U represents the number of original dry audio files), as the original sound source for the chorus sound effect.

[0137] The technical solution of this embodiment selects a first audio set whose sound quality meets preset sound quality conditions based on the sound quality of each dry audio in the dry audio set, so that the dry audio has good sound quality and ensures the audio quality after mixing multiple dry audios. At the same time, based on the pitch of each dry audio in the first audio set, a second audio set whose pitch meets preset pitch conditions is selected from the first audio set, so that the dry audio of the subsequent harmonious melody has good time alignment characteristics, ensuring that when mixing multiple dry audios, there will be no problem of noisy listening due to sound asynchrony.

[0138] In another embodiment, a first audio set is selected from the dry audio set based on the sound quality of each dry audio file in the dry audio set, including: removing dry audio files with sound quality abnormalities from the dry audio set to obtain the removed dry audio files; sound quality abnormalities include at least one of noise, pops, backsound, audio length less than a preset length threshold, and audio energy less than a preset energy threshold; inputting each removed dry audio file into a sound quality scoring model to obtain a sound quality score corresponding to each removed dry audio file; adding the removed dry audio files with sound quality scores greater than or equal to a preset score threshold to the first audio set.

[0139] In the specific implementation, during the process of filtering out the first audio set from the dry audio set based on the sound quality of each dry audio in the dry audio set, the server can use a sound quality detection tool to remove dry audio in the dry audio set that has abnormal sound quality such as noise, popping sounds, backtracking, audio length less than a preset length threshold, and audio energy less than a preset energy threshold, and obtain the removed dry audio.

[0140] The server can input each removed dry audio file into the sound quality scoring model (e.g., a sound quality assessment tool) to calculate the sound quality score corresponding to each removed dry audio file; then, it can extract the removed dry audio files with sound quality scores greater than or equal to a preset score threshold and add them to the first audio set.

[0141] For example, assuming that the sound quality score of the removed dry audio A is 90 and the sound quality score of the removed dry audio B is 10, and the preset score threshold is 60, then the server will add the removed dry audio A to the first audio set.

[0142] The technical solution of this embodiment removes at least one audio quality abnormality from the dry audio set, such as noise, popping, backtracking, audio length less than a preset length threshold, and audio energy less than a preset energy threshold. Each removed dry audio is then input into a sound quality scoring model to obtain a sound quality score for each removed dry audio. The removed dry audio with a sound quality score greater than or equal to a preset score threshold is added to the first audio set. This ensures that each dry audio in the first audio set meets the preset sound quality conditions and avoids at least one sound quality abnormality from the dry audio in the first audio set, such as noise, popping, backtracking, audio length less than a preset length threshold, and audio energy less than a preset energy threshold.

[0143] In another embodiment, selecting a second audio set from the first audio set based on the pitch accuracy of each dry audio in the first audio set includes: obtaining the rhythm sequence of each dry audio in the first audio set and determining the rhythm difference between the rhythm sequence of each dry audio in the first audio set and the reference rhythm sequence corresponding to the target song; obtaining the fundamental frequency of each dry audio in the first audio set and determining the fundamental frequency difference between the fundamental frequency of each dry audio in the first audio set and the reference fundamental frequency corresponding to the target song; determining at least one target pitch audio in the first audio set and adding each target pitch audio to the second audio set.

[0144] Among them, the rhythm difference corresponding to the target pitch audio is less than the preset rhythm difference level, and the fundamental frequency difference corresponding to the target pitch audio is less than the preset fundamental frequency difference level.

[0145] In practice, during the process of selecting a second audio set from the first audio set based on the pitch accuracy of each dry audio track in the first audio set, the server can obtain the rhythm sequence of each dry audio track in the first audio set and determine the rhythm difference between the rhythm sequence of each dry audio track in the first audio set and the reference rhythm sequence corresponding to the target song. Then, the server extracts the fundamental frequency (i.e., fundamental tone frequency) of each dry audio track in the first audio set and determines the fundamental frequency difference between the fundamental frequency of each dry audio track in the first audio set and the reference fundamental frequency corresponding to the target song. Finally, after determining the fundamental frequency difference and rhythm difference corresponding to each dry audio track in the first audio set, the server adds the dry audio tracks with rhythm differences less than a preset level and fundamental frequency differences less than a preset level of fundamental frequency difference as target pitch audio tracks to the second audio set.

[0146] For example, the server can obtain the rhythmic sequence (i.e., the dry note sequence) representing the rhythmic information of each dry audio file from the note or MIDI information of each dry audio file. Then, the server uses the same method to obtain the reference rhythmic sequence (i.e., the reference note sequence) corresponding to the target song, and determines the rhythmic difference between the rhythmic sequence of each dry audio file and the reference rhythmic sequence corresponding to the target song by comparing the start time difference between each dry note sequence and the reference note sequence. If the start time difference corresponding to any dry note sequence is within a preset time difference range, the server determines that the dry audio file corresponding to that dry note sequence meets the rhythmic consistency judgment standard. In practical applications, the time difference range can be set to less than 50ms.

[0147] Then, the server can extract the fundamental frequency curve of each dry audio file in the first audio set and calculate the curve similarity between the fundamental frequency curve of each dry audio file and the reference fundamental frequency curve corresponding to the target song; the fundamental frequency difference between the fundamental frequency of each dry audio file and the reference fundamental frequency corresponding to the target song is measured by cosine similarity and minimum mean square error.

[0148] Cosine similarity can be defined as:

[0149]

[0150] The minimum mean square error (MSE) can be defined as:

[0151]

[0152] For any dry audio file, if the cosine similarity between the fundamental frequency curve of the dry audio file and the reference fundamental frequency curve is greater than a preset similarity threshold, and / or the minimum mean square error between the fundamental frequency curve of the dry audio file and the reference fundamental frequency curve is less than a preset error threshold, then the server determines that the dry audio file meets the tone consistency judgment standard. In practical applications, the similarity threshold can be set to 0.9, and the error threshold can be set to 0.01.

[0153] The server can add dry audio that simultaneously meets the rhythm consistency judgment standard and the pitch consistency judgment standard from the first audio set to the second audio set as the target pitch audio.

[0154] The technical solution of this embodiment obtains the rhythm sequence of each dry audio in the first audio set and determines the rhythm difference between the rhythm sequence of each dry audio in the first audio set and the reference rhythm sequence corresponding to the target song; then obtains the fundamental frequency of each dry audio in the first audio set and determines the fundamental frequency difference between the fundamental frequency of each dry audio in the first audio set and the reference fundamental frequency corresponding to the target song; finally, determines at least one target pitch audio in the first audio set and adds each target pitch audio to the second audio set. In this way, based on the difference between the fundamental frequency and rhythm of each dry audio and the reference audio, it is possible to accurately determine whether the pitch of each dry audio meets the preset conditions.

[0155] In another embodiment, such as Figure 4 As shown, an audio processing method is provided, which is applied to... Figure 1 Taking the server in the example, the following steps are included:

[0156] Step S402: Obtain the dry audio set corresponding to the target song. The dry audio set is obtained by collecting the sound signal generated by at least one target object singing the target song.

[0157] Step S404: Remove dry audio files with abnormal sound quality from the dry audio set to obtain the removed dry audio files; abnormal sound quality includes at least one of the following: noise, popping sounds, backsound of accompaniment, audio length less than a preset length threshold, and audio energy less than a preset energy threshold.

[0158] Step S406: Input each removed dry audio file into the sound quality scoring model to obtain the sound quality score corresponding to each removed dry audio file.

[0159] Step S408: Add the dry audio files that have been removed and whose sound quality scores are greater than or equal to a preset score threshold to the first audio set.

[0160] Step S410: Obtain the rhythm sequence of each dry audio in the first audio set, and determine the rhythm difference between the rhythm sequence of each dry audio in the first audio set and the baseline rhythm sequence corresponding to the target song.

[0161] Step S412: Obtain the fundamental frequency of each dry audio in the first audio set, and determine the fundamental frequency difference between the fundamental frequency of each dry audio in the first audio set and the reference fundamental frequency corresponding to the target song.

[0162] Step S414: Determine at least one target pitch audio in the first audio set and add each target pitch audio to the second audio set; the rhythm difference corresponding to the target pitch audio is less than a preset rhythm difference level, and the fundamental frequency difference corresponding to the target pitch audio is less than a preset fundamental frequency difference level.

[0163] Step S416: Determine a preset number of dry audio files in the second audio set as at least one target dry audio file.

[0164] Step S418: Obtain the key of the target song and use the chord progression that matches the key of the song to perform pitch shifting on each target dry audio file to obtain the pitch-shifted dry audio file.

[0165] Step S420: Perform fusion processing on each pitch-shifted dry audio to obtain the target harmonic audio.

[0166] Step S422: Mix the target harmony audio with the accompaniment audio corresponding to the target song to output the target accompaniment audio with chorus sound effects.

[0167] It should be noted that the specific limitations of the above steps can be found in the specific limitations of an audio processing method described above.

[0168] In another embodiment, such as Figure 5 As shown, a schematic diagram of the architecture of an audio processing method is provided. Figure 5 The audio processing method includes the following stages:

[0169] Dry sound screening stage:

[0170] The server can select N target dry audio files from the dry audio set that meet the preset requirements in terms of sound quality, rhythm, and pitch, and use them as the original sound source for generating choir accompaniment audio.

[0171] Dry sound preprocessing stage:

[0172] The server uses Dynamic Range Control (DRC) to effectively compress the volume of each target dry audio signal, thereby obtaining the target dry audio signal with adjusted loudness. The preset loudness range can be set to approximately -14dB, ensuring that the filtered effective dry audio signals are kept within a similar loudness range, preventing disharmony in the synthesized sound volume due to individual sounds being too loud or too soft.

[0173] Choral generation stage:

[0174] The server can input the audio of the target song into the tonality detector to determine the key of the song. Then, the server determines the chord tones corresponding to the key of the song, and determines the harmonic position based on the chord tones. Using the harmonic tone indicated by the harmonic position, the server performs pitch shifting on each target dry audio file to obtain the pitch-shifted dry audio (i.e., the wet audio of the choir).

[0175] Post-processing stage of wet sound:

[0176] The server can perform post-processing on the target acoustic audio (i.e., wet acoustic audio). Specifically, the server can perform audio equalization (EQ), dynamic range control (DRC), and add preset reverb effects on the audio signal of the target acoustic audio, and output the target acoustic audio with adjusted sound effects to broaden the sound field spatial sense of the target acoustic audio.

[0177] Accompaniment mixing stage:

[0178] The server can prevent the dry vocals from overpowering the accompaniment by adjusting the energy ratio between the target harmonic audio and the corresponding accompaniment audio. Specifically, the server can adjust the loudness of the target harmonic audio so that the loudness of the adjusted target harmonic audio is less than that of the accompaniment audio. For example, the loudness of the adjusted target harmonic audio can be 3dB lower than that of the accompaniment audio.

[0179] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0180] Based on the same inventive concept, this application also provides an audio processing apparatus for implementing the audio processing method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more audio processing apparatus embodiments provided below can be found in the limitations of the audio processing method described above, and will not be repeated here.

[0181] Each module in the aforementioned audio processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0182] In one embodiment, a computer device is provided, which may be an electronic device, and its internal structure diagram may be as follows: Figure 6As shown. The computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores XX data. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements an audio processing method.

[0183] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0184] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, which, when executed by the processor, causes the processor to perform the steps of the audio processing method described above. The steps of the audio processing method described here may be steps from one of the audio processing methods in the various embodiments described above.

[0185] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, causes the processor to perform the steps of the audio processing method described above. The steps of the audio processing method described here may be steps from one of the audio processing methods in the various embodiments described above.

[0186] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0187] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0188] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0189] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An audio processing method, characterized in that, The method comprises: Obtain the dry audio set corresponding to the target song, and filter out at least one target dry audio that meets preset conditions from the dry audio set; the dry audio set is obtained by collecting the sound signal generated by at least one target object singing the target song; Obtain the key offset that matches the key of the song; the key offset is used to characterize the chord progression that matches the key of the song; Adjust the original pitch corresponding to the key of the song according to the key offset to determine the pitch adjustment target for the key shifting process; Using the aforementioned pitch adjustment targets, the dry audio of each target is subjected to pitch shifting processing to obtain the pitch-shifted dry audio; The modified dry audio files are fused together to obtain the target harmonic audio. The target harmonic audio is mixed with the accompaniment audio corresponding to the target song to output the target accompaniment audio with chorus sound effects.

2. The method according to claim 1, characterized in that, Before the step of performing pitch shifting processing on each of the target dry audio samples to obtain the pitch-shifted dry audio samples, the method further includes: The loudness of each target dry audio audio is adjusted to obtain the target dry audio audio with adjusted loudness; the loudness of each target dry audio audio with adjusted loudness is within a preset loudness range.

3. The method according to claim 1, characterized in that, The process of fusing the modulated dry audio files to obtain the target sound audio includes: Obtain the harmonic weights corresponding to each of the pitch-shifted dry audio samples; According to the harmonic weights corresponding to each of the pitch-shifted dry audio signals, the audio signals corresponding to each of the pitch-shifted dry audio signals are weighted and summed to output the target harmonic audio signal.

4. The method according to claim 1 or 3, characterized in that, After the step of fusing the pitch-shifted dry audio samples to obtain the target and ambient audio samples, the method further includes: The audio signal of the target and acoustic audio is processed with sound effects adjustment, and the sound effects adjusted target and acoustic audio is output. The sound effect adjustment processing includes at least one of audio equalization processing, dynamic range control processing, and reverb effect addition processing.

5. The method according to claim 1, characterized in that, The step of mixing the target harmonic audio with the accompaniment audio corresponding to the target song to output the target accompaniment audio with chorus effects includes: The loudness of the target and acoustic audio frequencies is adjusted so that the loudness of the adjusted target and acoustic audio frequencies is less than the loudness of the accompaniment audio frequencies; The adjusted target harmony audio is mixed with the accompaniment audio to output the target accompaniment audio. The loudness of the target accompaniment audio is adjusted so that the loudness of the adjusted target accompaniment audio is equal to a preset loudness threshold.

6. The method according to claim 1, characterized in that, The step of obtaining the key of the target song includes: The total cumulative duration of each note in the target song is identified; the total cumulative duration of each note is used to characterize the pitch distribution information of the target song. Based on the pitch distribution information and the tonality weights corresponding to each note, calculate the Pearson coefficient corresponding to at least one candidate song tonality. The key of the candidate song with the highest Pearson coefficient is taken as the key of the target song.

7. The method according to claim 1, characterized in that, The step of selecting at least one target dry audio file that meets preset conditions from the dry audio file set includes: Based on the sound quality of each dry audio file in the dry audio file set, a first audio file set is selected from the dry audio file set; the sound quality of each dry audio file in the first audio file set meets the preset sound quality conditions. Based on the pitch of each dry audio file in the first audio set, a second audio set is selected from the first audio set; the pitch of each dry audio file in the second audio set meets the preset pitch conditions. A preset number of dry audio files are determined from the second audio set as the at least one target dry audio file.

8. The method according to claim 7, characterized in that, The step of selecting a first audio set from the dry audio set based on the sound quality of each dry audio file in the dry audio set includes: Remove dry audio with abnormal sound quality from the dry audio set to obtain the removed dry audio; the abnormal sound quality includes at least one of the following: noise, popping sound, backslide of accompaniment, audio length less than a preset length threshold, and audio energy less than a preset energy threshold. The removed dry audio files are input into the sound quality scoring model to obtain the sound quality scores corresponding to each of the removed dry audio files. After removing the dry audio files whose sound quality scores are greater than or equal to a preset score threshold, add them to the first audio set.

9. The method according to claim 7, characterized in that, The step of selecting a second audio set from the first audio set based on the pitch of each dry audio file in the first audio set includes: Obtain the rhythm sequence of each dry audio in the first audio set, and determine the rhythm difference between the rhythm sequence of each dry audio in the first audio set and the reference rhythm sequence corresponding to the target song; Obtain the fundamental frequency of each dry audio file in the first audio set, and determine the fundamental frequency difference between the fundamental frequency of each dry audio file in the first audio set and the reference fundamental frequency corresponding to the target song; At least one target pitch audio is determined in the first audio set, and each target pitch audio is added to the second audio set; the rhythm difference corresponding to the target pitch audio is less than a preset rhythm difference level, and the fundamental frequency difference corresponding to the target pitch audio is less than a preset fundamental frequency difference level.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 9.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Audio correction method and device

    CN112309409A