Audio pitch correction method and apparatus, and electronic device
By combining speech recognition and pitch recognition technologies with audio description information to correct the boundaries of the initial pitch information and reconstruct an adaptive original pitch template, the problem of rigid pitch correction results in existing technologies is solved, and higher quality pitch correction effects are achieved.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SHANGHAI SOULGATE TECH CO LTD
- Filing Date
- 2025-07-31
- Publication Date
- 2026-05-07
AI Technical Summary
In existing technologies, vocal correction methods only use pitch information as a reference, resulting in rigid correction results, a lack of comprehensive information decision-making, low intelligence, and an inability to meet users' needs for high-quality music in karaoke scenarios.
Audio description information and initial pitch information are obtained through speech recognition and pitch recognition. The initial pitch information is corrected by combining the audio description information, and an adaptive original pitch template is reconstructed. Based on the adjusted pitch information, pitch correction is performed. A neural network model is used for noise reduction filtering and speed and pitch shifting to generate a pitch-corrected audio that is more consistent with the original song and the user's actual performance.
It improves the accuracy and quality of audio editing, ensuring that the pitch of the edited audio is more consistent with the original song and the user's actual singing performance, thus enhancing the quality of music in karaoke scenarios.
Smart Images

Figure CN2025111711_07052026_PF_FP_ABST
Abstract
Description
An audio sound correction method, apparatus and electronic device
[0001] This application claims priority to Chinese Patent Application No. 202411507119.8, filed on October 28, 2024, entitled "An Audio Correction Method, Apparatus and Electronic Device", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the technical field of audio processing, and in particular to an audio sound correction method, apparatus and electronic device. Background Technology
[0003] With the development of mobile internet, karaoke apps have gradually become an important form of leisure and entertainment. The application of vocal enhancement technology allows users to enjoy a better singing experience while karaoke, enabling even those with average singing skills to achieve satisfactory results.
[0004] Intelligent vocal correction technology is a basic algorithm processing capability in the context of internet-based entertainment karaoke. By analyzing the singer's vocal data and processing the vocal signal based on reference song information, it obtains vocal audio that is closer to the original singer's performance level, greatly improving the singer's performance quality and producing high-quality musical works.
[0005] In related technologies, after obtaining the user's singing audio data, the audio data is analyzed to obtain pitch information. The obtained pitch information is then compared with the spectrogram information to determine the pitch adjustment data for each note. Based on this adjustment data, the singing audio data is adjusted, resulting in an audio that more closely resembles the original recording in pitch trend. However, using only pitch information as a reference for pitch correction results in a rigid and inflexible tone, leading to generally poor audio quality after adjustment.
[0006] Therefore, how to provide a solution to improve the quality of audio correction has become a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0007] The purpose of this application is to provide an audio editing method, apparatus, and electronic device that can improve the audio editing results.
[0008] Firstly, an audio editing method is provided, including:
[0009] Obtain the singing audio to be processed;
[0010] The singing audio to be processed is subjected to speech recognition to obtain audio description information; and the singing audio to be processed is subjected to pitch recognition to obtain initial pitch information.
[0011] The initial pitch information is corrected for boundary information based on the audio description information to obtain the pitch information;
[0012] The original song pitch template corresponding to the singing audio to be processed is determined based on the audio description information, and the original song pitch template corresponding to the singing audio to be processed is corrected based on the pitch information to obtain the reference original song pitch template.
[0013] Based on the original pitch template and the pitch information, the pitch information to be adjusted is determined, and the pitch of the singing audio to be processed is corrected based on the pitch information to obtain the corrected audio.
[0014] In a preferred embodiment, this application can be further configured such that: the step of performing speech recognition on the singing audio to be processed to obtain audio description information includes: extracting the acoustic features of the singing audio to be processed;
[0015] Based on the acoustic features and the decoding diagram, the acoustic features are decoded to obtain a phoneme sequence, wherein the decoding diagram is generated based on a phoneme-level acoustic model, a lyrics-based language model, and a phoneme-lyrics conversion table;
[0016] The phoneme sequence is mapped to the lyrics text to obtain audio description information.
[0017] In a preferred embodiment, this application can be further configured such that the audio description information includes: lyrics information and phoneme boundary time information;
[0018] The step of correcting the initial pitch information based on the audio description information to obtain the pitch information includes:
[0019] Based on the boundary time information corresponding to the phonemes in the audio description information, filter out invalid pitch information in the initial pitch information;
[0020] Obtain the pitch information of the preceding and following audio segments of the singing audio to be processed;
[0021] Based on the preceding and / or following pitch information, the overall pitch of the filtered pitch information is adjusted.
[0022] In a preferred embodiment, this application can be further configured as follows: the step of correcting the pitch of the singing audio to be processed based on the adjusted pitch information to obtain the corrected audio includes:
[0023] The noise reduction filter is applied to the singing audio to be processed based on the first method and / or the second method to obtain the filtered audio.
[0024] Based on the adjusted pitch information, the filtered audio is corrected to obtain the corrected audio.
[0025] The first approach involves removing background music and environmental noise from the audio using the Speak-X module, and then extracting the human voice using a neural network model. The second approach is based on a multi-subband neural network model to eliminate background noise in the audio.
[0026] In a preferred embodiment, this application can be further configured as follows: the step of correcting the original pitch template corresponding to the singing audio to be processed based on the pitch information to obtain a reference original pitch template includes:
[0027] Based on the pitch information, determine the average pitch and the pitch of each individual note in the singing audio to be processed;
[0028] If the difference between the average pitch value and the average target pitch value corresponding to the original pitch template exceeds the first preset pitch threshold, the pitch of the original pitch template is adjusted as a whole to obtain the first original pitch template, and the difference between the average pitch value and the average target pitch value corresponding to the first original pitch template does not exceed the first preset pitch threshold.
[0029] If the difference between the pitch of the individual note and the pitch of the target note in the first original pitch template exceeds a second preset pitch threshold, then a first pitch set is determined, wherein the target note corresponds to the individual note, and the first pitch set is a set of pitches that differ from the target note by a preset pitch.
[0030] If there exists a pitch in the first pitch set whose pitch difference from the individual pitch is within a second preset pitch threshold, then the pitch difference from the individual pitch within the second preset pitch threshold will be used as the dynamic adjustment target for the target pitch.
[0031] If there is no pitch in the first pitch set that has a pitch difference of the second preset pitch threshold with respect to the individual pitch, then a second pitch set is determined, which is a set of pitches within the key of the original song; a dynamic adjustment target that meets the target conditions is selected from the second pitch set, the target conditions including: a pitch within the key that is closest to the target pitch and has a pitch difference of the second preset pitch threshold with respect to the individual pitch;
[0032] After determining the dynamic adjustment targets for all individual notes, the first original pitch template is adjusted according to the dynamic adjustment targets for all individual notes to obtain the reference original pitch template.
[0033] In a preferred embodiment, this application may be further configured as follows: after determining the dynamic adjustment targets for all individual notes, adjusting the first original pitch template according to the dynamic adjustment targets for all individual notes to obtain the reference original pitch template, the application further includes at least one of the following:
[0034] The reference target signal whose time domain variation coefficient in the reference original pitch template exceeds a preset multiple is processed by white noise insertion, while ensuring that the starting position remains aligned, to obtain the aligned reference original pitch template;
[0035] If the target note in the first original pitch template corresponding to a single note cannot be determined, the pitch of the first original pitch template is adjusted as a whole to a pitch within the key, and the starting point of the sound is adjusted to a preset time position.
[0036] In a preferred embodiment, this application can be further configured such that, after determining the adjusted pitch information based on the reference original pitch template and the pitch information, it also includes:
[0037] A two-level neural network vocoder is obtained, wherein the two-level neural network vocoder includes a first GAN network structure and a second GAN network structure;
[0038] Based on the singing audio to be processed, the first GAN network structure is used to perform speed and pitch shifting processing to obtain the initial audio.
[0039] Based on the initial audio, the sampling rate is increased using a second GAN network structure to obtain high-quality singing audio to be processed.
[0040] In a preferred embodiment, this application can be further configured such that, after determining the adjusted pitch information based on the reference original pitch template and the pitch information, it also includes:
[0041] Based on the individual words whose pitch information has been adjusted, determine the first musical score information corresponding to each word from the MIDI file of the original song;
[0042] Based on each phoneme of the single word whose pitch information is adjusted, determine the second musical score information corresponding to each phoneme from the first musical score information;
[0043] Based on each waveform of the phoneme, determine the third musical score information corresponding to each waveform from the second musical score information;
[0044] Based on the third score information corresponding to all waveforms, the final pitch adjustment information is determined to achieve smooth processing of the pitch adjustment information.
[0045] Secondly, an audio correction device is provided, comprising:
[0046] The audio acquisition module is used to acquire the singing audio to be processed;
[0047] The recognition module is used to perform speech recognition on the singing audio to be processed to obtain audio description information; and to perform pitch recognition on the singing audio to be processed to obtain initial pitch information.
[0048] The boundary correction module is used to correct the initial pitch information based on the audio description information to obtain the pitch information.
[0049] The dynamic template generation module is used to determine the original song pitch template corresponding to the singing audio to be processed based on the audio description information, and to correct the original song pitch template corresponding to the singing audio to be processed based on the pitch information to obtain the reference original song pitch template.
[0050] The pitch correction module is used to determine the adjusted pitch information based on the reference original pitch template and the pitch information, and to perform pitch correction on the singing audio to be processed based on the adjusted pitch information to obtain the pitch-corrected audio.
[0051] Thirdly, an electronic device is provided, comprising:
[0052] One or more processors;
[0053] Memory;
[0054] One or more applications, wherein the applications are stored in memory and configured to be executed by one or more processors, the applications being configured to: perform operations corresponding to the methods shown in any possible implementation of the first aspect.
[0055] Fourthly, a computer-readable storage medium is provided, the storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded by a processor and performs the steps of the method shown in any possible implementation of the first aspect.
[0056] Fifthly, a computer program product is provided, comprising a computer program that, when executed by a processor, implements operations corresponding to the methods shown in any possible implementation of the first aspect.
[0057] In summary, the audio correction method provided in this application includes the following beneficial technical effects: performing speech recognition and pitch recognition on the singing audio to be processed, obtaining audio description information and initial pitch information respectively, enabling simultaneous understanding of the audio content and pitch characteristics; using the audio description information to correct the boundary information of the initial pitch information, effectively removing noise or invalid parts in the pitch information and improving the accuracy of the pitch information; adaptively reconstructing the pitch template based on the pitch information to obtain a reference original song pitch template that allows for a range of speed and pitch changes; based on the corrected reference original song pitch template and pitch information, determining and adjusting the pitch information and performing correction, thereby achieving precise correction of the singing audio, making the corrected audio more consistent with the original song and the user's actual singing performance in terms of pitch, and improving the overall audio quality.
[0058] This application also provides an audio sound correction device and an electronic device, both of which have the aforementioned beneficial effects, which will not be elaborated here. Attached Figure Description
[0059] To more clearly illustrate the technical solutions of the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0060] Figure 1 is a schematic diagram of the process of vocal tone correction in related technologies;
[0061] Figure 2 is a schematic diagram of an application scenario of an audio correction method provided in the embodiments of this application;
[0062] Figure 3 is a schematic flowchart of an audio correction method provided in an embodiment of this application;
[0063] Figure 4 is a schematic flowchart of a smoothing process provided in an embodiment of this application;
[0064] Figure 5 is a schematic diagram of a specific audio correction scheme provided in an embodiment of this application;
[0065] Figure 6 is a comparison chart of the MOS score of a tone correction effect provided in an embodiment of this application;
[0066] Figure 7 is a schematic diagram of the structure of a device provided in an embodiment of this application;
[0067] Figure 8 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0068] This specific embodiment is merely an explanation of this application and is not intended to limit it. After reading this specification, those skilled in the art can make modifications to this embodiment without contributing any inventive step, but such modifications are protected by patent law as long as they are within the scope of the claims of this application.
[0069] It should be noted that, in the optional embodiments of this application, the data related to object information, when applied to specific products or technologies, requires the permission or consent of the object. Furthermore, the collection, use, and processing of this data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. In other words, if the embodiments of this application involve data related to an object, it must be obtained with the object's authorization and consent, the authorization and consent of relevant departments, and in accordance with the relevant laws, regulations, and standards of the country and region. If the embodiments involve personal information, the acquisition of all personal information requires the individual's consent. If sensitive information is involved, the separate consent of the information subject is required. The embodiments also need to be implemented with the object's authorization and consent.
[0070] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0071] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.
[0072] Intelligent vocal enhancement technology is now widely used in karaoke apps, serving as a fundamental algorithmic processing capability in the broader internet-based entertainment karaoke landscape. Through machine learning algorithms, it analyzes the singer's vocal data and processes the vocal signal based on reference song information to obtain audio that more closely resembles the original singer's performance, significantly improving the singer's performance quality and producing high-quality musical works. Simultaneously, intelligent vocal enhancement technology greatly lowers the barrier to entry for ordinary karaoke users, bringing a stable increase in active users to the corresponding apps.
[0073] The primary focus of vocal editing solutions when processing vocal data is pitch information, specifically the frequency information of the audio. They use frequency analysis tools to analyze the pitch information, compare it with the musical score, and adjust the pitch on a per-note basis. This adjustment process often employs acoustic modulation techniques, performing frequency domain audio transformations, processing the audio segments, and then splicing them together to produce vocal data that closely approximates the original score in pitch.
[0074] As shown in Figure 1, this is a flowchart illustrating the process of vocal tone correction in related technologies. The vocal tone correction scheme in these technologies has two core steps:
[0075] Vocal Analysis Stage: During the recording and acquisition of user vocal data, a third-party audio interface is typically used. The audio sampling rate is limited, often only reaching a maximum of 16kHz. Basic linear noise reduction can be employed to obtain the cleanest possible audio. After obtaining the audio data, pitch information is obtained using a pitch calculation tool, as shown in Figure 1, where A is represented on the time-frequency (pitch) coordinate system.
[0076] Vocal processing stage: Based on the score information (usually B in Figure 1, as shown in the illustration on the time-pitch coordinate system), the obtained audio pitch information is compared to determine the pitch adjustment data for each note, i.e., the adjustment information shown in Figure 1; after obtaining the adjustment information, the audio is adjusted in the frequency domain using a frequency domain transformation tool; the final output audio is closer to the original audio in the frequency dimension, as shown in C in Figure 1, as shown in the illustration on the time-frequency (pitch) coordinate system, which is the key information data of the adjusted audio, and the pitch trend is more in line with the original, thus obtaining the processed audio data.
[0077] Current pitch correction techniques rely solely on pitch information for decision-making, failing to make comprehensive judgments based on integrated information. This results in low intelligence and rigid, inflexible pitch correction outcomes. Other information, such as textual and modal information, is ignored and not extracted to improve intelligence and enable more rational pitch correction operations.
[0078] Therefore, improving the quality of user-generated music in karaoke scenarios, lowering the participation threshold, and making it easier for users to experience the joy and satisfaction of singing songs, as well as facilitating the production of high-quality and impactful audio recordings; and in music creation scenarios, making it easier for authors to verify the effects of their creations, and enabling them to produce high-quality original music without requiring extremely high singing skills, are areas that require the attention of those skilled in the art.
[0079] Based on this, referring to Figure 2, which is a schematic diagram of an application scenario of an audio correction method provided in the embodiments of this application, the audio correction method can be applied to an audio correction system.
[0080] In some embodiments, the audio editing system may consist only of a first electronic device for acquiring the singing audio to be processed and for editing the singing audio. The first electronic device may be a terminal device, such as a smartphone, tablet, laptop, desktop computer, karaoke microphone, etc., but is not limited to these.
[0081] In some embodiments, the audio editing system may include a data collection device and a second electronic device, etc. The data collection device is used to collect the singing audio to be processed, and the second electronic device is used to implement the audio editing method. The method provided in this application embodiment can be executed by the second electronic device, which is a server. The server can be an independent physical server, a server cluster composed of multiple physical servers, a distributed system, or a cloud server providing cloud computing services. The data collection device can be a terminal device, such as a smartphone, tablet, laptop, desktop computer, karaoke microphone, etc., but is not limited to these. The data collection device and the second electronic device can be directly or indirectly connected via wired or wireless communication, which is not limited in this application embodiment.
[0082] It is understood that the above is only one example, and this embodiment is not limited here.
[0083] The method for implementing audio correction using electronic devices (first electronic device / second electronic device) is further elaborated.
[0084] The electronic device may include a front-end speech recognition module, a singing pitch recognition module, and a pitch correction processing module that combines information from multiple sources to make pitch correction decisions and produce pitch-corrected audio. The initial pitch information obtained by the singing pitch recognition module is corrected by the audio description information obtained by the speech recognition module to obtain the initial pitch information. Based on the pitch information, the pitch correction processing module adaptively adjusts the original song pitch template to generate a new song template with a range of speed and pitch changes, so as to combine the pitch information and the new song template to jointly correct the singing data.
[0085] In some embodiments, the electronic device may further include a noise reduction filtering module, which may be a noise reduction filtering module with custom frequency band adjustment. The singing data is noise-reduced using a streaming parallel noise reduction filtering module. Further, the pitch correction processing module includes: an adaptive pitch correction template reconstruction unit, a smoothing unit, and a variable speed and pitch shifting unit.
[0086] The process involves the user recording their singing audio, which then passes through three data processing units: a speech recognition module, a pitch recognition module, and a noise reduction filter module with custom frequency band adjustment.
[0087] The speech recognition module generates audio description information, and the pitch recognition module generates initial pitch information. Pitch information is obtained based on the audio description information and the initial pitch information. Based on the pitch information, the adaptive pitch correction template reconstruction unit of the pitch correction processing module generates a reference original song pitch template within the allowed speed and pitch shifting range. Based on the pitch information, and using the reference original song pitch template, accurate pitch information is generated to help refine the speed and pitch shifting information to obtain adjusted pitch information. The refined speed and pitch shifting information (adjusted pitch information) passes through the smooth unit to generate a more reliable small-granular audio operation information sequence for use by the speed and pitch shifting unit. The noise reduction and filtering module outputs audio data as the operation material for the speed and pitch shifting algorithm in subsequent processes. Speed and pitch shifting, after pre-processing, processes the audio in parallel, completing audio processing with an effectiveness not exceeding RTF 0.03.
[0088] Specifically, this application provides an audio correction method, as shown in Figure 3, which includes:
[0089] S101. Obtain the singing audio to be processed;
[0090] The audio to be processed is the real-time audio of the user singing while singing karaoke.
[0091] S102. Perform speech recognition on the singing audio to be processed to obtain audio description information; and perform pitch recognition on the singing audio to be processed to obtain initial pitch information.
[0092] When performing speech recognition on the singing audio to be processed, machine learning algorithms can be used. By using speech recognition algorithms that improve the accuracy of boundary recognition, information in both semantic and boundary aspects is obtained. That is, the audio description information includes, but is not limited to, lyrics information and boundary time information corresponding to each phoneme. The time boundary information of each phoneme is the start and end time points of each phoneme.
[0093] Initial pitch information is the original pitch profile of a song's performance. Specific initial pitch information can be obtained using fundamental frequency detection algorithms (such as autocorrelation or cepstral methods).
[0094] S103. Correct the initial pitch information by boundary information based on the audio description information to obtain the pitch information;
[0095] In related technologies, it is difficult to perform precise processing in the time dimension during the pitch correction process. Only the pitch can be controlled, but the rhythm accuracy cannot be controlled. Furthermore, the basic unit of audio processing is an audio segment of a complete note length analyzed during the vocal analysis stage. The overall frequency multiplier adjustment is performed within this time length. Due to the imprecision in time, this audio segment may not be the audio that ideally needs frequency adjustment. Uniformly adjusting a fixed multiplier will cause some high-frequency sounds and audio segments without a stable fundamental frequency to be incorrectly adjusted. This is visually manifested as what we often call "pops", "audio static", "audio stuttering", and "audio lag". In this embodiment, the initial pitch information based on the pitch recognition output is combined with the boundary time information corresponding to the phonemes of the accurate speech recognition result to filter out invalid pitch information in the initial pitch information, such as pitch information corresponding to non-speech segments (e.g., silence, noise segments). These segments usually exhibit unstable pitch values or are extremely low / high, so as to calibrate the boundaries of the inherent pitch information and ensure that the pitch information of each effective speech segment (mainly vowel segments) is strictly aligned with the phoneme boundaries in the speech recognition result in the time dimension, so that the pitch data can be accurately truncated in the time dimension.
[0096] In this embodiment, the traditional pitch extraction algorithm and the speech recognition result are cross-validated to complete the accurate boundary information for the pitch information, making the obtained pitch information more accurate.
[0097] Furthermore, in order to accurately correct the pitch, in this embodiment of the application, the initial pitch information can be corrected again by combining the pitch information corresponding to the preceding and following audio of the singing audio to be processed. Specifically, invalid pitch information in the initial pitch information is filtered out according to the boundary time information corresponding to the phonemes in the audio description information; the preceding pitch information corresponding to the preceding singing audio and the following pitch information corresponding to the following singing audio are obtained; and the overall pitch is adjusted based on the preceding pitch information and / or the following pitch information.
[0098] An overall correction value is determined based on the preceding and / or following pitch information, and this overall correction value is used to further correct the filtered pitch information. In one possible implementation, the overall correction value can be obtained by subtracting the average pitch of the preceding or following audio from the average pitch of the audio to be processed. In another possible implementation, the average of the preceding and following audio pitches is calculated to obtain a composite average pitch value, and the overall correction value is obtained by subtracting the average pitch of the audio to be processed from the composite average pitch value. The specific method chosen in this application is not limited; users can choose according to their actual needs.
[0099] By utilizing the phoneme boundary time information in the audio description information, the invalid parts in the initial pitch information are effectively filtered out, reducing noise interference in pitch recognition. By combining the pitch information of the previous and subsequent singing audio, the filtered pitch is adjusted as a whole, making the corrected pitch information smoother and more in line with musical rules.
[0100] S104. Determine the original music pitch template corresponding to the singing audio to be processed based on the audio description information, and correct the original music pitch template corresponding to the singing audio to be processed based on the pitch information to obtain the reference original music pitch template.
[0101] Based on the audio description information and the reference description information, the target description information corresponding to the audio description information in the reference description information is determined. The lyrics information in the target description information corresponds to the lyrics information in the audio description information, and the time information of the phoneme boundaries in the target description information corresponds to the time information of the phoneme boundaries in the audio description information. Then, based on the target description information, the original song pitch template corresponding to the singing audio to be processed is determined. The original song pitch template refers to the ideal pitch sequence of each note or syllable in the original song, representing the standard pitch profile of the song.
[0102] The original song pitch template is used to represent a pitch template that is adjusted according to the singer's actual pitch while retaining the pitch characteristics of the original song. The original song pitch template is closer to the singer's actual performance.
[0103] S105. Based on the original song's pitch template and pitch information, determine the pitch information to be adjusted, and perform pitch correction on the singing audio to be processed based on the adjusted pitch information to obtain the corrected audio.
[0104] Based on the pitch difference between the original song's pitch template and pitch information, the pitch adjustment information is obtained. This adjusted pitch information is then used to correct the pitch of the singing audio being processed, resulting in a more reasonable and improved audio correction.
[0105] As can be seen, in this embodiment, speech recognition and pitch recognition are performed on the singing audio to be processed to obtain audio description information and initial pitch information, respectively, which can simultaneously understand the content and pitch characteristics of the audio; the initial pitch information is corrected by using the audio description information to effectively remove noise or invalid parts in the pitch information, thereby improving the accuracy of the pitch information; the pitch template is adaptively reconstructed based on the pitch information to obtain a reference original song pitch template that allows for a range of speed and pitch changes; based on the corrected reference original song pitch template and pitch information, the pitch information is adjusted and pitch correction is performed, thereby achieving precise pitch correction of the singing audio, making the pitch-corrected audio more consistent with the original song and the user's actual singing performance, thus improving the overall audio quality.
[0106] One possible implementation of this application embodiment is that S102 performs speech recognition on the singing audio to be processed to obtain audio description information, including: S1021-S1023, wherein:
[0107] S1021. Extract the acoustic features of the singing audio to be processed;
[0108] Extract standard features (F BANK), i.e., acoustic features, from the singing audio to be processed.
[0109] Furthermore, additional feature processing can be performed on the singing audio to be processed, such as handling changes in pitch and tone, before feature extraction. Moreover, the extracted acoustic features can be normalized to ensure they meet certain standards.
[0110] S1022. Based on the acoustic features and the decoding diagram, the acoustic features are decoded to obtain the phoneme sequence. The decoding diagram is generated based on the phoneme-level acoustic model, the lyrics-based language model, and the phoneme-lyrics conversion table.
[0111] In this embodiment of the application, an acoustic model with single-tone phoneme information and a speech model based on lyrics are adopted to improve the boundary recognition capability.
[0112] The redesigned speech recognition module achieved a boundary recognition accuracy of 93.2% at the 30ms level.
[0113] In some embodiments, to obtain more accurate speech recognition results, the resulting decoded map can more accurately recognize speech. In this embodiment, an acoustic model, a language model, and a pronunciation dictionary (phoneme-lyrics conversion table) are used when generating the decoded map to finally obtain the decoded map. The process of generating the decoded map may include:
[0114] A1. Acquire audio data, including: ordinary speech data and singing audio data. Ordinary speech data includes various daily conversations, speeches, etc.; singing audio data includes a large amount of singing audio data to ensure data diversity, including songs of different styles, languages and pitches.
[0115] For annotating transcribed text from ordinary speech data, phoneme-level annotation can also be performed; for annotating lyrics from singing audio data, phoneme-level annotation is also performed, taking into account that the pronunciation of singing may vary.
[0116] Standard features are extracted from audio data. Additional feature processing can be performed on singing data, such as handling variations in pitch and tone.
[0117] The extracted standard features are normalized to ensure that all data are trained under the same standard.
[0118] A2. Acoustic Model Training. The standard features of the pitch scale are aligned with the phoneme-level annotations using a phoneme alignment tool. The acoustic model is trained using phonemes as modeling units, combining phoneme labels with deep learning techniques (such as DNN, CNN, or LSTM). Because singing audio data is used in the training, the resulting acoustic model can handle audio with singing style.
[0119] A3. Language Model Training. Language models can include HMM (Hidden Markov Model) structures based on lyrics, where lyrics are used as hidden states, and the relationship between phoneme sequences and lyrics is learned through training; or n-gram language models based on lyrics or RNN-based language models that focus on the language structure and commonly used vocabulary of lyrics.
[0120] A4. Phoneme-vocabulary conversion table maps phoneme sequences to lyrics text, enabling a better integration of phoneme-level models and lyrics-level language models.
[0121] A5. Generate a decoding diagram based on the phoneme-level acoustic model, the lyrics-based language model, and the phoneme-lyrics conversion table.
[0122] In this embodiment of the application, after obtaining the decoding map, the audio to be processed is decoded according to the acoustic features and the decoding map to determine the recognition result of the words in the audio to be processed.
[0123] S1023. Map the phoneme sequence to the lyrics text to obtain audio description information.
[0124] Audio description information includes: lyrics information and boundary time information.
[0125] As can be seen, in this embodiment, by extracting the acoustic features of the singing audio to be processed, and using a decoding graph generated by a phoneme-level acoustic model, a lyrics-based language model, and a phoneme-lyrics conversion table, an accurate phoneme sequence is obtained. This not only improves the accuracy of speech recognition but also ensures that the recognition result accurately reflects the lyrics content in the audio. Then, the phoneme sequence is mapped to the lyrics text to obtain accurate audio description information containing lyrics information and phoneme boundary time information.
[0126] Furthermore, after obtaining the audio description information, for some songs where the ending notes may be relatively long, a post-processing workflow involving boundary operations can be performed. For example, boundary detection can be performed to obtain the separation between lines of lyrics and the temporal position information of natural breaks in words within the lyrics. Then, based on this temporal position information, the boundaries of the lyrics can be smoothed to make them more consistent with natural language segmentation. In addition, the impact of pitch and tone changes on the boundaries can be adjusted in the singing data to ensure the accuracy of the recognition results.
[0127] Furthermore, in karaoke scenarios, due to the uncertainty of the difference between the input singing audio to be processed and the original song template, it is difficult to determine a safe range for audio adjustment. In this application, the template is adaptively reconstructed based on the input singing audio to be processed, so that the adjustment range of the new template matches the capabilities of the speed and pitch shifting algorithm.
[0128] Specifically, the pitch template reconstruction module of this application, referring to the song's key and chord progression, calculates the pitch of alternative notes of different recommended levels at each point (note) to achieve the reconstruction of the score pitch template.
[0129] One possible implementation of this application embodiment involves correcting the original song pitch template corresponding to the singing audio to be processed based on pitch information to obtain a reference original song pitch template, including:
[0130] The average pitch and individual pitch of the singing audio to be processed are determined based on pitch information.
[0131] If the difference between the average pitch value and the average target pitch value corresponding to the original pitch template exceeds the first preset pitch threshold, the pitch of the original pitch template is adjusted as a whole to obtain the first original pitch template. The difference between the average pitch value and the average target pitch value corresponding to the first original pitch template does not exceed the first preset pitch threshold.
[0132] If the difference between the pitch of a single note and the pitch of the target note in the first original pitch template exceeds the second preset pitch threshold, then the first pitch set is determined, wherein the target note corresponds to the single note, and the first pitch set is the set of pitches that differ from the target note by a preset pitch.
[0133] If there exists a pitch in the first pitch set whose pitch difference from that of a single pitch is within the second preset pitch threshold, then the pitch difference from that of a single pitch within the second preset pitch threshold will be used as the dynamic adjustment target for the target pitch.
[0134] If there is no pitch in the first pitch set that has a pitch difference of the second preset pitch threshold with respect to a single pitch, then a second pitch set is determined. The second pitch set is a set of pitches within the key of the original piece. From the second pitch set, a dynamic adjustment target that meets the target conditions is selected. The target conditions include: a pitch within the key that is closest to the target pitch and has a pitch difference of the second preset pitch threshold with respect to a single pitch.
[0135] After determining the dynamic adjustment targets for all individual notes, the first original pitch template is adjusted according to the dynamic adjustment targets for all individual notes to obtain the reference original pitch template.
[0136] Specifically, when the pitch to be adjusted deviates too much or there are significant changes in the time domain, potentially exceeding the design capabilities of the speed and pitch shifting module, the pitch correction reference target is dynamically adjusted to match the capabilities of the speed and pitch shifting module.
[0137] The algorithm determines whether the difference (in absolute value form) between the average pitch of the entire segment and the average target pitch of the song segment exceeds a first preset pitch threshold. The first preset pitch threshold can be set according to actual needs. For example, the first preset pitch threshold can be 10 semitones. If it exceeds the first preset pitch threshold, the input audio cannot be forcibly moved closer to the target audio, which would exceed the upper limit of the algorithm's capabilities and result in severe sound quality degradation. In this embodiment, the original song pitch template is adjusted by an integer number of octaves. Adjusting by octaves will not affect the use of the accompaniment and other music resources until the difference in the average pitch of the entire segment just does not exceed the first preset pitch threshold, thus obtaining the first original song pitch template. If the average pitch is greater than the average target pitch, the original song pitch template is adjusted; if the average pitch is not greater than the average target pitch, the original song pitch template is adjusted lower; if it does not exceed the first preset pitch threshold, the original song pitch template is used as the first original song pitch template.
[0138] For each individual note in the entire segment, determine whether the difference (in absolute value form) between the pitch of the individual note and the pitch of the target note in the corresponding first original pitch template exceeds a second preset pitch threshold. The second preset pitch threshold can be set according to actual needs. For example, the second preset pitch threshold is 8 semitones, and the second preset pitch threshold is less than the first preset pitch threshold.
[0139] If the pitch does not exceed the second preset pitch threshold, the pitch of the target note in the first original pitch template remains unchanged;
[0140] If the pitch exceeds the second preset pitch threshold, the pitches of the target sound that are far from the preset pitch are selected as the first pitch set. The preset pitch can be set according to actual needs. For example, the preset pitch can be 2 degrees. If there are pitches in the first pitch set whose pitch difference with the individual sound is within the second preset pitch threshold, the pitches whose pitch difference with the individual sound is within the second preset pitch threshold are used as the dynamic adjustment target of the target sound. For example, if the pitch of the individual sound is a and the pitch of the target sound is b = a + 9 semitones, then the determined first pitch set is a + 11 and a + 7. In this case, b is modified to a + 7.
[0141] If the first set of pitches does not contain a pitch whose difference from the individual note falls within the second preset pitch threshold, then the tones within the key are used as the second set. From this second set, several tones within the first key that have a pitch difference from the individual note within the second preset pitch threshold are selected. From these first tones, the second tone within the key that is closest to the target note is selected as the dynamic adjustment target. Specifically, the tones within the second set can be traversed from closest to furthest pitch difference from the individual note to find the note closest to the target note whose pitch difference is within 8 semitones of the individual note, as the dynamic adjustment target.
[0142] After determining the dynamic adjustment targets for all individual notes, the pitch template of the first original track is adjusted according to these targets to obtain the reference original track pitch template. Using the DTW (Dynamic Time Warping) algorithm, a distance calculation expression is constructed based on the key signature to obtain the most economical new template, thus providing different pitch correction strategies for different audio files.
[0143] As can be seen, in the embodiments of this application, when the pitch deviation of the target pitch correction is too large, which may far exceed the design capability of the speed and pitch shifting module, the pitch correction reference target is dynamically adjusted to match the capability of the speed and pitch shifting module. Specifically, the overall template is first adjusted based on the average pitch value, and after the adjustment, a second adjustment is made based on the difference between the individual notes and the target notes to achieve the reconstruction of the score.
[0144] One possible implementation of this application embodiment, after determining the dynamic adjustment targets for all individual notes, adjusts the first original pitch template according to the dynamic adjustment targets for all individual notes to obtain the reference original pitch template, and further includes:
[0145] The reference target signal whose time domain variation coefficient in the original pitch template exceeds a preset multiple is processed by white noise insertion, while ensuring that the starting position remains aligned, to obtain the aligned original pitch template.
[0146] The time-domain variation coefficient can be customized by the user; for example, the time-domain variation coefficient is 2.
[0147] In the embodiments of this application, when the pitch correction target is significantly altered in the time domain to the extent that it may exceed the design capabilities of the variable speed and pitch module, the pitch correction reference target is dynamically adjusted to achieve the adjustment of the musical score.
[0148] As can be seen, by processing the reference target signal whose time-domain variation coefficient in the original pitch template exceeds a preset multiple using white noise insertion, abnormal or abrupt pitch changes in the template are effectively smoothed out, avoiding the adverse effects of extreme pitch changes on subsequent pitch correction. Simultaneously, by ensuring the alignment of the white noise insertion's starting position with the original signal, the continuity and consistency of the processed template on the time axis are guaranteed, making the aligned original pitch template more consistent with the natural laws of music.
[0149] One possible implementation of this application embodiment, after determining the dynamic adjustment targets for all individual notes, adjusts the first original pitch template according to the dynamic adjustment targets for all individual notes to obtain the reference original pitch template, and further includes:
[0150] If the target note in the first original pitch template corresponding to a single note cannot be determined, the pitch of the first original pitch template is adjusted as a whole to the pitch within the key, and the starting point of the sound is adjusted to the preset time position.
[0151] In cases where lyrics are sung incorrectly and the target pitch cannot be determined, the pitch is adjusted to the corresponding pitch within the paragraph, and the starting point of the pronunciation is adjusted to an integer multiple of the 1 / 16 beat position.
[0152] It is evident that, when the target note in the original pitch template corresponding to a single note cannot be determined, global adjustments can improve the overall pitch harmony of the template to a certain extent. Furthermore, adjusting the position of the articulation point regulates the temporal characteristics of the target note.
[0153] In this embodiment of the application, the audio adjustment range can be controlled within a certain range by using the template reconstruction algorithm. All audio can be processed by the audio editing module, so that the output sound quality damage rate is lower than the acceptable range (0.005%) of streaming music products. After all audio is processed by audio editing, all audio is placed in the accompaniment, and its rhythmic harmony is fully guaranteed.
[0154] In one possible implementation of this application, after obtaining the pitch adjustment information, it is necessary to perform pitch correction on the singing audio to be processed. At this time, the quality of the singing audio to be processed is crucial to the effect of pitch correction. In order to improve the pitch correction effect, in this application embodiment, sound enhancement processing can be performed to obtain cleaner vocal data.
[0155] For sound enhancement technology, a noise reduction module that eliminates sudden and continuous noise (Method 2) can be used, or an echo cancellation module that eliminates leaked background music (Method 1) can be used. Alternatively, the sound can be enhanced by first using an echo cancellation module to eliminate leaked background music, and then by using a noise reduction module to eliminate sudden and continuous noise. This application's embodiments are not limited to these methods; users can configure them according to their actual needs. After obtaining the noise-reduced singing audio to be processed, the audio after pitch correction is obtained based on the noise-reduced audio and adjusted pitch information, further improving the pitch correction effect.
[0156] In one possible implementation, the sound enhancement process by eliminating the echo cancellation module that removes the external background music may include: processing the background music and environmental noise of the singing audio to be processed based on the Speak-X module to obtain the first singing audio; and using a neural network model to extract the human voice from the first singing audio to obtain the second singing audio.
[0157] Then, based on the second vocal audio and adjusting the pitch information, the corrected audio is obtained.
[0158] The algorithm employs a pre-speaker-X and post-neural network structure to eliminate background music generated by the playback device, resulting in a clearer vocal, or second vocal audio. This is particularly important in scenarios where the device is used with external speakers, as it can significantly improve speech clarity.
[0159] The Speak-X module is a complex filter or acoustic model focused on processing background music. In this embodiment, the pre-processor Speak-X module is used to process background music and ambient noise to generate a clear audio signal, i.e., the first vocal audio.
[0160] The training process of the subsequent neural network will be further elaborated, including:
[0161] 1. Data preparation.
[0162] This process collects audio data from various environments, especially background music and clear vocals from playback devices, and gathers audio data recorded for different types of music and vocals. It ensures the data sampling rate covers application requirements, and higher sampling rates (such as 48kHz or 32kHz) can be used to retain sufficient audio detail. In related technologies, third-party recording and noise reduction services are typically used, forcibly locking the audio sampling rate at 16kHz. While this sampling rate meets the needs in conversational scenarios, in karaoke scenarios, the texture of high-frequency vocals is significantly amplified by sound effects, exposing the sound quality issues, particularly the lack of high-frequency information, in 16kHz audio. Therefore, in this embodiment, a high sampling rate can be used for model training, and also when recording the singing audio to be processed, to ensure sufficient detail and reduce the loss of high-frequency information.
[0163] The audio data is annotated in detail, including background music, vocals, and the separation between them.
[0164] Feature extraction is performed on audio data, including Short-Time Fourier Transform (STFT) or other time-frequency domain representations.
[0165] Furthermore, data augmentation can be performed, such as adding different types of noise or adjusting the volume, to simulate variations in various real-world usage scenarios.
[0166] The collected audio data is divided into training, validation, and test sets. Ensure the training set includes a variety of background music and vocal combinations.
[0167] 2. Model design and training.
[0168] The post-processor neural network model can employ a hybrid architecture combining Convolutional Neural Networks (CNNs) and a transformer to further process the output of the pre-processor module, enhance the clarity of the human voice, remove residual background noise or echoes, and extract clear human voices. The embodiments of this application do not limit the number of layers and nodes in the neural network model, as long as it can handle complex audio signals.
[0169] Supervised learning methods are used to train neural network models, and loss functions (such as mean squared error, MSE) are optimized to minimize the gap between predicted noise and actual noise. Based on this gap, the hyperparameters of the model, such as learning rate, batch size, and number of network layers, are adjusted to achieve the best model performance.
[0170] Use a validation set to evaluate the performance of the trained neural network model, focusing on metrics such as echo cancellation effectiveness, speech intelligibility, and sound quality. Based on the validation results, adjust the model architecture or training parameters to improve performance.
[0171] 3. Model testing and evaluation.
[0172] The actual performance of the validated neural network model is evaluated using a test set to obtain the echo cancellation effect, speech intelligibility, and background music suppression effect. The test data should include audio from various device external playback scenarios to verify the model's generalization ability; and the echo cancellation effect, speech intelligibility, and background music suppression effect are evaluated according to the set standards.
[0173] Furthermore, feedback can be collected from actual users to understand the performance of the neural network model in real-world use, and the model can be further optimized based on user feedback.
[0174] In this embodiment, after obtaining the trained model, it is integrated into a practical application, which could be a device for a karaoke scenario, a conferencing system, or a real-time communication system. It is ensured that the deployment environment can handle high-sampling-rate audio data. Performance optimization can also be performed in practical applications, such as accelerating the inference process and reducing computational resource consumption.
[0175] As can be seen, in the embodiments of this application, real-time audio is processed in the scenario of external speaker, the front-end Speak-X module is used to process the background music, and then the clarity of the human voice is enhanced by the rear-end neural network module, and the clear human voice after echo cancellation and background music suppression is output, ensuring that the audio quality is better, and the sound correction effect is better after adjusting the pitch information.
[0176] In another possible implementation, the process of enhancing sound by eliminating sudden and continuous noise through a noise reduction module may include:
[0177] By using a multi-subband neural network model to eliminate background noise in the singing audio to be processed, a third singing audio is obtained.
[0178] Based on the third vocal audio and adjusting the pitch information, the corrected audio is obtained.
[0179] Specifically, acquire audio data from karaoke scenarios, including singing audio and corresponding clean audio, ensuring the data encompasses various styles, pitches, and background noise levels, and that the sampling rate supports a range higher than 32kHz to meet high-frequency requirements. Audio data can also undergo data augmentation processing, including noise addition and volume adjustment, to simulate the actual karaoke environment and sound quality variations. Label lyrics and background noise in the audio data; convert the audio data to a time-frequency domain representation using short-time Fourier transform or other suitable transforms, ensuring that feature extraction retains rich frequency information for high-sampling-rate data.
[0180] A multi-subband neural network model is constructed. The number of subbands is not limited in this application embodiment and can be 24, 20, etc. Each subband is responsible for processing speech signals in a specific frequency band. The hidden layer configuration of the multi-subband neural network model can be fixed at 512 and 384 nodes to increase the network's expressive power and ability to process complex audio signals. The multi-subband neural network model includes multiple convolutional layers (for time-frequency feature extraction), fully connected layers (for feature fusion), and activation functions (such as ReLU).
[0181] The audio data is divided into training, validation and test sets to ensure that the training set contains rich karaoke scene data.
[0182] Supervised learning methods are used to train the model, and the loss function (such as mean squared error, MSE) is optimized to minimize the gap between the predicted noise and the real noise. The learning rate, batch size and regularization parameters are adjusted to obtain the best model performance.
[0183] Use a validation set to evaluate the model's noise reduction performance and speech quality, focusing on metrics such as signal-to-noise ratio (SNR), speech intelligibility, and sound quality. Based on the validation results, adjust the network architecture or training parameters to improve model performance.
[0184] Evaluate the model's actual performance using unseen test set data. The test data should include audio from various karaoke scenarios to test the model's generalization ability.
[0185] The test results evaluate the overall performance of the model. The neural network model that passes the test can handle various noise and sound quality problems in practical applications.
[0186] Integrate the trained model into karaoke applications or other speech processing systems. Ensure that the integration process can handle high-sampling-rate audio data and perform real-time processing. Furthermore, optimize performance in practical applications, such as accelerating the inference process and reducing computational resource consumption.
[0187] In this embodiment, the singing audio to be processed in the karaoke scene is input to a multi-subband neural network model for noise reduction, and the enhanced audio signal, i.e., the third singing audio, is output, ensuring clear sound quality and good noise suppression effect.
[0188] To further enhance sound quality, the output audio signal can be further adjusted as needed, such as pitch correction, dynamic range adjustment, and / or, to address potential audio boundary issues and ensure audio coherence and naturalness.
[0189] In another possible implementation, the sound enhancement process, which involves first passing the sound through an echo cancellation module to eliminate background music and then through a noise reduction module to eliminate sudden and continuous noise, may include: processing the background music and environmental noise of the singing audio to be processed based on the Speak-X module to obtain a first singing audio; extracting the human voice from the first singing audio using a neural network model to obtain a second singing audio; eliminating background noise in the second singing audio using a multi-subband neural network model to obtain a fourth singing audio; and adjusting the pitch information based on the fourth singing audio to obtain the tuned audio.
[0190] Furthermore, in the process of audio editing, pitch shifting and tempo shifting are two crucial technical means, which play an indispensable role in improving audio quality, meeting specific musical needs, and creating unique auditory effects.
[0191] In some embodiments, pitch and speed shifting can be achieved in various ways: such as neural network vocoders (two-level neural network vocoders, single-layer GAN (Generative Adversarial Networks) vocoders, GPT (Generative Pre-trained Transformer) structure sound generators, Diffusion structure sound generators, and other neural network structures) and traditional acoustic techniques (Tdpsola algorithm, sola, wpsola). Furthermore, the available computing power varies depending on the service deployment. Traditional acoustic techniques such as the Tdpsola algorithm use traditional acoustic techniques for pitch and speed shifting, offering fast processing speed and low runtime environment requirements (it can run on servers with only CPUs), making it suitable for most pitch correction algorithm applications and deployments. The two-layer neural network vocoder uses a two-layer GAN network structure. The first layer performs speed and pitch shifting, and the second layer produces high-sampling-rate audio. It requires additional GPU (Graphics Processing Unit) resources on the server to run under the condition of meeting the real-time design requirements of the algorithm. It has better voice preservation effect, higher sound quality, and the ability to cope with situations can be steadily improved as the training data is expanded.
[0192] In one possible scenario, when the electronic device has sufficient computing resources to meet the computational requirements of a neural network vocoder, a neural network vocoder is used for speed and pitch shifting; otherwise, traditional acoustic techniques are used for speed and pitch shifting. In another scenario, the user specifies the speed and pitch shifting method.
[0193] One possible implementation of this application embodiment, taking the Tdpsola algorithm as an example, will be further illustrated.
[0194] The analysis of pitch period in the processed singing audio is called pitch period analysis. Pitch period refers to the periodic part in the speech signal, which usually corresponds to a sound period. Common methods include autocorrelation function (ACF) or periodogram-based techniques.
[0195] By analyzing the periodicity of the signal, the position and length of the pitch period can be determined;
[0196] The singing audio to be processed is divided into multiple frames according to pitch period, and these frames are overlapped on the time axis for subsequent overlap addition processing.
[0197] Pitch can be changed by adjusting the length and overlap of frames. For example, increasing the frame length will lower the pitch, while decreasing the frame length will raise the pitch. The playback speed of the audio signal can be changed by adjusting the frame sampling frequency. Time scaling can affect pitch and speech rate. During time scaling or pitch adjustment, the frame length and overlap ratio can be adjusted to achieve the desired effect.
[0198] Overlapping addition is applied to the processed frames to synthesize the final audio signal, ensuring a smooth transition of the audio signal and reducing artifacts caused by frame processing.
[0199] One possible implementation of the embodiments of this application, taking the two-level neural network vocoder method as an example, will be further illustrated as follows:
[0200] Obtain a two-level neural network vocoder, which includes a first GAN network structure and a second GAN network structure.
[0201] Based on the singing audio to be processed, the first GAN network structure is used to perform speed and pitch shifting to obtain the initial audio.
[0202] Based on the initial audio, the sampling rate is increased using a second GAN network structure to obtain high-quality singing audio to be processed.
[0203] The two-level neural network vocoder is an improved Hifigan, where:
[0204] Acquire a large-scale audio dataset, ensuring audio quality and diversity. The audio duration should be sufficiently long, covering different speech, music, or sound effects. The audio data needs to be processed into the model input format, i.e., converting the raw audio signal into a Mel spectrogram. This can be done using librosa or other audio processing libraries. Further data preprocessing can be performed, including normalization, segmentation, and noise reduction, to ensure data quality and consistency.
[0205] The audio dataset is divided into training, validation, and test sets to ensure that the model has good generalization ability on different types of data.
[0206] Select HiFi-GAN model configuration: Configure the model's hyperparameters, such as learning rate, batch size, optimizer, and loss function, based on the characteristics of the dataset and computing resources.
[0207] Large-scale training is performed using GPUs / TPUs. During training, the model learns how to generate high-quality audio from Mel spectrograms. The model performance is evaluated periodically using a validation set to prevent overfitting. The quality of the model can be monitored through loss curves and audio examples.
[0208] Model optimization and fine-tuning: Adjusting the model's hyperparameters or network structure based on training results to improve audio quality or training speed. Model fine-tuning uses specific audio datasets to further optimize the model's performance in specific scenarios.
[0209] Save the trained HiFi-GAN model in the format required for deployment (such as a PyTorch model file). Export the model in ONNX format for inference deployment on different platforms.
[0210] As can be seen, in this embodiment, a two-level neural network vocoder including a first GAN network structure and a second GAN network structure is introduced. The first GAN network structure is responsible for speed and pitch shifting of the singing audio to be processed, which can flexibly adjust the playback speed and pitch of the audio. The second GAN network structure increases the sampling rate of the initial audio, which effectively improves the clarity and quality of the audio, reduces the sound quality loss caused by insufficient sampling rate, and makes the audio processing result more excellent.
[0211] Furthermore, the granularity of information generated by speed and pitch changes is at the note level, which can lead to abrupt changes and unnatural transitions when used directly. Therefore, this application's embodiment designs a smooth transition module. After obtaining the basic pitch correction information, it progressively degrades the granularity to obtain smaller operation units that correspond to the reference information (MIDI). During the decomposition process, geometric curves are used to maintain the audio frequency, ensuring that it does not lose its original continuity with each level of decomposition. Finally, the processing unit is decomposed to the waveform level, reducing the time granularity to below 5ms. At this fine granularity, the transitions between different phonemes are natural, and the original vocal tone and vocal characteristics are preserved. Specifically, one possible implementation of this application's embodiment, after determining the adjusted pitch information based on the original pitch template and pitch information, further includes:
[0212] Based on the individual words whose pitch information is adjusted, determine the first musical score information corresponding to each word from the original MIDI file;
[0213] Based on each phoneme of the single word whose pitch information is adjusted, determine the second musical score information corresponding to each phoneme from the first musical score information;
[0214] Based on each waveform of the phoneme, determine the corresponding third musical score information from the second musical score information;
[0215] Based on the third score information corresponding to all waveforms, the final pitch adjustment information is determined to achieve smooth processing of the pitch adjustment information.
[0216] Referring to Figure 4, which is a flowchart illustrating a smoothing process provided in an embodiment of this application, it shows the gradual decomposition of a human voice signal: from words to phonemes to waveform signals. Here, full-text ASR (Automatic Speech Recognition); MIDI information is general musical notation used to represent musical scores; Note is pitch information extracted from an audio segment; Pitch constitutes a more subdivided unit of Note, where pitch is continuous and notes are discrete. In this gradual decomposition process, smaller units are used step-by-step to align with a general musical score, allowing for audio adjustments at a finer granularity.
[0217] As can be seen, in this embodiment, based on the single word of the pitch information, the corresponding first score information is extracted from the MIDI file of the original song, which provides an accurate score reference for subsequent pitch adjustment; then, at the level of each phoneme, the second score information corresponding to each phoneme is determined from the first score information. This detailed processing at the phoneme level ensures the accuracy and consistency of pitch adjustment; then, based on each waveform of the phoneme, the third score information corresponding to each waveform is determined from the second score information, which fully considers the changes in the waveform within the phoneme, making the pitch adjustment smoother and more natural.
[0218] Based on any of the above embodiments, referring to Figure 5, Figure 5 is a schematic diagram of a specific audio correction scheme provided by an embodiment of this application. It is a comprehensive scheme that integrates three parts: a pre-processed speech recognition and vocal pitch recognition module, a streaming parallel noise reduction and filtering module, and a correction processing module that combines multiple information sources to make correction decisions and produce corrected audio. The entire process is a 32K high sampling rate scheme, but it can also be 28K or 22.5K.
[0219] Based on the karaoke scenario, the audio correction algorithm used in this embodiment adds a speech recognition module that enhances boundary capabilities, a smooth transition module for speed and pitch shifting, an adaptive audio correction template reconstruction module, and a speed and pitch shifting module that processes waveform segments as units, and readjusts the data flow direction compared to related audio correction technologies.
[0220] The audio editing method includes:
[0221] Speech recognition steps: Perform speech recognition on the singing audio to be processed to obtain audio description information.
[0222] Pitch recognition steps: Perform pitch recognition on the singing audio to be processed to obtain initial pitch information;
[0223] Noise reduction filtering steps: Perform noise reduction filtering on the singing audio to be processed to obtain the noise-reduced singing audio;
[0224] The steps for obtaining related information are as follows: Obtain the previous pitch information and / or the next pitch information corresponding to the singing audio to be processed, i.e., historical and future information;
[0225] Steps to determine the correspondence between audio and reference: Based on the audio description information, determine the target lyrics in the reference lyrics that correspond to the audio description information;
[0226] The pitch adjustment steps are as follows: First, determine the original song pitch template corresponding to the singing audio to be processed based on the target lyrics. Second, filter invalid pitch information from the initial pitch information based on the boundary time information corresponding to the phonemes in the audio description information. Third, obtain the preceding pitch information corresponding to the preceding singing audio and the following pitch information corresponding to the following singing audio. Fourth, adjust the overall pitch of the filtered pitch information based on the preceding and / or following pitch information to obtain the pitch information. Fifth, correct the original song pitch template corresponding to the singing audio to be processed based on the pitch information to obtain a reference original song pitch template. Sixth, determine the adjusted pitch information based on the reference original song pitch template and the pitch information.
[0227] The smoothing step: smooths out the pitch information.
[0228] Speed and pitch shifting steps: The noise-reduced singing audio will be processed by speed and pitch shifting to obtain a high-quality singing audio to be processed. Based on the pitch adjustment information, the high-quality singing audio to be processed will be pitch-processed to obtain the audio after pitch correction.
[0229] As can be seen, in this embodiment, the accurate pronunciation information and reference information generated by the speech recognition module work together to generate a new song template with a permissible speed and pitch range in the adaptive pitch correction template reconstruction module, forming basic speed and pitch information. The accurate pitch information produced by the pitch recognition module helps to improve the speed and pitch information. The improved speed and pitch information is then processed by the smoothing module to generate a more reliable small-granular audio operation information sequence for use by the speed and pitch module. The noise reduction and filtering module produces audio data as the operation material for the speed and pitch algorithm in subsequent processes. After processing in the pre-processing stage, the speed and pitch are processed in parallel using the audio operation information sequence, completing the audio processing with an effectiveness not exceeding 0.03 RTF.
[0230] Furthermore, the audio editing technology used in this application embodiment was compared with the audio editing effect of the WeSing app using MOS scores, as shown in Figure 6. Through the intelligent audio editing technology designed in this application embodiment, different audios receive different processing solutions, and all recorded audios are processed uniformly. The processed audios perform excellently in terms of rhythm and pitch control, reducing the usage threshold for users in karaoke scenarios, and enabling most novice users to produce singing works that conform to musical aesthetics.
[0231] The following describes an audio correction device provided in an embodiment of this application. The device described below corresponds to the method described above. The device in this embodiment is installed in an electronic device. Referring to FIG7, FIG7 is a structural block diagram of an audio correction device according to one embodiment of this application. The audio correction device 700 includes:
[0232] The audio acquisition module 710 is used to acquire the singing audio to be processed.
[0233] The recognition module 720 is used to perform speech recognition on the singing audio to be processed to obtain audio description information; and to perform pitch recognition on the singing audio to be processed to obtain initial pitch information.
[0234] The boundary correction module 730 is used to correct the boundary information of the initial pitch information according to the audio description information to obtain the pitch information.
[0235] The dynamic template generation module 740 is used to determine the original music pitch template corresponding to the singing audio to be processed based on the audio description information, and to correct the original music pitch template corresponding to the singing audio to be processed based on the pitch information to obtain the reference original music pitch template.
[0236] The pitch correction module 750 is used to determine the pitch information to be adjusted based on the original song's pitch template and pitch information, and to perform pitch correction on the singing audio to be processed based on the adjusted pitch information to obtain the pitch-corrected audio.
[0237] In one possible implementation, the identification module 720 is specifically used for:
[0238] Extract the acoustic features of the singing audio to be processed;
[0239] Based on the acoustic features and the decoding map, the acoustic features are decoded to obtain the phoneme sequence. The decoding map is generated based on the phoneme-level acoustic model, the lyrics-based language model, and the phoneme-lyrics conversion table.
[0240] The phoneme sequence is mapped to the lyrics text to obtain audio description information.
[0241] In one feasible approach, the audio description information includes: lyrics information and phoneme boundary time information;
[0242] Boundary correction module 730 is specifically used for:
[0243] Based on the boundary time information corresponding to the phonemes in the audio description information, filter out invalid pitch information in the initial pitch information;
[0244] Obtain the pitch information of the preceding and following audio segments of the singing audio to be processed;
[0245] Based on the preceding and / or following pitch information, the overall pitch of the filtered pitch information is adjusted.
[0246] In one feasible implementation, the pitch correction module 750 is specifically used for:
[0247] The noise reduction filter is applied to the singing audio to be processed based on the first method and / or the second method to obtain the filtered audio.
[0248] The filtered audio is corrected based on the pitch information to obtain the corrected audio.
[0249] The first approach involves removing background music and environmental noise from the audio using the Speak-X module, and then extracting the human voice using a neural network model. The second approach is based on a multi-subband neural network model to eliminate background noise in the audio.
[0250] In one possible implementation, the dynamic template generation module 740 is specifically used for:
[0251] The average pitch and individual pitch of the singing audio to be processed are determined based on pitch information.
[0252] If the difference between the average pitch value and the average target pitch value corresponding to the original pitch template exceeds the first preset pitch threshold, the pitch of the original pitch template is adjusted as a whole to obtain the first original pitch template. The difference between the average pitch value and the average target pitch value corresponding to the first original pitch template does not exceed the first preset pitch threshold.
[0253] If the difference between the pitch of a single note and the pitch of the target note in the first original pitch template exceeds the second preset pitch threshold, then the first pitch set is determined, wherein the target note corresponds to the single note, and the first pitch set is the set of pitches that differ from the target note by a preset pitch.
[0254] If there exists a pitch in the first pitch set whose pitch difference from that of a single pitch is within the second preset pitch threshold, then the pitch difference from that of a single pitch within the second preset pitch threshold will be used as the dynamic adjustment target for the target pitch.
[0255] If there is no pitch in the first pitch set that has a pitch difference of the second preset pitch threshold with respect to a single pitch, then a second pitch set is determined. The second pitch set is a set of pitches within the key of the original piece. From the second pitch set, a dynamic adjustment target that meets the target conditions is selected. The target conditions include: a pitch within the key that is closest to the target pitch and has a pitch difference of the second preset pitch threshold with respect to a single pitch.
[0256] After determining the dynamic adjustment targets for all individual notes, the first original pitch template is adjusted according to the dynamic adjustment targets for all individual notes to obtain the reference original pitch template.
[0257] In one feasible approach, the dynamic template generation module 740 is further configured to: process the reference target signal whose time-domain variation coefficient in the reference original pitch template exceeds a preset multiple by inserting white noise, while ensuring that the starting position remains aligned, to obtain an aligned reference original pitch template;
[0258] If the target note in the first original pitch template corresponding to a single note cannot be determined, the pitch of the first original pitch template is adjusted as a whole to the pitch within the key, and the starting point of the sound is adjusted to the preset time position.
[0259] In one possible implementation, it also includes: a speed and pitch shifting module, used for:
[0260] Obtain a two-level neural network vocoder, which includes a first GAN network structure and a second GAN network structure.
[0261] Based on the singing audio to be processed, the first GAN network structure is used to perform speed and pitch shifting to obtain the initial audio.
[0262] Based on the initial audio, the sampling rate is increased using a second GAN network structure to obtain high-quality singing audio to be processed.
[0263] In one possible implementation, a smoothing module is also included, for:
[0264] Based on the individual words whose pitch information is adjusted, determine the first musical score information corresponding to each word from the original MIDI file;
[0265] Based on each phoneme of the single word whose pitch information is adjusted, determine the second musical score information corresponding to each phoneme from the first musical score information;
[0266] Based on each waveform of the phoneme, determine the corresponding third musical score information from the second musical score information;
[0267] Based on the third score information corresponding to all waveforms, the final pitch adjustment information is determined to achieve smooth processing of the pitch adjustment information.
[0268] This application provides an electronic device, as shown in FIG8. The electronic device 800 shown in FIG8 includes a processor 801 and a memory 803. The processor 801 and the memory 803 are connected, for example, via a bus 802. Optionally, the electronic device 800 may further include a transceiver 804. It should be noted that in practical applications, the transceiver 804 is not limited to one type, and the structure of this electronic device 800 does not constitute a limitation on the embodiments of this application.
[0269] Processor 801 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 801 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0270] Bus 802 may include a pathway for transmitting information between the aforementioned components. Bus 802 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 802 can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in Figure 8, but this does not indicate that there is only one bus or one type of bus.
[0271] The memory 803 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.
[0272] The memory 803 stores application code that executes the scheme of this application, and its execution is controlled by the processor 801. The processor 801 executes the application code stored in the memory 803 to implement the content shown in the foregoing method embodiments.
[0273] The electronic device shown in Figure 8 is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0274] This application provides a computer-readable storage medium storing a computer program that, when run on a computer, enables the computer to execute the corresponding content in the aforementioned method embodiments.
[0275] This application provides a computer program product, including a computer program that, when executed by a processor, implements the corresponding content in the aforementioned method embodiments.
[0276] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0277] The above are only some embodiments of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. An audio sound correction method, characterized in that, include: Obtain the singing audio to be processed; The singing audio to be processed is subjected to speech recognition to obtain audio description information; The pitch of the singing audio to be processed is then identified to obtain initial pitch information. The initial pitch information is corrected for boundary information based on the audio description information to obtain the pitch information; The original song pitch template corresponding to the singing audio to be processed is determined based on the audio description information, and the original song pitch template corresponding to the singing audio to be processed is corrected based on the pitch information to obtain the reference original song pitch template. Based on the original pitch template and the pitch information, the pitch information to be adjusted is determined, and the pitch of the singing audio to be processed is corrected based on the pitch information to obtain the corrected audio.
2. The method according to claim 1, characterized in that, The step of performing speech recognition on the singing audio to be processed to obtain audio description information includes: extracting the acoustic features of the singing audio to be processed; Based on the acoustic features and the decoding diagram, the acoustic features are decoded to obtain a phoneme sequence, wherein the decoding diagram is generated based on a phoneme-level acoustic model, a lyrics-based language model, and a phoneme-lyrics conversion table; The phoneme sequence is mapped to the lyrics text to obtain audio description information.
3. The method according to claim 2, characterized in that, The audio description information includes: lyrics information and phoneme boundary time information; The step of correcting the initial pitch information based on the audio description information to obtain the pitch information includes: Based on the boundary time information corresponding to the phonemes in the audio description information, filter out invalid pitch information in the initial pitch information; Obtain the pitch information of the preceding and following audio segments of the singing audio to be processed; Based on the preceding and / or following pitch information, the overall pitch of the filtered pitch information is adjusted.
4. The method according to claim 1, characterized in that, The process of adjusting the pitch of the singing audio to be processed based on the adjusted pitch information to obtain the adjusted audio includes: The noise reduction filter is applied to the singing audio to be processed based on the first method and / or the second method to obtain the filtered audio. Based on the adjusted pitch information, the filtered audio is corrected to obtain the corrected audio. The first approach involves removing background music and environmental noise from the audio using the Speak-X module, and then extracting the human voice using a neural network model. The second approach is based on a multi-subband neural network model to eliminate background noise in the audio.
5. The method according to any one of claims 1 to 4, characterized in that, The step of correcting the original pitch template corresponding to the singing audio to be processed based on the pitch information to obtain a reference original pitch template includes: Based on the pitch information, determine the average pitch and the pitch of each individual note in the singing audio to be processed; If the difference between the average pitch value and the average target pitch value corresponding to the original pitch template exceeds the first preset pitch threshold, the pitch of the original pitch template is adjusted as a whole to obtain the first original pitch template, and the difference between the average pitch value and the average target pitch value corresponding to the first original pitch template does not exceed the first preset pitch threshold. If the difference between the pitch of the individual note and the pitch of the target note in the first original pitch template exceeds a second preset pitch threshold, then a first pitch set is determined, wherein the target note corresponds to the individual note, and the first pitch set is a set of pitches that differ from the target note by a preset pitch. If there exists a pitch in the first pitch set whose pitch difference from the individual pitch is within a second preset pitch threshold, then the pitch difference from the individual pitch within the second preset pitch threshold will be used as the dynamic adjustment target for the target pitch. If there is no pitch in the first pitch set that has a pitch difference of the second preset pitch threshold with respect to the individual pitch, then a second pitch set is determined, which is a set of pitches within the key of the original song; a dynamic adjustment target that meets the target conditions is selected from the second pitch set, the target conditions including: a pitch within the key that is closest to the target pitch and has a pitch difference of the second preset pitch threshold with respect to the individual pitch; After determining the dynamic adjustment targets for all individual notes, the first original pitch template is adjusted according to the dynamic adjustment targets for all individual notes to obtain the reference original pitch template.
6. The method according to claim 5, characterized in that, After determining the dynamic adjustment targets for all individual notes, the first original pitch template is adjusted according to the dynamic adjustment targets for all individual notes to obtain the reference original pitch template, and then at least one of the following is also included: The reference target signal whose time domain variation coefficient in the reference original pitch template exceeds a preset multiple is processed by white noise insertion, while ensuring that the starting position remains aligned, to obtain the aligned reference original pitch template; If the target note in the first original pitch template corresponding to a single note cannot be determined, the pitch of the first original pitch template is adjusted as a whole to a pitch within the key, and the starting point of the sound is adjusted to a preset time position.
7. The method according to claim 1, characterized in that, After determining the adjusted pitch information based on the reference original pitch template and the pitch information, the process further includes: A two-level neural network vocoder is obtained, which includes a first GAN network structure and a second GAN network structure. Based on the singing audio to be processed, the first GAN network structure is used to perform speed and pitch shifting processing to obtain an initial audio. Based on the initial audio, the second GAN network structure is used to increase the sampling rate to obtain a high-quality singing audio to be processed.
8. The method according to claim 7, characterized in that, After determining the adjusted pitch information based on the reference original pitch template and the pitch information, the process further includes: Based on the individual words of the pitch adjustment information, the first musical score information corresponding to each word is determined from the MIDI file of the original song; based on each phoneme of the individual words of the pitch adjustment information, the second musical score information corresponding to each phoneme is determined from the first musical score information; based on each waveform of the phoneme, the third musical score information corresponding to each waveform is determined from the second musical score information; based on the third musical score information corresponding to all waveforms, the final pitch adjustment information is determined to achieve smooth processing of the pitch adjustment information.
9. An audio tone correction device, characterized in that, include: The audio acquisition module is used to acquire the singing audio to be processed; The recognition module is used to perform speech recognition on the singing audio to be processed to obtain audio description information; The pitch of the singing audio to be processed is then identified to obtain initial pitch information. The boundary correction module is used to correct the initial pitch information based on the audio description information to obtain the pitch information. The dynamic template generation module is used to determine the original song pitch template corresponding to the singing audio to be processed based on the audio description information, and to correct the original song pitch template corresponding to the singing audio to be processed based on the pitch information to obtain the reference original song pitch template. The pitch correction module is used to determine the adjusted pitch information based on the reference original pitch template and the pitch information, and to perform pitch correction on the singing audio to be processed based on the adjusted pitch information to obtain the pitch-corrected audio.
10. An electronic device, characterized in that, include: One or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, the one or more applications being configured to: perform the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method and apparatus for determining pitch deviation of audio content
CN108206026A
Automatic sound correction system and sound correction method
CN112447182A
Sound correction method and device, equipment and storage medium
CN113066462A
Sound correction method, computer equipment and computer readable storage medium
CN115101080A
Audio correction method and device and electronic equipment
CN119028323A