A method, device and medium for beautifying singing voice

By adopting a singing enhancement method based on singing scores and timbre inference models, the problems of correction bias and timbre misjudgment caused by template dependence are solved, and highly accurate singing enhancement is achieved.

CN119479684BActive Publication Date: 2025-11-14TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411656754.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-19
Publication Date
2025-11-14
Estimated Expiration
2044-11-19

AI Technical Summary

Technical Problem

Existing vocal enhancement technologies rely on the accuracy of templates, which can lead to correction deviations and misjudgments of timbre spectrum characteristics, thus affecting the enhancement effect.

Method used

The system determines the region to be inferred based on the singing score, calculates the octave shift pitch, trims the lead vocal feature file, and uses a pre-trained timbre inference model to fuse the user's timbre vector, outputting the beautified audio.

Benefits of technology

It achieves vocal enhancement without relying on template accuracy, improves the accuracy of timbre and pitch, and retains the user's unique vocal enhancement effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119479684B_ABST
    Figure CN119479684B_ABST
Patent Text Reader

Abstract

This application discloses a method, device, and medium for enhancing vocal performance, relating to the field of vocal conversion technology. The method includes: upon receiving an enhancement service instruction for a target recorded song, determining a region to be inferred that matches the enhancement service instruction based on the singing score of the target recorded song; determining an octave offset pitch based on the vocal range distribution of the target recorded song, and calculating the target inferred pitch based on the octave offset pitch and the target user's singing pitch; trimming the lead vocal feature file based on the region to be inferred to obtain a target feature file; the lead vocal feature file is the feature file of the original song; inputting the target inferred pitch, the target feature file, and the target user's timbre vector into a target timbre inference model to output the enhanced audio; and sending the enhanced audio to the user. The vocal inference effect of this application does not rely on the accuracy of traditional pitch correction templates, ensuring the accuracy of vocal enhancement while preserving some of the user's own characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vocal conversion technology, and in particular to a method, device and medium for vocal enhancement. Background Technology

[0002] Currently, most vocal enhancement solutions are based on signal processing, and the enhancement effect of signal processing depends on the accuracy of the template. When the accuracy of the template is low, the signal processing process may encounter problems such as correction deviations and misjudgments of timbre spectrum characteristics, resulting in the inability to guarantee the vocal enhancement effect. Therefore, these problems urgently need to be solved by those skilled in the art. Summary of the Invention

[0003] In view of this, the purpose of this invention is to provide a method, device, and medium for enhancing singing voice, so that the singing voice inference effect does not depend on the accuracy of traditional pitch correction templates, and the accuracy of singing voice enhancement is guaranteed while retaining some of the user's own characteristics. The specific solution is as follows:

[0004] Firstly, this application discloses a method for enhancing singing voice, including:

[0005] When a beautification service instruction is received for a target recorded song, the region to be inferred in the target recorded song that matches the beautification service instruction is determined based on the singing score of the target recorded song; the target recorded song is obtained by the target user singing the target track;

[0006] Based on the vocal range distribution of the target recorded song, the octave offset pitch of the region to be inferred is determined, and the target inference pitch is calculated based on the octave offset pitch and the singing pitch of the target user.

[0007] The vocal feature file is cropped based on the region to be inferred to obtain the target feature file; wherein, the vocal feature file is the feature file of the original song of the target track;

[0008] The target inference pitch, the target feature file, and the timbre vector of the target user are input into the target timbre inference model, so as to output the target beautified audio through the target timbre inference model;

[0009] The audio of the target device is then beautified and sent to the user's device.

[0010] Optionally, the singing enhancement method further includes:

[0011] Obtain the user's dry audio file from the target user;

[0012] The user's dry audio file is subjected to noise reduction processing to obtain a noise-reduced user dry audio file;

[0013] Multiple songs that meet preset conditions are selected from the user's dry audio file after noise reduction;

[0014] A sample training set is constructed based on the multiple songs, and the timbre model to be trained is trained based on the sample training set to obtain a target timbre inference model whose training loss satisfies the loss condition.

[0015] Optionally, if the enhancement service instruction is a preview clip enhancement service instruction, then determining the region to be inferred in the target recorded song that matches the enhancement service instruction based on the singing score of the target recorded song includes:

[0016] If the singing score of each sung sentence of the target recorded song is not zero, then the target song segment is obtained based on the sung sentence with the lowest singing score.

[0017] The target song segment is cropped according to the first cropping strategy to obtain the first listening segment, and the time area corresponding to the first listening segment is determined as the region to be inferred. The first cropping strategy is obtained based on the lyrics timestamp information, the reserved time on the left side of the listening segment and the expected duration of the listening segment. The reserved time on the left side of the listening segment is a reserved time to prevent cropping deviation.

[0018] Optionally, if the enhancement service instruction is a preview clip enhancement service instruction, then determining the region to be inferred in the target recorded song that matches the enhancement service instruction based on the singing score of the target recorded song includes:

[0019] If the performance score of each singing sentence in the target recorded song is all zero, or the number of performance scores of the target recorded song is different from the length of the sentence list, then the preset duration segment of the target recorded song is determined as the target song segment; the preset duration segment includes the start timestamp to the target timestamp of the target recorded song, and the length of the sentence list is the number of sentences in the singing sentence;

[0020] The target song segment is cropped according to the second cropping strategy to obtain a second listening segment, and the time region corresponding to the second listening segment is determined as the region to be inferred; the second cropping strategy is obtained based on the lyrics timestamp information and the expected duration of the listening segment.

[0021] Optionally, if the enhancement service instruction is a complete song enhancement service instruction, then determining the region to be inferred in the target recorded song that matches the enhancement service instruction based on the singing score of the target recorded song includes:

[0022] The region to be inferred is determined based on whether all the singing sentences in the target recorded song fail to meet the standard; the singing score of any dimension of the singing score of the singing sentence that fails to meet the standard is lower than the corresponding scoring threshold, and the singing score includes multiple dimensions of scoring.

[0023] Optionally, determining the region to be inferred based on whether all sung phrases in the target recorded song fail to meet the standard includes:

[0024] If all the sung sentences in the target recorded song fail to meet the standard, then the time area corresponding to the target recorded song is determined as the area to be inferred.

[0025] If any part of the singing in the target recorded song does not meet the standard, then the time area corresponding to the lead vocal feature file is determined as the area to be inferred.

[0026] If all the sung sentences in the target recorded song meet the standard, no processing is required.

[0027] Optionally, if the region to be inferred is the time region corresponding to the vocal feature file, then the step of inputting the target inferred pitch, the target feature file, and the timbre vector of the target user into the target timbre inference model, so as to output the target beautified audio through the target timbre inference model, includes:

[0028] The target inference pitch, the vocal feature file, and the timbre vector of the target user are input into the target timbre inference model to obtain the audio after the first inference.

[0029] The unqualified singing sentences in the target recorded song are identified as the first singing sentences, and the qualified singing sentences in the target recorded song are identified as the second singing sentences.

[0030] The third singing sentence corresponding to the first singing sentence is selected from the audio after the first reasoning, and the second singing sentence and the third singing sentence are concatenated to obtain the concatenated result;

[0031] The spliced ​​result is cropped according to the third cropping strategy and the target recorded song to obtain the cropped result, and the cropped result is directly determined as the target beautified audio.

[0032] Optionally, after cropping the spliced ​​result according to the third cropping strategy and the target recorded song to obtain the cropped result, the method further includes:

[0033] The updated vocal feature file is obtained based on the cropped result;

[0034] The target inferred pitch, the updated vocal feature file, and the timbre vector of the target user are input into the target timbre inference model to obtain the target beautified audio.

[0035] Optionally, sending the beautified audio of the target device to the user terminal includes:

[0036] If the enhancement service instruction is the audition segment enhancement service instruction, then the singing segment corresponding to the target enhanced audio is determined from the target recorded song, and the original singing segment and the target enhanced audio are sent to the user terminal; wherein, the singing segment is the segment corresponding to the region to be inferred;

[0037] If the enhancement service instruction is the complete song enhancement service instruction, then the target enhanced audio will be sent to the user terminal.

[0038] Secondly, this application discloses a singing voice enhancement device, comprising:

[0039] The region to be inferred module is used to determine the region to be inferred in the target recorded song that matches the beautification service instruction based on the singing score of the target recorded song when a beautification service instruction is received for the target recorded song; the target recorded song is obtained by the target user singing the target track;

[0040] The inference pitch determination module is used to determine the octave offset pitch of the region to be inferred based on the vocal range distribution of the target recorded song, and to calculate the target inference pitch based on the octave offset pitch and the singing pitch of the target user.

[0041] The target feature file determination module is used to trim the vocal lead feature file based on the region to be inferred to obtain the target feature file; wherein, the vocal lead feature file is the feature file of the original song of the target track;

[0042] The audio enhancement module is used to input the target inferred pitch, the target feature file, and the timbre vector of the target user into the target timbre inference model, so as to output the target enhanced audio through the target timbre inference model;

[0043] The beautified audio delivery module is used to deliver the target beautified audio to the user terminal.

[0044] Thirdly, this application discloses an electronic device, including:

[0045] Memory, used to store computer programs;

[0046] A processor is used to execute the computer program to implement the aforementioned disclosed method for enhancing singing voice.

[0047] Fourthly, this application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned voice enhancement method.

[0048] Fifthly, this application discloses a computer program product, including a computer program / instruction that, when executed by a processor, implements the aforementioned singing enhancement method.

[0049] As can be seen, this application proposes a singing enhancement method, including: when receiving an enhancement service instruction for a target recorded song, determining the region to be inferred in the target recorded song that matches the enhancement service instruction based on the singing score of the target recorded song; the target recorded song is obtained by a target user singing a target track; determining the octave offset pitch of the region to be inferred based on the vocal range distribution of the target recorded song, and calculating the target inferred pitch based on the octave offset pitch and the singing pitch of the target user; cropping the vocal feature file based on the region to be inferred to obtain a target feature file; wherein, the vocal feature file is the feature file of the original song of the target track; inputting the target inferred pitch, the target feature file, and the timbre vector of the target user into a target timbre inference model, so as to output the target enhanced audio through the target timbre inference model; and sending the target enhanced audio to the user terminal.

[0050] Beneficial Effects: For different enhancement service instructions, this application determines different regions to be inferred based on the singing score of the target recorded song, and trims the feature file of the original song of the target track according to the region to be inferred to obtain a target feature file. The target inference pitch, the target feature file, and the timbre vector of the target user are then input into a pre-trained target timbre inference model, which outputs the enhanced audio. This achieves a fusion of timbre, pitch, and original vocal characteristics, resulting in a final enhanced audio output with good timbre performance and pitch accuracy. Furthermore, for the target inference pitch used during inference, this application calculates the octave difference between the target recorded song and the original song and performs corresponding offset synthesis to make the enhanced audio more closely match the user's original singing habits. This ensures that the final enhanced audio retains the charm of the original song while highlighting the target user's singing characteristics and providing targeted enhancement. In summary, the audio enhancement method of this application does not rely on the accuracy of the template, thus solving the problems of correction deviation and misjudgment of timbre spectrum characteristics caused by inaccurate templates in traditional technologies. Attached Figure Description

[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0052] Figure 1 This is a flowchart of a singing voice enhancement method disclosed in this application;

[0053] Figure 2 This application discloses a flowchart for enhancing audio-visual segments.

[0054] Figure 3 This application discloses a complete song enhancement flowchart;

[0055] Figure 4 This application discloses a specific flowchart for voice enhancement.

[0056] Figure 5 This is an interactive schematic flowchart disclosed in this application;

[0057] Figure 6 This application discloses a specific method for enhancing singing voice as shown in the flowchart.

[0058] Figure 7 This application discloses a flowchart of a complete sentence replacement strategy in song enhancement.

[0059] Figure 8 This is a schematic diagram of the structure of a singing enhancement device disclosed in this application;

[0060] Figure 9 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0061] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0062] Currently, most vocal enhancement solutions are based on signal processing, and the enhancement effect of signal processing-based vocal enhancement depends on the accuracy of the template. When the accuracy of the template is low, problems such as correction deviation and misjudgment of timbre spectrum characteristics may occur during signal processing, resulting in the enhancement effect not being guaranteed.

[0063] Therefore, this application proposes a vocal enhancement scheme that makes the vocal enhancement effect independent of the accuracy of traditional pitch correction templates and ensures the accuracy of vocal enhancement.

[0064] This application discloses a method for beautifying singing voices; see [link to relevant documentation]. Figure 1 As shown, the method includes:

[0065] Step S11: When a beautification service instruction for a target recorded song is received, the region to be inferred in the target recorded song that matches the beautification service instruction is determined based on the singing score of the target recorded song; the target recorded song is obtained by the target user singing the target track.

[0066] In this embodiment, the user-selected enhancement service instruction is obtained through the singing enhancement service settings interface on the interactive interface. Based on the user-selected enhancement service instruction and the singing score of the target recorded song, the region to be inferred in the target recorded song that matches the enhancement service instruction is determined. It should be noted that the singing score includes multiple dimensions of scoring, which constitute a dimension scoring array. The dimension scoring array includes, but is not limited to, breath control scoring, pitch scoring, and rhythm scoring.

[0067] The following explains how to obtain the performance score of the target recorded song:

[0068] First, analyze the timestamps of the song sentences to obtain... ;in, Indicates the line number. This is the timestamp of the beginning of that lyric. This is the timestamp indicating the end of the lyric. Further, it checks if the number of performance scores (the length of the dimension scoring array) matches the length of the sentence list (the number of sentences in the song). For example, if the sentence list length is 20, it checks if the number of performance scores is 20. If not, it selects the audition segment-based determination strategy, which will be described later. If it matches, it stores the dimension scores corresponding to the sentences, in the following format: ;in, , , Dimension express Dimensional scoring. It's understandable that since the target user might only sing a few lines of lyrics, this embodiment needs to retain the sentence score for the target user's performance and filter out invalid sentence scores from other parts. Since the target user might start singing from the middle of a line of lyrics, it's necessary to handle abnormal situations at the left and right boundaries. First, the left boundary is processed: if the entire line of lyrics is to the left of the start time of the target user's performance segment, the line is discarded; if the entire line of lyrics is to the right of the start time, the line is retained, and its right boundary is determined in subsequent steps. If the line just crosses the left boundary, the start timestamp of the line is truncated based on the start time of the performance segment. Further, the right boundary is processed: if the entire line of lyrics is to the right of the end time of the target user's performance segment, the line is discarded; if the entire line of lyrics is to the left of the end time, it is directly retained. If the line just crosses the right boundary, it's necessary to determine if the effective performance rate threshold for the sung sentence has been reached. The threshold can be preset to A. If it doesn't reach A, it's directly discarded; otherwise, it's necessary to determine if the sentence score is non-zero. If the sentence score is non-zero, it's retained; otherwise, it's discarded. Among them, the effective singing rate of a sentence is the ratio of the time range of a user singing part of a certain lyric to the time range of the complete lyric.

[0069] The following explains how to determine the region to be reasoned about:

[0070] In the first implementation, the user requires a preview segment enhancement service. Based on this, if the performance scores of each sung phrase in the target recorded song are not zero, then the target song segment is obtained based on the sung phrase with the lowest performance score. This target song segment is then trimmed according to a first trimming strategy to obtain a first preview segment. The time region corresponding to the first preview segment is then determined as the region to be inferred. The first trimming strategy is based on lyrics timestamp information, the reserved time on the left side of the preview segment, and the expected duration of the preview segment. The reserved time on the left side of the preview segment is reserved to prevent trimming deviations. For example, if the target recorded song contains three sung phrases, it is determined whether the performance scores of these three phrases are not zero. It should be noted that a performance score not being zero means that multiple dimensions of the score are not all zero. For example, in the first sung phrase, the breath control score is not zero; in the second sung phrase, the pitch score is not zero; and in the third sung phrase, the rhythm score is not zero. These can be described as all three sung phrases having a non-zero performance score. Furthermore, if any singing sentence has a non-zero performance score and the scores for each dimension of the performance score are lower than the corresponding threshold, then the singing sentence is considered a substandard singing sentence. The singing sentence with the lowest performance score is selected from the substandard singing sentences and identified as the target song segment. Then, the target song segment is trimmed according to the first trimming strategy to obtain the first listening segment, and the time region corresponding to the first listening segment is identified as the region to be inferred.

[0071] The first cropping strategy is explained in detail below:

[0072] ;

[0073] ;

[0074] in, This is the timestamp indicating the start of the first audio clip. The timestamp of the start of the target song segment. Reserve time on the left side of the audition clip. The timestamp indicating the start of the user's performance. This is the end timestamp of the first audio clip. The expected duration of the audition clip. This is the end timestamp of the user's performance segment. For example, It can be 80 milliseconds. It can be 10 seconds.

[0075] In the second implementation, the user requires a preview segment enhancement service. Based on this, if the performance score of each sung line in the target recorded song is all zero, or if the number of performance scores for the target recorded song is different from the length of the sentence list, a preset duration segment of the target recorded song is determined as the target song segment. This segment is then trimmed according to a second trimming strategy to obtain a second preview segment. The time region corresponding to the second preview segment is then determined as the region to be inferred. The preset duration segment includes the start timestamp to the target timestamp of the target recorded song. The second trimming strategy is based on the lyrics timestamp information and the expected duration of the preview segment. The preset duration segment can be 10 seconds; for example, the first 10 seconds of the target recorded song can be directly determined as the target song segment, and this segment can be trimmed to obtain the second preview segment. In this embodiment, trimming the target song segment is to prevent the last word in the segment from being truncated. It should be noted that this implementation is a fallback strategy for determining the preview segment when the performance score of each sung line by the user is zero.

[0076] The second cropping strategy is explained in detail below:

[0077] ;

[0078] ;

[0079] The meanings of each part in the second trimming strategy are as described above and will not be repeated here.

[0080] In the third implementation, the user requires a complete song enhancement service. Based on this, the region to be inferred is determined according to whether all the vocal phrases in the target recorded song fail to meet the requirements. If all the vocal phrases in the target recorded song fail to meet the requirements, the time region corresponding to the target recorded song is determined as the region to be inferred; if only some vocal phrases in the target recorded song fail to meet the requirements, the time region corresponding to the vocal feature file is determined as the region to be inferred; if all the vocal phrases in the target recorded song meet the requirements, no processing is performed. In other words, if a user requires a complete song enhancement service, when all the singing sentences fail to meet the standards, inference is performed on each singing sentence. At this time, the time area corresponding to the target recorded song is determined as the inference area. When several singing sentences fail to meet the standards, the time area corresponding to the vocal guide feature file is determined as the inference area. Based on the inference results, the sentences that fail to meet the standards are replaced. Details about the replacement will be explained later. It should also be noted that the vocal guide feature file is the feature file of the original song of the target track, which is obtained in the following way: the original song is downloaded and the original dry vocals are separated. The reverb / harmony / other noise in the original dry vocals are removed, and PPG (phonetic posterior grams) & F0 (fundamental frequency of the speech signal) are extracted to obtain the vocal guide feature file. The vocal guide feature file is stored in the CFS (Cloud File Storage) shared disk for sharing among various services.

[0081] Step S12: Determine the octave offset pitch of the region to be inferred based on the vocal range distribution of the target recorded song, and calculate the target inference pitch based on the octave offset pitch and the singing pitch of the target user.

[0082] Considering that male singers often sing female-accompanied songs an octave lower, and female singers may sing male-accompanied songs an octave higher, this embodiment calculates the octave difference between the target recorded song and the original song to make the enhanced vocals closer to the user's original singing habits, and performs corresponding offset synthesis. Specifically, the vocal audio file of the target recorded song is denoised to obtain the denoised version of the song, and the pitch distribution of the denoised version is calculated to obtain the median pitch. An octave offset pitch is calculated based on the median pitch, and the target inferred pitch is calculated by adding the octave offset pitch to the user's singing pitch. The user's singing pitch can be obtained from the relevant recording information of the target recorded song.

[0083] ;

[0084] in, Represents the median of the vocal range.

[0085] Step S13: Based on the region to be inferred, the vocal feature file is cropped to obtain the target feature file; wherein, the vocal feature file is the feature file of the original song of the target track.

[0086] In this embodiment, the vocal feature file is cropped based on the region to be inferred. If the region to be inferred is 2s to 5s of the target recorded song, then 2s to 5s of the vocal feature file is determined as the target feature file.

[0087] Step S14: Input the target inferred pitch, the target feature file and the target user's timbre vector into the target timbre inference model, so as to output the target beautified audio through the target timbre inference model.

[0088] In this embodiment, the user's dry vocal file is downloaded from the streaming media system to a shared network disk for file storage, and the effective vocal duration of the user's dry vocal file is verified to reach a first threshold. The user's dry vocal file is a vocal recording of the target user. If the effective vocal duration of the user's dry vocal file reaches the first threshold, the user's dry vocal file is denoised to obtain a denoised user's dry vocal file. Multiple songs are selected from the denoised user's dry vocal file based on the user's headphone wearing status, instrument digital interface scoring, and the effective vocal duration. A sample training set is then constructed based on these multiple songs, and the timbre model to be trained is trained using the sample training set to obtain a target timbre inference model whose training loss satisfies the loss condition.

[0089] Specifically, this embodiment first collects the user's recorded dry audio PCM (Pulse Code Modulation) data via a mobile client or web page, with a sampling rate of 44100 or 48000, and the audio channel can be mono or stereo, which is not limited here. This data is then encoded into AAC (Advanced Audio Coding, a file compression format designed specifically for audio data) format and uploaded to a streaming media system for storage. The streaming media system includes, but is not limited to, COS (Cloud Object Storage) or other cloud storage. Simultaneously, corresponding recording information is recorded, such as whether only a portion of the audio is sung, the start timestamp of the sung portion (relative to the accompaniment), the end timestamp of the sung portion (relative to the accompaniment), headphone plug / unplug information, lyrics version information, multiple-dimensional scoring, MIDI (Musical Instrument Digital Interface) scoring, etc. Furthermore, the duration of the encoded file (i.e., the user's dry audio file) is verified to meet a second threshold. If so, it is downloaded to a CFS shared network disk to avoid repeated downloads and reduce development complexity. Furthermore, the effective vocal duration of the user's dry vocal files is verified to meet a first threshold, which can be 30 seconds. If it does not meet the threshold, the files are discarded. Noise reduction is then applied to the dry vocal files whose effective vocal duration meets the first threshold to minimize interference from noise caused by environmental factors, equipment malfunctions, or abnormal operations during recording. Multiple dry vocal tracks are selected from the noise-reduced user dry vocal files based on the user's headphone wearing status, instrument digital interface scores, and effective vocal duration. For example, tracks sung with headphones are given priority, as this effectively avoids environmental noise and accompaniment backslides, resulting in significantly better sound pickup than external speakers. Songs with higher instrument digital interface scores and longer effective vocal durations are also prioritized. Finally, a sample training set is constructed based on the selected dry vocal tracks, and the timbre model to be trained is trained using this sample training set to obtain a target timbre inference model whose training loss meets the loss condition. It should be noted that, in order to avoid the large number of dry audio samples obtained from the screening, which would lead to a large amount of computation and a long time consumption, this embodiment can preset a screening threshold. For example, the top 5 dry audio samples in the overall ranking can be used as part of the training set, and the other dry audio samples can be discarded.

[0090] The following explains how to train the target timbre inference model: (1) Divide the training set into a training subset and a validation subset according to a certain ratio. The training subset is used for the training process of the model, and the validation subset is used to evaluate the performance of the model during the training process so as to adjust the model parameters and training strategies in a timely manner. (2) Select the architecture of the timbre model to be trained, such as convolutional neural network (CNN), long short-term memory network (LSTM), gated recurrent unit (GRU), etc. (3) Define the loss function. Commonly used loss functions include cross-entropy loss function. (4) Set the training parameters. First, select an initial learning rate, such as 0.001. If it is found that the loss value decreases too slowly or oscillates during the model training process, the learning rate can be appropriately reduced or increased. (5) Set the number of training rounds and decide whether to increase or decrease the number of rounds based on the performance evaluation results on the validation subset. (6) Determine the batch size. Furthermore, the training subset data is fed into the timbre model to be trained in batches according to a set batch size. For each batch of data, the model's output value for that batch is calculated, the loss value for that batch is calculated using a defined loss function, and the backpropagation algorithm is used to differentiate the model's parameters based on the loss value, obtaining the gradient of each parameter. Finally, based on the obtained gradients and the set learning rate, the model's parameters are updated to reduce the loss value in the next training iteration, until training is completed for all batches of data in the entire training subset. After each round of training, the model's performance is evaluated using a validation subset. When the model's performance evaluation result on the validation subset meets the set loss condition, this model is considered the target timbre inference model whose training loss meets the loss condition.

[0091] In this embodiment, one user corresponds to one timbre inference model. Different users' voices have unique physiological characteristics. Training different timbre inference models for different users can better adapt to these physiological characteristics and achieve accurate beautification of each user's voice.

[0092] After obtaining the target timbre inference model, the target inferred pitch, target feature file and target user's timbre vector are input into the target timbre inference model to output the target beautified audio through the target timbre inference model. The user's timbre vector is the average value of multiple dry timbre vectors of the user in each dimension.

[0093] Step S14: Send the beautified audio of the target device to the user terminal.

[0094] In this embodiment, different content is sent to the user terminal according to the service instructions input by the user.

[0095] See Figure 2As shown, if the user inputs a preview segment enhancement service command, the original singing segment corresponding to the target enhanced audio is determined from the target recorded song. The original singing segment and the target enhanced audio (here referring to the enhanced audio obtained after reasoning the region to be reasoned based on the preview segment) are sent to the user terminal, allowing the user to intuitively compare the singing segment and the enhanced segment.

[0096] See Figure 3 As shown, if the user inputs a command for a paid full song enhancement service, the enhanced audio (here referring to the enhanced audio of the full song) will be sent to the user's device so that the user can enjoy a more professional and comprehensive song enhancement service. It should be noted that... Figure 3 The term "merging sentences" refers to merging adjacent sentences that do not meet the criteria in order to reduce damage to the original audio.

[0097] See Figure 4 As shown, after a user performs a song, they enter the mixing console, open the Super Audio Editing order panel, and choose to initiate a preview (i.e., preview clip enhancement) or pay to activate Super Audio Editing (i.e., full song enhancement). If they choose a preview, a preview clip enhancement task is initiated. The process generates the lowest quality original audio clip and the edited version of that clip, resulting in two audio files: the lowest quality original clip and the edited version. If they choose to pay to activate Super Audio Editing, a full song enhancement task is initiated. The process generates the enhanced full-length vocal track, resulting in one audio file: the complete vocal track. See also... Figure 5 As shown, Figure 5 The left side is the preview interface, and the right side is the complete enhancement interface. Users can click the "Before Processing" button to play the user's performance segment, and click the "After Processing" button to play the enhanced preview segment. If interested, users can click the bottom button to pay and then start the full enhancement process. Furthermore, this embodiment also supports visualizing the vocal enhancement process through technical language, allowing users to view the processing progress and method in real time, and play the enhanced full version upon completion.

[0098] In addition, the enhanced audio can be encrypted and uploaded to the streaming media system for distribution to users.

[0099] As can be seen, this application proposes a singing enhancement method, including: when receiving an enhancement service instruction for a target recorded song, determining the region to be inferred in the target recorded song that matches the enhancement service instruction based on the singing score of the target recorded song; the target recorded song is obtained by a target user singing a target track; determining the octave offset pitch of the region to be inferred based on the vocal range distribution of the target recorded song, and calculating the target inferred pitch based on the octave offset pitch and the singing pitch of the target user; cropping the vocal feature file based on the region to be inferred to obtain a target feature file; wherein, the vocal feature file is the feature file of the original song of the target track; inputting the target inferred pitch, the target feature file, and the timbre vector of the target user into a target timbre inference model, so as to output the target enhanced audio through the target timbre inference model; and sending the target enhanced audio to the user terminal.

[0100] Beneficial Effects: For different enhancement service instructions, this application determines different regions to be inferred based on the singing score of the target recorded song, and trims the feature file of the original song of the target track according to the region to be inferred to obtain a target feature file. The target inference pitch, the target feature file, and the timbre vector of the target user are then input into a pre-trained target timbre inference model, which outputs the enhanced audio. This achieves a fusion of timbre, pitch, and original vocal characteristics, resulting in a final enhanced audio output with good timbre performance and pitch accuracy. Furthermore, for the target inference pitch used during inference, this application calculates the octave difference between the target recorded song and the original song and performs corresponding offset synthesis to make the enhanced audio more closely match the user's original singing habits. This ensures that the final enhanced audio retains the charm of the original song while highlighting the target user's singing characteristics and providing targeted enhancement. In summary, the audio enhancement method of this application does not rely on the accuracy of the template, thus solving the problems of correction deviation and misjudgment of timbre spectrum characteristics caused by inaccurate templates in traditional technologies.

[0101] The following explains a scenario where a user needs a complete song enhancement service, but only a portion of the sung lines in the target recorded song are substandard, thus requiring the replacement of the substandard sung lines:

[0102] This application discloses a specific method for enhancing singing voice. Compared to the previous embodiment, this embodiment further explains and optimizes the technical solution. See also... Figure 6 As shown, it specifically includes:

[0103] Step S21: Input the target inference pitch, the vocal feature file, and the timbre vector of the target user into the target timbre inference model to obtain the audio after the first inference.

[0104] When certain singing phrases fail to meet the standard, the time region corresponding to the vocal guide feature file is determined as the region to be inferred. Based on the inference results, the failing phrases are replaced; this is the singing phrase replacement strategy. Specifically, the target inferred pitch, the vocal guide feature file, and the target user's timbre vector are first input into the target timbre inference model to obtain the audio after the first inference.

[0105] Step S22: Identify the unqualified singing sentences in the target recorded song as the first singing sentences, and identify the qualified singing sentences in the target recorded song as the second singing sentences.

[0106] See Figure 7 As shown, lines 1, 3, and 4 in the user's performance are qualified segments, which are the second performance sentences in this embodiment (the sentences after merging adjacent unqualified performance sentences). Line 2 in the user's performance is an unqualified segment, which is the first performance sentence in this embodiment.

[0107] Step S23: Select the third singing sentence corresponding to the first singing sentence from the audio after the first reasoning, and concatenate the second singing sentence and the third singing sentence to obtain the concatenated result.

[0108] See Figure 7 As shown, lines 0, 2, and 5 in the audio after the first inference correspond to the substandard segments in the target recorded song, which are the third singing sentences corresponding to the first singing sentence. It can be understood that lines 0, 2, and 5 in the audio after the first inference are the segments that have met the standards after inference. In this embodiment, the substandard segments in the target recorded song, i.e., the second singing sentence, are spliced ​​with the substandard segments, i.e., the third singing sentence, after inference, to obtain the spliced ​​result, which is the length of the complete song.

[0109] Step S24: Based on the third cropping strategy and the target recorded song, crop the spliced ​​result to obtain the cropped result, and directly determine the cropped result as the target beautified audio.

[0110] In this embodiment, since the target user may not sing the entire song, it is necessary to trim the spliced ​​result according to the region of the complete song where the target user's current singing segment is located, to obtain the trimmed result, which is the third trimming strategy. Furthermore, the trimmed result is the target beautified audio.

[0111] It should be noted that, to reduce the seamlessness caused by replacement splicing, this embodiment introduces secondary inference. Specifically, an updated vocal feature file is obtained based on the trimmed result, and the target inferred pitch, the updated vocal feature file, and the target user's timbre vector are input into the target timbre inference model to obtain the target enhanced audio. In this embodiment, the pitch used for secondary inference is 0, consistent with the pitch of the updated vocal feature file. Furthermore, the enhanced audio is encrypted and uploaded to the streaming media system for subsequent distribution to the user.

[0112] In summary, this embodiment beautifies and repairs sentences that do not meet the singing standards, while retaining sentences that meet the standards. This not only preserves some of the user's excellent singing skills, but also beautifies sentences of slightly lower quality to an effect comparable to the original singer. Furthermore, it makes the transitions more natural and smooth through secondary reasoning, greatly improving the accuracy of reasoning and the user experience.

[0113] As can be seen, this application proposes a singing analysis technology based on multi-dimensional scoring and a singing enhancement technology based on vocal transformation, which can be applied in multiple singing enhancement scenarios. In asynchronous recording scenarios, users can obtain the singing enhancement effect of a complete song. Furthermore, this embodiment covers the scenario of secondary creation of works. For historically accumulated works, through enhancement processing, combined with professional record-level mixing, video trimming, special effects rendering, and other technologies, high-quality works are created and distributed to users in consumer scenarios or offered as customized service packages.

[0114] Next, we will explain the scenario where a user needs a complete song reasoning service, but all the sung phrases in the target recorded song fail to meet the standards, thus requiring reasoning to be performed on all the non-compliant sung phrases:

[0115] When all singing phrases fail to meet the standard, inference needs to be performed on each singing phrase. In this case, the time region corresponding to the target recorded song is determined as the region to be inferred. Furthermore, considering that male singers often sing female accompaniment songs an octave lower, and female singers may sing male accompaniment songs an octave higher, this embodiment needs to calculate the octave difference between the target recorded song and the original song and perform corresponding offset synthesis to better match the beautified singing voice with the user's original singing habits. Specifically, the vocal audio file of the target recorded song is denoised to obtain the denoised singing song, and the vocal range distribution of the denoised singing song is calculated to obtain the median of the vocal range. Then, the octave offset pitch is calculated based on the median of the vocal range, and the target inferred pitch is calculated based on the octave offset pitch and the user's singing pitch. The specific calculation formula for the target inferred pitch is shown in the aforementioned disclosed embodiment and will not be elaborated here. Finally, the target inferred pitch, the target feature file, and the timbre vector of the target user are input into the target timbre inference model to output the target beautified audio through the target timbre inference model.

[0116] Accordingly, this application also discloses a singing voice enhancement device, see [link to relevant documentation]. Figure 8 As shown, the device includes:

[0117] The region to be inferred module 11 is used to determine the region to be inferred in the target recorded song that matches the beautification service instruction based on the singing score of the target recorded song when a beautification service instruction is received for the target recorded song; the target recorded song is obtained by the target user singing the target song;

[0118] The inference pitch determination module 12 is used to determine the octave offset pitch of the region to be inferred based on the vocal range distribution of the target recorded song, and to calculate the target inference pitch based on the octave offset pitch and the singing pitch of the target user.

[0119] The target feature file determination module 13 is used to trim the vocal feature file based on the region to be inferred to obtain the target feature file; wherein, the vocal feature file is the feature file of the original song of the target track;

[0120] Audio enhancement module 14 is used to input the target inference pitch, the target feature file and the timbre vector of the target user into the target timbre inference model, so as to output the target enhanced audio through the target timbre inference model;

[0121] The beautified audio delivery module 15 is used to deliver the target beautified audio to the user terminal.

[0122] For more detailed information on the working process of each of the above modules, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.

[0123] As can be seen, this application proposes a singing enhancement method, including: when receiving an enhancement service instruction for a target recorded song, determining the region to be inferred in the target recorded song that matches the enhancement service instruction based on the singing score of the target recorded song; the target recorded song is obtained by a target user singing a target track; determining the octave offset pitch of the region to be inferred based on the vocal range distribution of the target recorded song, and calculating the target inferred pitch based on the octave offset pitch and the singing pitch of the target user; cropping the vocal feature file based on the region to be inferred to obtain a target feature file; wherein, the vocal feature file is the feature file of the original song of the target track; inputting the target inferred pitch, the target feature file, and the timbre vector of the target user into a target timbre inference model, so as to output the target enhanced audio through the target timbre inference model; and sending the target enhanced audio to the user terminal.

[0124] Beneficial Effects: For different enhancement service instructions, this application determines different regions to be inferred based on the singing score of the target recorded song, and trims the feature file of the original song of the target track according to the region to be inferred to obtain a target feature file. The target inference pitch, the target feature file, and the timbre vector of the target user are then input into a pre-trained target timbre inference model, which outputs the enhanced audio. This achieves a fusion of timbre, pitch, and original vocal characteristics, resulting in a final enhanced audio output with good timbre performance and pitch accuracy. Furthermore, for the target inference pitch used during inference, this application calculates the octave difference between the target recorded song and the original song and performs corresponding offset synthesis to make the enhanced audio more closely match the user's original singing habits. This ensures that the final enhanced audio retains the charm of the original song while highlighting the target user's singing characteristics and providing targeted enhancement. In summary, the audio enhancement method of this application does not rely on the accuracy of the template, thus solving the problems of correction deviation and misjudgment of timbre spectrum characteristics caused by inaccurate templates in traditional technologies.

[0125] Furthermore, embodiments of this application also provide an electronic device. Figure 9 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0126] Figure 9This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a display screen 23, an input / output interface 24, a communication interface 25, a power supply 26, and a communication bus 27. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the vocal enhancement method disclosed in any of the foregoing embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be a computer.

[0127] In this embodiment, the power supply 26 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 25 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 24 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0128] Furthermore, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk, or optical disk, etc. The resources stored thereon may include computer programs 221, and the storage method may be temporary storage or permanent storage. The computer programs 221 may include, in addition to computer programs capable of performing the singing enhancement method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, computer programs capable of performing other specific tasks.

[0129] Furthermore, embodiments of this application also disclose a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned voice enhancement method.

[0130] For the specific steps of this method, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.

[0131] The various embodiments in this application are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. For the same or similar parts between the various embodiments, refer to each other. As for the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and relevant parts can be referred to in the method section.

[0132] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0133] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0134] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0135] The above provides a detailed description of a singing voice enhancement method, device, and storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for beautifying singing voice, characterized in that, include: When a beautification service instruction is received for a target recorded song, the region to be inferred in the target recorded song that matches the beautification service instruction is determined based on the singing score of the target recorded song. The target recorded song is obtained by the target user singing the target track; Based on the vocal range distribution of the target recorded song, the octave offset pitch of the region to be inferred is determined, and the target inference pitch is calculated based on the octave offset pitch and the singing pitch of the target user. The vocal feature file is cropped based on the region to be inferred to obtain the target feature file; wherein, the vocal feature file is the feature file of the original song of the target track; The target inference pitch, the target feature file, and the timbre vector of the target user are input into the target timbre inference model, so as to output the target beautified audio through the target timbre inference model; The enhanced audio of the target device is then sent to the user's device. When the enhancement service instruction is a preview segment enhancement service instruction, if the singing score of each singing sentence in the target recorded song is not zero, the target song segment is obtained based on the singing sentence with the lowest singing score, and the area to be inferred is obtained based on the target song segment; if the singing score of each singing sentence in the target recorded song is all zero or the number of singing scores of the target recorded song is different from the length of the sentence list, the preset duration segment of the target recorded song is determined as the target song segment, and the area to be inferred is obtained based on the target song segment. When the enhancement service instruction is a complete song enhancement service instruction, if all the singing sentences in the target recorded song fail to meet the standard, the time area corresponding to the target recorded song will be determined as the area to be inferred; if only some of the singing sentences in the target recorded song fail to meet the standard, the time area corresponding to the vocal feature file will be determined as the area to be inferred.

2. The method for enhancing singing voice according to claim 1, characterized in that, Also includes: Obtain the user's dry audio file from the target user; The user's dry audio file is subjected to noise reduction processing to obtain a noise-reduced user dry audio file; Multiple songs that meet preset conditions are selected from the user's dry audio file after noise reduction; A sample training set is constructed based on the multiple songs, and the timbre model to be trained is trained based on the sample training set to obtain a target timbre inference model whose training loss satisfies the loss condition.

3. The method for enhancing singing voice according to claim 1, characterized in that, The enhancement service instruction is a preview segment enhancement service instruction, and the singing scores of each sung phrase of the target recorded song are not zero. Therefore, obtaining the region to be inferred based on the target song segment includes: The target song segment is cropped according to the first cropping strategy to obtain the first listening segment, and the time area corresponding to the first listening segment is determined as the region to be inferred. The first cropping strategy is obtained based on the lyrics timestamp information, the reserved time on the left side of the listening segment and the expected duration of the listening segment. The reserved time on the left side of the listening segment is a reserved time to prevent cropping deviation.

4. The method for enhancing singing voice according to claim 1, characterized in that, The enhancement service instruction is a preview segment enhancement service instruction, and the singing scores of each sung phrase of the target recorded song are not zero. Therefore, obtaining the region to be inferred based on the target song segment includes: The target song segment is trimmed according to the second trimming strategy to obtain a second listening segment, and the time region corresponding to the second listening segment is determined as the region to be inferred; the second trimming strategy is obtained based on the lyrics timestamp information and the expected duration of the listening segment; the preset duration segment includes the start timestamp to the target timestamp of the target recorded song, and the length of the sentence list is the number of sentences in the singing sentence.

5. The method for beautifying singing voice according to claim 1, characterized in that, If a sentence fails to meet the performance standard, any dimension of the performance score is lower than the corresponding scoring threshold. The performance score includes multiple dimensions of scoring.

6. The method for beautifying singing voice according to claim 5, characterized in that, Also includes: If all the sung sentences in the target recorded song meet the standard, no processing is required.

7. The method for beautifying singing voice according to claim 6, characterized in that, The region to be inferred is the time region corresponding to the vocal feature file. The step of inputting the target inferred pitch, the target feature file, and the timbre vector of the target user into the target timbre inference model, and outputting the target beautified audio through the target timbre inference model, includes: The target inference pitch, the vocal feature file, and the timbre vector of the target user are input into the target timbre inference model to obtain the audio after the first inference. The unqualified singing sentences in the target recorded song are identified as the first singing sentences, and the qualified singing sentences in the target recorded song are identified as the second singing sentences. The third singing sentence corresponding to the first singing sentence is selected from the audio after the first reasoning, and the second singing sentence and the third singing sentence are concatenated to obtain the concatenated result; The spliced ​​result is cropped according to the third cropping strategy and the target recorded song to obtain the cropped result, and the cropped result is directly determined as the target beautified audio.

8. The method for enhancing singing voice according to claim 7, characterized in that, After cropping the spliced ​​result according to the third cropping strategy and the target recorded song to obtain the cropped result, the method further includes: The updated vocal feature file is obtained based on the cropped result; The target inferred pitch, the updated vocal feature file, and the timbre vector of the target user are input into the target timbre inference model to obtain the target beautified audio.

9. The method for enhancing singing voice according to any one of claims 1 to 8, characterized in that, The step of sending the beautified audio of the target device to the user terminal includes: If the enhancement service instruction is the audition segment enhancement service instruction, then the original singing segment corresponding to the target enhanced audio is determined from the target recorded song, and the original singing segment and the target enhanced audio are sent to the user terminal; wherein, the original singing segment is the segment corresponding to the region to be inferred; If the enhancement service instruction is the complete song enhancement service instruction, then the target enhanced audio will be sent to the user terminal.

10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the singing enhancement method as described in any one of claims 1 to 9.

11. A computer-readable storage medium, characterized in that, Used to store computer programs; wherein, when the computer programs are executed by a processor, they implement the singing enhancement method as described in any one of claims 1 to 9.

12. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the singing enhancement method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method and system for intelligently adjusting sound effects

    CN109905806A

  • Audio adjustment method, computer device and computer program product

    CN114743526A