Correction method, electronic device and computer storage medium
By receiving the singer's voice and accompaniment sound in the reference song, the user's singing voice is corrected and mixed, which solves the problem of poor singing sound editing effect in the prior art, and achieves a singing effect closer to the sound quality of professional singers, improving the user's singing experience.
Patent Information
- Application Number
- CN202111215480.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-19
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2041-10-19
AI Technical Summary
The existing singing and sound refining technology corrects the singing voice of ordinary people and cannot effectively remove defects such as puffing and squirting, breathing and off-tuning.
By receiving the singer's voice and accompaniment sound in the reference song, the user's singing voice is corrected using a preset correction algorithm or neural network model, including corrections of audio amplitude, spectrum envelope, harmonics and spectral diagrams, and mixing it with instrument accompaniment to form a singing voice closer to the reference song.
It improves the correction effect of singing sound, makes the user's singing sound closer to the sound quality of professional singers, enhances the singing effect and sound effects, and improves the user's singing experience.
Smart Images

Figure CN115995223B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a singing correction technology in an electronic device, and in particular to a correction method, an electronic device and a computer storage medium. Background Art
[0002] At present, music-related entertainment applications such as online karaoke and short videos are becoming more and more popular. Professional singers have received professional training, so their singing sounds better; but for most ordinary people, there are always more or less defects in their singing, such as puffing sounds, breathing sounds, out-of-tune singing, and insufficient music.
[0003] In order to make ordinary people's singing voices sound better, many singing tuning technologies have emerged. However, these tuning technologies mainly use traditional signal processing methods on songs that users have already sung. For example, certain special frequency bands are reduced or expanded to solve noises such as breathing, puffing, and mic spraying. However, this method only uses some generally applicable methods to correct songs that have already been sung, which results in poor results for the corrected songs. It can be seen that the existing song correction methods have technical problems with poor results. Summary of the Invention
[0004] The embodiments of the present application provide a correction method, an electronic device, and a computer storage medium, which can improve the effect of a song after correction.
[0005] The technical solution of this application is achieved as follows:
[0006] In a first aspect, an embodiment of the present application provides a correction method, comprising:
[0007] For a reference song, receiving a user's singing voice;
[0008] Acquire the singing voice of the singer and the accompaniment sound of the instrument in the reference song;
[0009] Correcting the user's singing voice according to the singer's singing voice to obtain a corrected singing voice of the user;
[0010] The corrected singing voice of the user and the accompaniment sound of the instrument are mixed to obtain the singing song of the user.
[0011] In a second aspect, an embodiment of the present application provides an electronic device, including:
[0012] A receiving module, configured to receive a user's singing voice for a reference song;
[0013] An acquisition module, configured to acquire the singing voice of a singer and the accompaniment sounds of musical instruments in the reference song;
[0014] a correction module, configured to correct the singing voice of the user according to the singing voice of the singer, and obtain a corrected singing voice of the user;
[0015] The mixing module is used to mix the corrected singing voice of the user and the accompaniment sound of the instrument to obtain the singing song of the user.
[0016] In a third aspect, an embodiment of the present application provides an electronic device comprising: a processor and a storage medium storing instructions executable by the processor; the storage medium relies on the processor to perform operations through a communication bus, and when the instructions are executed by the processor, the correction method described in one or more of the above embodiments is executed.
[0017] In a fourth aspect, an embodiment of the present application provides a computer storage medium storing executable instructions. When the executable instructions are executed by one or more processors, the processors execute the correction method described in one or more of the above embodiments.
[0018] The embodiment of the present application provides a correction method, electronic device and computer storage medium, including: receiving a user's singing voice for a reference song, obtaining the singer's singing voice and the accompaniment sound of an instrument in the reference song, correcting the user's singing voice according to the singer's singing voice to obtain a corrected user's singing voice, and mixing the corrected user's singing voice and the accompaniment sound of the instrument to obtain the user's singing song; that is, in the embodiment of the present application, after receiving the user's singing voice for the reference song, the user's singing voice is corrected according to the singer's singing voice in the reference song to obtain a corrected user's singing voice. In this way, the user's singing voice is corrected based on the singer's singing voice as the standard, making the correction of the user's singing voice more targeted. Compared with using a general correction method, the corrected user's singing voice is closer to the singer's singing voice. Based on this, the corrected user's singing voice and the accompaniment sound of the instrument are mixed to obtain the user's singing song, making the user's singing song closer to the reference song. When the user sings the reference song, the correction effect of the singing voice is improved, thereby improving the effect of the user singing the reference song. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 A schematic flow chart of an optional correction method provided in an embodiment of the present application;
[0020] Figure 2 A flowchart of Example 1 of an optional correction method provided in an embodiment of the present application;
[0021] Figure 3 A flowchart of Example 2 of an optional correction method provided in an embodiment of the present application;
[0022] Figure 4 A schematic diagram of the structure of an optional electronic device provided in an embodiment of the present application;
[0023] Figure 5 A schematic structural diagram of another optional electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0024] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.
[0025] Example 1
[0026] The embodiment of the present application provides a correction method, Figure 1 A flow chart of an optional correction method provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the correction method may include:
[0027] S101: Receive a user's singing voice for a reference song;
[0028] At present, a general method is usually used to correct the collected singing voice. However, the correction effect obtained by the general method is poor and it is difficult to meet the user's requirements. In order to improve the correction effect, an embodiment of the present application provides a correction method, which is applied to an electronic device. Moreover, for each reference song, the electronic device can store the singing voice of the singer of the reference song and the accompaniment sound of the instrument. For each reference song, the type of instrument included in the accompaniment sound of the instrument is related to the reference song. For example, the accompaniment sound of the instrument can include: the accompaniment sound of the piano, the accompaniment sound of the guitar, the accompaniment sound of the erhu, the accompaniment sound of the guzheng, etc.
[0029] Specifically, for the reference song, the electronic device receives the user's singing voice, that is, the electronic device receives the user's singing voice for the reference song. For example, the user uses a singing application in the mobile phone to first select the song to be sung, which is the reference song, and then the user sings the reference song, so that the mobile phone collects the user's singing voice for the reference song.
[0030] S102: Acquire the singing voice of the singer and the accompaniment sound of the instrument in the reference song;
[0031] Since the electronic device stores the singing voice of the singer and the accompaniment sound of the instrument of the reference song, after receiving the singing voice of the user for the reference song, the singing voice of the reference song and the accompaniment sound of the instrument can be directly obtained, and the singing voice of the singer and the accompaniment sound of the instrument can also be separated from the reference song in real time. Here, the embodiment of the present application does not make specific limitations on this.
[0032] S103: Correcting the user's singing voice according to the singer's singing voice to obtain a corrected user's singing voice;
[0033] After obtaining the singer's singing voice and the user's singing voice, since the singer's singing voice generally has accurate pitch and good sound quality, the singer's singing voice can be used as a reference. That is, the user's singing voice can be corrected based on the singer's singing voice to obtain a corrected user's singing voice. To correct the user's singing voice, a preset correction algorithm can be used to correct the user's singing voice, or a neural network model can be used to correct the user's singing voice. This embodiment of the present application does not specifically limit this.
[0034] In order to correct the user's singing voice, in an optional embodiment, S103 may include:
[0035] According to the singing voice of the singer, the user's singing voice is corrected by using a preset singing voice correction algorithm to obtain the corrected singing voice of the user.
[0036] Specifically, a preset singing voice correction algorithm can be used to correct the user's singing voice based on the singer's singing voice.
[0037] Here, it should be pointed out that, in correcting the user's singing voice according to the singer's singing voice using a preset singing voice correction algorithm, the singer's singing voice and the user's singing voice can be segmented first, and for each corresponding segment, according to the segmentation of the singer's singing voice, the preset singing voice correction algorithm can be used to correct the segmentation of the user's singing voice corresponding to the segment of the singer's singing voice, until each segment is corrected, thereby obtaining the corrected user's singing voice, wherein the duration of the above segments can be selected according to actual conditions.
[0038] Of course, the entire singer's singing voice can also be corrected using a preset singing voice correction algorithm. Here, the embodiment of the present application does not make any specific limitations on this.
[0039] The preset singing voice correction algorithm can be used to correct different waveform features of the singing voice. In an optional embodiment, the user's singing voice is corrected using the preset singing voice correction algorithm based on the singer's singing voice to obtain the corrected user's singing voice, including:
[0040] According to the singing voice of the singer, the audio amplitude in the user's singing voice is corrected using a preset singing voice correction algorithm to obtain a corrected singing voice of the user.
[0041] Specifically, each singing voice has a waveform of audio amplitude. Here, the audio amplitude waveform of the user's singing voice can be corrected based on the audio amplitude waveform of the singer's singing voice using a preset singing voice correction algorithm. The singing voice correction algorithm uses a smoothing method to correct the audio amplitude waveform, so that the corrected audio amplitude waveform approaches the singer's audio amplitude waveform, thereby obtaining the corrected singing voice of the user.
[0042] In this way, the audio amplitude of the user's singing voice can be smoothly corrected, so that the breathing sound and microphone spraying sound in the corrected user's singing voice can be removed, thereby improving the effect of the corrected user's singing voice.
[0043] In the correction of the user's singing voice, in an optional embodiment, the user's singing voice is corrected using a preset singing voice correction algorithm based on the singer's singing voice to obtain the corrected user's singing voice, including:
[0044] According to the singing voice of the singer, a preset singing voice correction algorithm is used to correct the spectrum envelope in the user's singing voice to obtain the corrected singing voice of the user.
[0045] Specifically, each singing voice also has a waveform of a spectrum envelope. Here, the waveform of the spectrum envelope of the user's singing voice can be corrected based on the waveform of the spectrum envelope of the singer's singing voice using a preset singing voice correction algorithm. The singing voice correction algorithm uses a smoothing method to correct the waveform of the spectrum envelope, so that the waveform of the corrected spectrum envelope approaches the waveform of the singer's spectrum envelope, thereby obtaining the corrected singing voice of the user.
[0046] In this way, the smooth correction of the spectral envelope of the user's singing voice can be achieved. Similar to the smooth correction of the audio amplitude mentioned above, the breathing sound and microphone spraying sound in the corrected user's singing voice can be removed, thereby improving the effect of the corrected user's singing voice.
[0047] In actual applications, after using the preset singing voice correction algorithm to correct the audio amplitude and spectrum envelope, many noises in singing such as breathing and microphone spraying can basically be removed, thereby greatly improving the sound effect of the user's singing voice.
[0048] In addition, the singing voice correction algorithm can also correct the harmonics in the singing voice. In an optional embodiment, the user's singing voice is corrected using a preset singing voice correction algorithm based on the singer's singing voice to obtain the corrected user's singing voice, including:
[0049] According to the singing voice of the singer, the harmonics in the user's singing voice are corrected using a preset singing voice correction algorithm to obtain a corrected singing voice of the user.
[0050] In practical applications, pleasant sounds often contain rich harmonics, so harmonic correction is also very important for correcting singing voices.
[0051] In order to correct the harmonics in the singing voice, specifically, each singing voice also has a harmonic waveform diagram. Here, the harmonic waveform diagram of the user's singing voice can be corrected according to the harmonic waveform diagram of the singer's singing voice using a preset singing voice correction algorithm. Among them, the singing voice correction algorithm uses a smoothing method to correct the harmonic waveform diagram, so that the corrected harmonic waveform diagram approaches the singer's harmonic waveform diagram, thereby obtaining the corrected singing voice of the user.
[0052] When a singer has multiple harmonics, the user's preferred harmonics can be selected according to the user's preferences. Based on the user's preferred harmonics, a preset singing voice correction algorithm is used to correct the harmonics in the user's singing voice to obtain the corrected user's singing voice.
[0053] In this way, the harmonics in the user's singing voice can be corrected, and the sound effect of the corrected user's singing voice can be improved.
[0054] In addition, in the singing voice correction algorithm, the spectrogram in the singing voice can also be corrected. In an optional embodiment, the user's singing voice is corrected using a preset singing voice correction algorithm based on the singer's singing voice to obtain the corrected user's singing voice, including:
[0055] According to the singing voice of the singer, the spectrogram in the singing voice of the user is corrected using a preset singing voice correction algorithm to obtain the corrected singing voice of the user.
[0056] Specifically, each singing voice also has a spectrogram waveform. Here, the waveform of the spectrogram of the user's singing voice can be corrected based on the waveform of the spectrogram of the singer's singing voice using a preset singing voice correction algorithm. The spectrogram contains features of the audio time domain and frequency domain. In correcting the waveform of the spectrogram, it can include correction of the time domain, correction of the frequency domain, or correction of both the time domain and the frequency domain. In addition, the singing voice correction algorithm uses a smoothing method to correct the waveform of the spectrogram, so that the waveform of the corrected spectrogram is close to the waveform of the singer's spectrogram, thereby obtaining the corrected singing voice of the user.
[0057] In this way, the sound effect of the corrected user's singing voice can be further improved.
[0058] Some received user singing voices may be out of tune. To eliminate the out-of-tune, the pitch of the user's singing voice may be corrected. In an optional embodiment, the user's singing voice is corrected using a preset singing voice correction algorithm based on the singer's singing voice to obtain the corrected user's singing voice, including:
[0059] According to the singing voice of the singer, the pitch of the user's singing voice is corrected using a preset singing voice correction algorithm to obtain a preliminarily corrected singing voice of the user;
[0060] The pitch of the user's singing voice after preliminary correction is corrected according to the accompaniment sound of the musical instrument to obtain the corrected user's singing voice.
[0061] Specifically, the pitch of the user's singing voice can be corrected according to the singer's singing voice using a preset singing voice correction algorithm to obtain the corrected user's singing voice. The pitch of the user's singing voice can also be corrected according to the accompaniment sound of the musical instrument to obtain the corrected user's singing voice. The pitch of the user's singing voice can also be corrected according to the singer's singing voice using a preset singing voice correction algorithm to obtain the preliminarily corrected user's singing voice. The pitch of the preliminarily corrected user's singing voice can then be corrected according to the accompaniment sound of the musical instrument to obtain the corrected user's singing voice. Here, the embodiments of the present application do not make specific limitations on this.
[0062] First, based on the singing voice of the singer, the pitch of the user's singing voice is corrected by using a preset singing voice correction algorithm to obtain the singing voice of the user after preliminary correction. Then, based on the accompaniment sound of the musical instrument, the pitch of the singing voice of the user after preliminary correction is corrected to obtain the corrected singing voice of the user, each singing voice has a waveform of the pitch. Here, based on the waveform of the pitch of the singer's singing voice, the waveform of the pitch of the user's singing voice can be corrected by using a preset singing voice correction algorithm. Among them, the singing voice correction algorithm adopts a smoothing method to correct the waveform of the pitch, so that the waveform of the corrected pitch is close to the waveform of the singer's pitch, thereby obtaining the singing voice of the user after preliminary correction.
[0063] Since the accompaniment sound of the musical instrument can often determine the pitch of the reference song at certain nodes, here, the user's singing voice after preliminary correction is corrected according to the accompaniment sound of the musical instrument to further correct the out-of-tune part of the user's singing voice and obtain the corrected user's singing voice.
[0064] Here, it should be noted that the above-mentioned preset singing voice correction algorithm mainly includes the correction of the waveform graphs of the following different characteristics of the singing voice: correction of the waveform graph of the audio amplitude, correction of the waveform graph of the spectrum envelope, correction of the waveform graph of the spectrogram, correction of the waveform graph of the pitch, etc. The singing voice correction algorithm may include the correction of the waveform graph of one or more of the above characteristics, and the order of correction of the waveform graph of each characteristic is not limited.
[0065] In addition, in order to correct the user's singing voice, a neural network model may be used. In an optional embodiment, S103 may include:
[0066] The singer's singing voice and the user's singing voice are input into a pre-trained machine learning model to obtain the corrected user's singing voice.
[0067] Here, a machine learning model needs to be pre-trained. Based on the trained machine learning model, the singer's singing voice and the user's singing voice are input into the trained machine learning model, so that the corrected user's singing voice can be output. In order to obtain the trained machine learning model, in an optional embodiment, the above method further includes:
[0068] Get a sample dataset;
[0069] The sample data set is input into the machine learning model for training to obtain a trained machine learning model.
[0070] Specifically, a sample data set is first obtained, wherein the sample data set is: the singing voice of the singer in the collected song, the singing voices of at least two users for the collected song, and the corrected singing voice of each of the at least two users; that is, the song is first collected, and then the singing voice of the collected song is separated, and then the singing voices of at least two users for the song are collected. A preset singing voice correction algorithm can be used to obtain the corrected singing voice of each of the at least two users, and the singing voice of the singer in the collected song, the singing voice of at least two users for the collected song, and the corrected singing voice of each of the at least two users form a sample data set. It should be noted that the more singing voices of users in the sample data set, the better the correction effect of the trained machine learning model.
[0071] After obtaining the sample data set, the sample data set is input into the machine learning model for training to obtain a trained machine learning model. Here, after obtaining the trained machine learning model, the machine learning model can also be tested, and the machine training model that passes the test is determined as the trained machine learning model, and the machine learning model that fails the test continues to be trained until the test passes.
[0072] In addition, for some users, even if the above singing voice correction algorithm or the trained machine learning model is used, the sound effect of the user's singing voice after correction is still not good. In order to obtain a better sound effect, in an optional embodiment, the above method further includes:
[0073] The vocal features of the singer in the singing voice of the singer are replaced with the vocal features of the user to obtain a corrected singing voice of the user.
[0074] Specifically, based on the sound feature extraction technology, the singer's voice features and the user's voice features can be extracted, and then the singer's voice features are changed into the user's voice features in the singer's singing voice to obtain the corrected user's singing voice, so as to realize voice changing, improve the corrected sound effect, and thus improve the user experience.
[0075] S104: Mixing the corrected singing voice of the user and the accompaniment sound of the instrument to obtain the user's singing song.
[0076] After the user's singing voice is corrected, the corrected user's singing voice is obtained, and the corrected user's singing voice and the accompaniment sound of the instrument are mixed to obtain the user's singing song. In order to further improve the sound effect of the user's singing song, in an optional embodiment, S104 may include:
[0077] Selecting a portion of the corrected singing voice of the user that does not meet a preset singing quality condition;
[0078] Finding a singing voice corresponding to the selected portion from the singing voice of the singer, replacing the voice features of the singer in the singing voice corresponding to the selected portion with the voice features of the user, to obtain a replaced singing voice;
[0079] Replacing the portion of the corrected user's singing voice that does not meet the preset singing quality condition with the replaced singing voice, thereby obtaining the replaced user's singing voice;
[0080] The replaced singing voice of the user and the accompaniment sound of the musical instrument are mixed to obtain the user's singing song.
[0081] In actual applications, there will still be some parts of the user's singing songs with poor sound effects after correction. In order to avoid these problems, singing quality conditions are preset here. For example, some software scoring methods can be used to score the mixed songs in sections, and the song parts with lower scores are selected from the singer's singing voice. Then, the singing voices corresponding to the selected parts are changed, that is, the singer's voice characteristics are replaced with the user's voice characteristics, thereby obtaining the replaced singing voice. Finally, the parts of the corrected user's singing songs that do not meet the preset singing quality conditions are replaced with the replaced singing voice to obtain the replaced user's singing voice. The replaced user's singing voice and the accompaniment sounds of musical instruments are mixed to obtain the user's singing songs.
[0082] In this way, the user's voice characteristics can be further utilized to perform voice change to replace the poorer parts of the corrected user's singing voice, thereby improving the sound quality of the user's singing songs.
[0083] Finally, in order to improve the sound effect of the user's singing song, in an optional embodiment, S104 may include:
[0084] performing smoothing processing on the corrected singing voice of the user to obtain a smoothed singing voice of the user;
[0085] The user's singing voice and the accompaniment sound of the musical instrument after smoothing are mixed to obtain the user's singing song.
[0086] That is to say, in order to prevent the user's singing songs from having fluctuating singing voices due to the correction, the corrected user's singing voice is smoothed here. For example, the waveform of the audio amplitude in the user's singing voice can be smoothed, the waveform of the spectrum envelope in the user's singing voice can be smoothed, the waveform of the spectrogram in the user's singing voice can be smoothed, and the waveform of the tone in the user's singing voice can be smoothed. In actual applications, one or more of the above waveforms can be smoothed, and the order of smoothing multiple waveforms is not specifically limited.
[0087] The following examples are used to illustrate the correction method described in one or more of the above embodiments.
[0088] Figure 2 A flowchart of an example 1 of an optional correction method provided in an embodiment of the present application is shown as follows: Figure 2 As shown, the original music (equivalent to the above-mentioned reference song) is input into the vocal separation module to extract the original singer (equivalent to the singing voice of the above-mentioned singer), and the remaining accompaniment enters the instrument separation module to extract various instrument sounds (equivalent to the accompaniment sounds of the above-mentioned instruments) as needed. The separated original singer enters the original singer analysis module to analyze the audio characteristics in segments or as a whole, such as the fundamental frequency change law, harmonic law, spectrum high and low frequency difference, sound amplitude and other singing characteristics. At the same time, combined with the instrument sounds separated on demand, such as drum sounds, guitar sounds and piano sounds, it can better match various song characteristics such as rhythm and style.
[0089] These feature information enters the solo singing analysis module, where the original and solo singing (equivalent to the singing voice of the above-mentioned user) are segmented or analyzed for audio features according to segmentation information similar to the original singing. The most similar fragments are matched through multiple matching rules. The matched fragments are corrected in the solo singing tuning module according to the analysis features of the original singing, including time adjustment, out-of-tune correction, amplitude correction, spectrum enhancement or attenuation, harmonic correction, etc.
[0090] Among them, the vocal separation module and the instrument separation module can be traditional methods or more advanced neural network methods. For the neural network method, the two modules can often use the same model to complete the separation.
[0091] In addition, the instrument analysis module, original singer analysis module, individual singing analysis module, and individual singing tuning module can be implemented using traditional signal processing methods or neural network methods. For the neural network method, the same model can often be used to complete the analysis and correction.
[0092] Figure 3A flow chart of Example 2 of an optional correction method provided in an embodiment of the present application is shown as follows: Figure 3 As shown, the original vocals (equivalent to the singing voice of the singer) and the original solo singing (equivalent to the singing voice of the user) are segmented according to a certain duration and the sound is corrected segment by segment.
[0093] The audio amplitude calculation module calculates the audio amplitude, calculating the amplitude of each segment of the original vocal, and then adjusts the amplitude of each segment of the original vocal to approximate the audio amplitude of the original vocal. The spectrum envelope calculation module analyzes the differences between the spectrum envelopes of the original vocal and the original vocal, and performs amplitude correction on frequency bands where the envelope amplitude differs significantly. Using these two modules, many specific singing defects, such as breathing and mic popping, can be largely eliminated.
[0094] Pleasant sounds often contain rich harmonics, so the harmonic tuning module is also very important for singing tuning. In the harmonic analysis module, the fundamental frequency of the original singer's voice and the original solo singing is calculated, and their harmonics are analyzed. The original solo singing is corrected according to the harmonic distribution of the original singer's voice (additional harmonic correction can also be performed according to the harmonic characteristics of different types of pleasant sounds and user preferences).
[0095] Spectrograms contain audio features in both the time and frequency domains and are widely used in various audio deep learning networks. Due to their importance, they are also a key feature for vocal tuning. In the spectrogram analysis module, spectrograms of the original vocalist and the original solo performance are calculated. For any significant differences in spectrograms, the original solo performance's spectrogram is corrected to approximate the original vocalist's spectrogram.
[0096] Non-professional singers often sing out of tune, too fast or too slow. We combine the rhythm and musical style of the accompaniment with the pitch of the original singer's voice to obtain accurate pitch analysis results. We analyze the original singing pitch to determine if it's off-tune and make fine-tuning corrections.
[0097] In order to avoid discontinuity in the adjustments of each segment, it is necessary to gradually correct the song to approximate the original song through time smoothing.
[0098] In the vocal feature extraction module, the user's voice features are extracted, and the original singer's voice can be directly changed into the user's singing voice.
[0099] All of the above modules can utilize traditional signal processing methods or popular deep learning approaches, each with varying computational loads. While the modules are shown in the diagram in a sequential order, in practice, the order of the modules and the complexity of the tuning algorithms can be adjusted based on the target platform's resources. Furthermore, the tuning strength of each module can be adjusted, and corresponding modules can be added or removed, depending on the selected tuning mode.
[0100] In the mixing module, for the output of the voice change result, the quality of the user's original singing program is judged. The poor part can be directly replaced by the voice change result. The voice change result can also be fully adopted, partially adopted, or directly prohibited according to the user's choice of voice correction mode.
[0101] Through the above examples, by analyzing the clean original singing, we can obtain the targeted singing characteristics of the current song, and perform targeted tuning, and the tuning result is more ideal.
[0102] An embodiment of the present application provides a correction method, comprising: receiving a user's singing voice for a reference song, obtaining the singer's singing voice and the accompaniment sound of an instrument in the reference song, correcting the user's singing voice according to the singer's singing voice to obtain a corrected user's singing voice, and mixing the corrected user's singing voice and the accompaniment sound of an instrument to obtain the user's singing song; that is, in an embodiment of the present application, after receiving the user's singing voice for the reference song, the user's singing voice is corrected according to the singer's singing voice in the reference song to obtain a corrected user's singing voice. In this way, the user's singing voice is corrected based on the singer's singing voice, making the correction of the user's singing voice more targeted. Compared with the general correction method, the corrected user's singing voice is closer to the singer's singing voice. Based on this, the corrected user's singing voice and the accompaniment sound of the instrument are mixed to obtain the user's singing song, making the user's singing song closer to the reference song. When the user sings the reference song, the correction effect of the singing voice is improved, thereby improving the effect of the user singing the reference song.
[0103] Example 2
[0104] Based on the same inventive concept, an embodiment of the present application provides an electronic device, Figure 4 A schematic diagram of the structure of an optional electronic device provided in an embodiment of the present application is shown in FIG. Figure 4 As shown, the electronic device includes: a receiving module 41, an acquisition module 42, a correction module 43 and a mixing module 44; wherein,
[0105] The receiving module 41 is used to receive the user's singing voice for the reference song;
[0106] An acquisition module 42 is used to acquire the singing voice of the singer and the accompaniment sound of the instrument in the reference song;
[0107] A correction module 43 is used to correct the user's singing voice according to the singer's singing voice to obtain a corrected user's singing voice;
[0108] The mixing module 44 is used to mix the corrected singing voice of the user and the accompaniment sound of the instrument to obtain the user's singing song.
[0109] In an optional embodiment, the correction module 43 is specifically configured to:
[0110] According to the singing voice of the singer, the user's singing voice is corrected by using a preset singing voice correction algorithm to obtain the corrected singing voice of the user.
[0111] In an optional embodiment, the correction module 43 corrects the user's singing voice according to the singer's singing voice using a preset singing voice correction algorithm, and obtains the corrected user's singing voice, including:
[0112] According to the singing voice of the singer, the audio amplitude in the user's singing voice is corrected using a preset singing voice correction algorithm to obtain a corrected singing voice of the user.
[0113] In an optional embodiment, the correction module 43 corrects the user's singing voice according to the singer's singing voice using a preset singing voice correction algorithm, and obtains the corrected user's singing voice, including:
[0114] According to the singing voice of the singer, a preset singing voice correction algorithm is used to correct the spectrum envelope in the user's singing voice to obtain the corrected singing voice of the user.
[0115] In an optional embodiment, the correction module 43 corrects the user's singing voice according to the singer's singing voice using a preset singing voice correction algorithm, and obtains the corrected user's singing voice, including:
[0116] According to the singing voice of the singer, the harmonics in the user's singing voice are corrected using a preset singing voice correction algorithm to obtain a corrected singing voice of the user.
[0117] In an optional embodiment, the correction module 43 corrects the user's singing voice according to the singer's singing voice using a preset singing voice correction algorithm, and obtains the corrected user's singing voice, including:
[0118] According to the singing voice of the singer, the spectrogram in the singing voice of the user is corrected using a preset singing voice correction algorithm to obtain the corrected singing voice of the user.
[0119] In an optional embodiment, the correction module 43 corrects the user's singing voice according to the singer's singing voice using a preset singing voice correction algorithm, and obtains the corrected user's singing voice, including:
[0120] According to the singing voice of the singer, the pitch of the user's singing voice is corrected using a preset singing voice correction algorithm to obtain a preliminarily corrected singing voice of the user;
[0121] The pitch of the user's singing voice after preliminary correction is corrected according to the accompaniment sound of the musical instrument to obtain the corrected user's singing voice.
[0122] In an optional embodiment, the correction module 43 is specifically configured to:
[0123] The singer's singing voice and the user's singing voice are input into a pre-trained machine learning model to obtain the corrected user's singing voice.
[0124] In an optional embodiment, the electronic device is further used for:
[0125] Obtain a sample data set; wherein the sample data set includes: the singing voice of the singer in the collected song, the singing voices of at least two users of the collected song, and the corrected singing voice of each user of the at least two users;
[0126] The sample data set is input into the machine learning model for training to obtain a trained machine learning model.
[0127] In an optional embodiment, the electronic device is further used for:
[0128] The vocal features of the singer in the singing voice of the singer are replaced with the vocal features of the user to obtain a corrected singing voice of the user.
[0129] In an optional embodiment, the audio mixing module 44 is specifically configured to:
[0130] Selecting a portion of the corrected singing voice of the user that does not meet a preset singing quality condition;
[0131] Finding a singing voice corresponding to the selected portion from the singing voice of the singer, replacing the voice features of the singer in the singing voice corresponding to the selected portion with the voice features of the user, to obtain a replaced singing voice;
[0132] Replacing the portion of the corrected user's singing voice that does not meet the preset singing quality condition with the replaced singing voice, thereby obtaining the replaced user's singing voice;
[0133] The replaced singing voice of the user and the accompaniment sound of the musical instrument are mixed to obtain the user's singing song.
[0134] In an optional embodiment, the audio mixing module 44 is specifically configured to:
[0135] performing smoothing processing on the corrected singing voice of the user to obtain a smoothed singing voice of the user;
[0136] The user's singing voice after smoothing and the accompaniment sound of the instrument are mixed to obtain the user's singing song.
[0137] Figure 5 A schematic diagram of the structure of another optional electronic device provided in an embodiment of the present application, such as Figure 5 As shown, an embodiment of the present application provides an electronic device 500, including: a processor 51 and a storage medium 52 storing instructions executable by the processor; the storage medium 52 relies on the processor 51 to perform operations through a communication bus 53, and when the instructions are executed by the processor, the correction method executed on the processor side in one or more of the above embodiments is executed.
[0138] It should be noted that in actual application, the various components in the terminal are coupled together through the communication bus 53. It is understandable that the communication bus 53 is used to realize the connection and communication between these components. In addition to the data bus, the communication bus 53 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, Figure 5 Various buses are labeled as communication buses 53.
[0139] An embodiment of the present application provides a computer storage medium storing executable instructions. When the executable instructions are executed by one or more processors, the processors execute the correction method described in one or more of the above embodiments.
[0140] Among them, the computer-readable storage medium can be a magnetic random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory (Flash Memory), a magnetic surface storage device, an optical disc, or a compact disc read-only memory (CD-ROM) and other memories.
[0141] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Furthermore, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage, etc.) containing computer-usable program code.
[0142] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0143] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0144] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0145] The above description is merely a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application.
Claims
1. A correction method, characterized in that: include: For a reference song, receiving a user's singing voice; Acquire the singing voice of the singer and the accompaniment sound of the instrument in the reference song; Modifying waveforms of different features of the user's singing voice based on the singer's singing voice to obtain a modified singing voice of the user; wherein the different features include at least one of the following: audio amplitude, spectrum envelope, spectrogram, harmonics, and pitch; Mixing the corrected singing voice of the user and the accompaniment sound of the instrument to obtain a song sung by the user; The mixing of the corrected singing voice of the user and the accompaniment sound of the instrument to obtain the singing song of the user comprises: selecting a portion of the corrected singing voice of the user that does not meet a preset singing quality condition; Finding a singing voice corresponding to the selected portion from the singing voice of the singer, and replacing the voice features of the singer in the singing voice corresponding to the selected portion with the voice features of the user to obtain a replaced singing voice; replacing the portion of the corrected singing voice of the user that does not meet the preset singing quality condition with the replaced singing voice, thereby obtaining the replaced singing voice of the user; The replaced singing voice of the user and the accompaniment sound of the instrument are mixed to obtain the singing song of the user.
2. The method according to claim 1, characterized in that The method of modifying the user's singing voice according to the singer's singing voice to obtain the modified user's singing voice includes: According to the singing voice of the singer, the singing voice of the user is corrected by using a preset singing voice correction algorithm to obtain the corrected singing voice of the user.
3. The method according to claim 2, characterized in that The method of correcting the user's singing voice by using a preset singing voice correction algorithm based on the singer's singing voice to obtain the corrected singing voice of the user includes: According to the singing voice of the singer, the audio amplitude in the singing voice of the user is corrected using a preset singing voice correction algorithm to obtain the corrected singing voice of the user.
4. The method according to claim 2, characterized in that The method of correcting the user's singing voice by using a preset singing voice correction algorithm based on the singer's singing voice to obtain the corrected singing voice of the user includes: According to the singing voice of the singer, a preset singing voice correction algorithm is used to correct the spectrum envelope in the singing voice of the user to obtain the corrected singing voice of the user.
5. The method according to claim 2, characterized in that The method of correcting the user's singing voice by using a preset singing voice correction algorithm based on the singer's singing voice to obtain the corrected singing voice of the user includes: According to the singing voice of the singer, the harmonics in the singing voice of the user are corrected using a preset singing voice correction algorithm to obtain the corrected singing voice of the user.
6. The method according to claim 2, characterized in that The method of correcting the user's singing voice by using a preset singing voice correction algorithm based on the singer's singing voice to obtain the corrected singing voice of the user includes: According to the singing voice of the singer, a preset singing voice correction algorithm is used to correct the spectrogram in the singing voice of the user to obtain the corrected singing voice of the user.
7. The method according to claim 2, characterized in that The method of correcting the user's singing voice by using a preset singing voice correction algorithm based on the singer's singing voice to obtain the corrected singing voice of the user includes: According to the singing voice of the singer, the pitch of the singing voice of the user is corrected using a preset singing voice correction algorithm to obtain the singing voice of the user after preliminary correction; According to the accompaniment sound of the musical instrument, the pitch of the user's singing voice after preliminary correction is corrected to obtain the corrected singing voice of the user.
8. The method according to claim 1, characterized in that The method of modifying the user's singing voice according to the singer's singing voice to obtain the modified user's singing voice includes: The singing voice of the singer and the singing voice of the user are input into a pre-trained machine learning model to obtain the corrected singing voice of the user.
9. The method according to claim 8, characterized in that The method further comprises: Obtaining a sample data set; wherein the sample data set includes: the singing voice of the singer in the collected song, the singing voices of at least two users of the collected song, and the corrected singing voice of each user of the at least two users; The sample data set is input into the machine learning model for training to obtain the trained machine learning model.
10. The method according to claim 1, characterized in that The method further comprises: The voice features of the singer in the singing voice of the singer are replaced with the voice features of the user to obtain a corrected singing voice of the user.
11. The method according to claim 1, wherein The mixing process of the corrected singing voice of the user and the accompaniment sound of the instrument to obtain the singing song of the user comprises: performing smoothing processing on the corrected singing voice of the user to obtain the smoothed singing voice of the user; The user's singing voice and the accompaniment sound of the instrument after smoothing are mixed to obtain the user's singing song.
12. An electronic device, characterized in that: include: A receiving module, configured to receive a user's singing voice for a reference song; An acquisition module, configured to acquire the singing voice of a singer and the accompaniment sounds of musical instruments in the reference song; a correction module, configured to correct waveforms of different features of the user's singing voice based on the singer's singing voice, thereby obtaining a corrected singing voice of the user; wherein the different features include at least one of the following: audio amplitude, spectrum envelope, spectrogram, harmonics, and pitch; A mixing module, configured to mix the corrected singing voice of the user and the accompaniment sound of the instrument to obtain the user's singing song; Wherein, the mixing module is used to: selecting a portion of the corrected singing voice of the user that does not meet a preset singing quality condition; Finding a singing voice corresponding to the selected portion from the singing voice of the singer, and replacing the voice features of the singer in the singing voice corresponding to the selected portion with the voice features of the user to obtain a replaced singing voice; replacing the portion of the corrected singing voice of the user that does not meet the preset singing quality condition with the replaced singing voice, thereby obtaining the replaced singing voice of the user; The replaced singing voice of the user and the accompaniment sound of the instrument are mixed to obtain the singing song of the user.
13. An electronic device, characterized in that: include: A processor and a storage medium storing instructions executable by the processor; The storage medium relies on the processor to perform operations through a communication bus, and when the instructions are executed by the processor, the correction method described in any one of claims 1 to 11 is executed.
14. A computer storage medium, characterized in that Executable instructions are stored, and when the executable instructions are executed by one or more processors, the processors execute the correction method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Singing generation method and apparatus, terminal, and storage medium
CN108831437A
Audio signal processing method and device, electronic equipment and storage medium
CN110675886A
Audio processing method and device, electronic equipment and storage medium
CN112216294A
Audio correction method and device
CN112309409A
Karaoke apparatus, karaoke program, and karaoke system
JP2016071209A