Song processing method and related device

By adjusting the original dry sound signal of the song with constant speed and constant tone and formant information, the problem of color distortion of human voice in the prior art is solved, and the high-fidelity modulation processing is achieved.

CN116153277BActive Publication Date: 2025-06-10TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310192052.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-02
Publication Date
2025-06-10
Estimated Expiration
2043-03-02

AI Technical Summary

Technical Problem

The existing method of processing songs with variable tone causes timbre distortion in the vocal part, which cannot effectively maintain the timbre characteristics of the singer themselves.

Method used

By obtaining the original dry sound signal in the original song, variable speed and unchanging tone and resampling processing are performed to change the tone characteristics; at the same time, the formant peak distribution of the new dry sound signal is adjusted based on the formant peak information of the original dry sound signal to obtain the target dry sound signal after timbre correction.

Benefits of technology

It has achieved a high vocal fidelity effect on the sound signal after changing tones, avoiding the problem of vocal distortion caused by changing tones of the song, and improving the tone retention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116153277B_ABST
    Figure CN116153277B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a song processing method and related devices. The method includes: obtaining an original dry voice signal in an original song; performing pitch-shifting without pitch-changing and resampling processing on the original dry voice signal to change the pitch characteristics of the original dry voice signal to obtain a new dry voice signal; adjusting the formant distribution of the new dry voice signal based on the formant information of the original dry voice signal to obtain a target dry voice signal with corrected timbre. The embodiments of the present application can not only adjust the pitch of the original dry voice signal, but also adjust the spectral information of the new dry voice signal based on the formant information of the original dry voice signal, so that the pitch-shifted voice signal has a high original singer vocal fidelity effect, that is, the timbre is highly maintained, thus avoiding voice distortion after the song is pitch-shifted and affecting the listening experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of audio technology, and in particular, to a method for processing songs and related devices. Background Art

[0002] In daily life applications, different users have different pitch preferences or requirements for songs.

[0003] Currently, the mainstream pitch adjustment solution is to directly use a pitch change tool based on time-scale modification (TSM) to perform unified pitch adjustment on songs. However, inevitably, in the vocal part (which can be called dry voice or unaccompanied pure singing voice) processed by this pitch adjustment solution, there is still a chipmunk effect like the voice of a minion. The adverse manifestations of this effect are that the vocal tract of the vocal part has serious timbre distortion. In terms of listening experience, the processed vocal sound is very different from the timbre characteristics of the singer's own voice.

[0004] In view of this, the related art does not provide an effective solution. Summary of the Invention

[0005] The embodiments of the present application provide a method for processing songs and related devices, which are used to solve the technical problem of timbre distortion in the dry voice part of existing pitch-changed songs.

[0006] The first aspect of the embodiments of the present application provides a method for processing songs, including:

[0007] Obtaining the original dry voice signal in the original song;

[0008] Performing variable speed without pitch change and resampling processing on the original dry voice signal to change the pitch characteristics of the original dry voice signal to obtain a new dry voice signal;

[0009] Adjusting the formant distribution of the new dry voice signal based on the formant information of the original dry voice signal to obtain a target dry voice signal with corrected timbre.

[0010] The second aspect of the embodiments of the present application provides an electronic device, including:

[0011] A central processing unit, a memory, and an input / output interface;

[0012] The memory is a transient storage memory or a persistent storage memory;

[0013] The central processing unit is configured to communicate with the memory and execute the instruction operations in the memory to execute the method described in the first aspect or any specific implementation manner of the first aspect of the embodiments of the present application.

[0014] The third aspect of the embodiments of the present application provides a computer-readable storage medium, including instructions, which when running on a computer, cause the computer to execute the method described in the first aspect or any specific implementation manner of the first aspect of the embodiments of the present application.

[0015] The fourth aspect of the embodiments of the present application provides a computer program product including instructions or a computer program, which when running on a computer, cause the computer to execute the method described in the first aspect or any specific implementation manner of the first aspect of the embodiments of the present application.

[0016] As can be seen from the above technical solutions, the embodiments of the present application have at least the following advantages:

[0017] The embodiments of the present application can not only adjust the pitch of the original dry voice signal, but also adjust the spectral information of the new dry voice signal based on the formant information of the original dry voice signal, so that the pitch-adjusted voice signal has a high original singer voice fidelity effect, that is, the timbre is highly maintained, avoiding the voice distortion caused by the pitch change of the song and affecting the listening experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings.

[0019] Figure 1 It is a schematic diagram of the system architecture of the embodiments of the present application;

[0020] Figure 2 It is a schematic flowchart of a song processing method according to an embodiment of the present application;

[0021] Figure 3 It is another schematic flowchart of a song processing method according to an embodiment of the present application;

[0022] Figure 4 It is another schematic flowchart of a song processing method according to an embodiment of the present application;

[0023] Figure 5 It is another schematic flowchart of a song processing method according to an embodiment of the present application;

[0024] Figure 6 It is a signal spectrum diagram of a song processing method according to an embodiment of the present application;

[0025] Figure 7 It is a schematic structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe this application in detail with reference to the accompanying drawings. The described embodiments should not be construed as limitations on this application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.

[0027] Terms such as "first," "second," "third," "fourth," etc. (if any) in the specification, claims, and drawings of this application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order different from that illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0028] In the following description, expressions such as "a specific embodiment" or "a specific example" are used, which describe subsets of all possible embodiments. However, it can be understood that "a specific embodiment" or "a specific example" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict. In the following description, the term "a plurality" refers to at least two. When it is said in this application that a certain value reaches a threshold (if any), in some specific examples, it may include the case where the former is greater than the latter.

[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0030] For ease of understanding and explanation, before further elaborating on this application in detail, the nouns and terms involved in the embodiments of this application will be explained. The nouns and terms involved in the embodiments of this application are subject to the following explanations.

[0031] STFT: Short-Time Fourier Transform.

[0032] PhaseVocoder: Phase Vocoder, which is used to change the speed of the waveform without changing the pitch, that is, variable speed without pitch change.

[0033] TSM: The full name is Time-scale modification, which is translated as time-domain companding or variable-speed pitch-invariant algorithm. As the name implies, TSM is an algorithm that can change the "speech rate" of audio without changing its pitch. Simply put, TSM divides an audio signal into frames of different lengths, then performs a series of processes on each frame, such as stretching or compressing, and then overlays these frames into a synthesized signal. In many TSM methods, there is often an overlap part when two frames are overlaid. Of course, this overlap part needs to be processed through a series of processes to reduce the influence caused by factors such as phase discontinuity and amplitude fluctuation.

[0034] Wsola: It is a time stretching (TSM) algorithm optimized from ola to sola. Briefly speaking, it realizes the elongation or shortening of the signal length by repeating or deleting frame segments, so as to achieve the purpose of variable speed.

[0035] Phasevocoder: Different from wsola, it completes the stretching and shrinking of time-domain frames by adjusting the phase between frames (frequency domain), that is, it mainly focuses on adjusting frequency-domain information.

[0036] Pitch change: It means adjusting the key of a song, such as changing it to the key of C or D, etc., the main tone of the music.

[0037] The song processing method provided by the embodiments of this application can be applied to an application environment as Figure 1 shown. Taking the audio of this application as an example, it can specifically contain a dry-sound song (including an a cappella pure vocal work). Among them, the terminal 102 communicates with the server 101 through the network. The data storage system 100 can store the data that the server 101 needs to process; the data storage system 100 can be integrated on the server 101, or can be placed in the cloud or other network servers. The terminal 102 can obtain the original song input by the user, specifically, it can obtain the original dry-sound signal in the original song, and can also obtain the original accompaniment signal, and send the obtained sound signal to the server 101. The server 101 can perform processing such as pitch change and timbre correction for the dry sound based on the obtained sound signal. In addition, the server 101 can also send the processed sound signal to the terminal 102 for playback. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The server 101 can be implemented by an independent server or a server cluster composed of multiple servers. It should be noted that the method provided by the embodiments of this application can be jointly implemented by the terminal device and the server as described above, or can be all implemented on the server side, or can also be all implemented on the terminal device side, which can be specifically determined according to the actual application scenario and is not limited here.

[0038] The method of the present application will be further described in detail below.

[0039] Please refer to Figure 2 , a specific embodiment of the song processing method is provided in the first aspect of the present application. The embodiment includes the following operation steps:

[0040] 21. Obtain the original dry voice signal in the original song.

[0041] In practical applications, in order to provide users with song library works that meet their pitch perception, the pitch of the dry voice signal in the original song can be obtained and adjusted. In this embodiment, the dry voice signal specifically refers to the pure human voice signal emitted by the singer's vocal tract in the case of unaccompanied aliasing, such as the humming sound of the lyrics "ah a".

[0042] 22. Perform pitch-shifting without pitch-changing and resampling processing on the original dry voice signal to change the pitch characteristics of the original dry voice signal to obtain a new dry voice signal.

[0043] In the face of the user's demand for adjusting the pitch of the song, the pitch-shifting without pitch-changing and resampling scheme can be combined to process the original dry voice signal to serially adjust the pitch of the original dry voice signal, thereby obtaining the required new dry voice signal. Specifically, the pitch-shifting without pitch-changing processing can be implemented using the TSM algorithm. According to different pitch-changing requirements of raising or lowering the pitch, the resampling processing can be specifically divided into downsampling or upsampling.

[0044] 23. Adjust the formant distribution of the new dry voice signal based on the formant information of the original dry voice signal to obtain the target dry voice signal with corrected timbre.

[0045] Considering that the new dry voice signal obtained by pitch-changing is prone to the chipmunk effect or timbre distortion, such as the vocal sound played sounds similar to the voice of a Minion and loses the timbre characteristics of the singer itself. And the formant information of the sound signal can largely determine the timbre of the sound signal. Therefore, the formant distribution of the new dry voice signal can be adjusted based on the formant information of the original dry voice signal to obtain the target dry voice signal with corrected timbre. In other words, while promoting the pitch-changing of the original dry voice signal, its original timbre characteristics can be retained, so as to highly meet the user's adjustment requirements for the song and improve the user experience.

[0046] In summary, the embodiment of the present application can not only adjust the pitch of the original dry voice signal, but also adjust the spectral information of the new dry voice signal based on the formant information of the original dry voice signal, so that the pitch-changed sound signal has a high original vocal fidelity effect, that is, the timbre is highly maintained, and the influence of vocal distortion caused by song pitch-changing on the listening experience is avoided.

[0047] Based on the above example description, some specific possible implementation examples will be provided below. In practical applications, the implementation contents between these examples can be combined as needed according to the corresponding functional principles and application logics.

[0048] Please refer to Figures 3 to 6 , another specific embodiment of the song processing method provided by this application, the embodiment includes the following operation steps:

[0049] 30. Separate the original accompaniment signal and the original dry voice signal from the original song.

[0050] The original song (or called the original singer) often mixes with the accompaniment sound and the dry voice signal, that is, there is x msc (n)=x voc (n)+x acc (n), x msc (n), x voc (n), x acc (n) can respectively represent the original singer, the dry voice signal and the accompaniment signal; Therefore, in order to better perform corresponding pitch shifting and subsequent mixing processing on the song, and avoid the distortion sound effect similar to the Minions' tone when the dry voice is uniformly pitch shifted as the accompaniment sound in the traditional scheme. Therefore, as Figure 4 shown, the accompaniment track and the singing track, that is, the original accompaniment signal x acc (n) and the original dry voice signal x voc (n), can be separated from the original song by using a voice and accompaniment separation tool, such as the open-source spleeter tool or the self-developed separation tool, and then the two are respectively pitch shifted and other processed, and finally mixed to obtain the required high-quality pitch shifted song.

[0051] 31. Obtain the original dry voice signal in the original song.

[0052] I. Perform pitch shifting and timbre preservation processing on the dry voice signal (please refer to steps 32 to 33 for details):

[0053] 32. Perform variable speed without pitch shifting and resampling processing on the original dry voice signal to change the pitch characteristics of the original dry voice signal to obtain a new dry voice signal.

[0054] In actual operation, the essence of audio pitch shifting is to change its frequency, that is, frequency conversion or frequency shift; The pitch shifting in the embodiment of this application can be realized by connecting two processing modules of variable speed without pitch shifting + resampling in series. Specifically, please refer to the pitch shifting processing sequence table in Table 1. Pscale in Table 1 represents the pitch change degree Pitch-shift-Scale, which can be expressed by the log relationship. Pscale greater than 0 can represent pitch up, and less than 0 can represent pitch down; The resampling processing can be specifically divided into downsampling or upsampling according to the effect.

[0055] Ascending tone (Pscale > 0) Stretching or expanding frame segments → Downsampling Descending tone (Pscale < 0) Upsampling → Shortening or compressing frame segments

[0056] Table 1 Pitch shifting processing sequence table

[0057] (1) In the face of the pitch-up demand, that is, if a target dry voice signal with a higher pitch than the original dry voice signal is to be obtained, step 32 may specifically include the following operation steps:

[0058] Use the signal frame expansion algorithm to stretch the frame segments of the original dry voice signal to obtain a variable-speed dry voice signal that is decelerated but not pitch-changed relative to the original dry voice signal; downsample the time-frequency point signals in the variable-speed dry voice signal to obtain a target dry voice signal with the same frame segment length as the original dry voice signal but with a higher pitch.

[0059] Exemplarily, any one of the signal frame expansion algorithms such as ESOLA, WSOLA, phasevocoder, TSM_Toolbox_MATLAB, and rubberband can be used to achieve variable speed without pitch change of the voice signal. As Figure 5 shown, for the original dry voice signal composed of 10 frame segments, the ESOLA algorithm can be used to splice the replicated frames of each frame behind each frame to obtain a 20-frame stretched or expanded signal. The obtained 20-frame signal is a variable-speed dry voice signal that is decelerated but not pitch-changed relative to the previous 10-frame signal. Then, downsample this 20-frame signal, which is equivalent to discarding the repeated frames and retaining the valid samples to the greatest extent, and a target dry voice signal with the same frame segment length as the original dry voice signal but with a higher pitch can be obtained. In short, the frame segment length of the original dry voice signal can be doubled first and then shortened back to the original frame segment length in the reverse direction, so as to comprehensively achieve the purpose of finally raising the pitch of the voice signal. Among them, a sample can be considered as a discretized representation of an analog audio signal, and a valid sample is a quantized signal that retains the original audio information to a certain extent or to the greatest extent; the series use of the two processing modules of variable speed without pitch change and resampling in (1) can be regarded as stretching and then compressing the frame segments in combination. The final frame segment length remains unchanged, or it can be regarded as one module decelerating and the other module accelerating in the opposite direction to restore the sound speed, so the sound speed of the voice signal is not finally changed, and the audio pitch change can be jointly achieved while the sound speed can be maintained at the original speed.

[0060] (2) In the face of the pitch-down demand, that is, if a target dry voice signal with a lower pitch than the original dry voice signal is to be obtained, step 32 may specifically include the following operation steps:

[0061] Upsample the time-frequency point signals in the original dry voice signal to obtain a variable-pitch dry voice signal that is longer than the frame segment of the original dry voice signal but with a lower pitch; use the signal frame compression algorithm to shorten the frame segments of the variable-pitch dry voice signal to obtain a target dry voice signal with the same frame segment length as the original dry voice signal but with a lower pitch.

[0062] It should be noted that the above algorithms such as ESOLA, WSOLA, phasevocoder, TSM_Toolbox_MATLAB, and rubberband are all based on the TSM scheme and can be specifically divided into those for expanding signal frames or compressing signal frames according to the application effect. Therefore, in the face of the need for pitch lowering, any of the above algorithms can also be used as the signal frame compression algorithm. Compared with the pitch raising operation in (1) above, the ESOLA algorithm can be used to upsample the original dry voice signal composed of 10-frame segments. This is equivalent to inserting a duplicate frame of each frame behind each frame (equivalent to decelerating), and then compressing the 20-frame signal obtained by the insertion to the original frame segment length (equivalent to accelerating back to the original speed) to comprehensively achieve the final purpose of lowering the pitch of the voice signal. Of course, in some specific examples, the Figure 5 execution order of the TSM algorithm (for changing speed without changing pitch) module and the resample module shown in Table 1 can be unrestricted, but in the face of the actual needs of pitch raising or lowering, preferably the execution order shown in Table 1, Figure 5 as shown, is adopted.

[0063] 33. Adjust the formant distribution of the new dry voice signal based on the formant information of the original dry voice signal to obtain the target dry voice signal with corrected timbre.

[0064] In some specific examples, the formant information can largely determine or affect the timbre of the voice signal. Therefore, step 33 can specifically include the following operation steps (which can be collectively referred to as formant processing or timbre preservation processing):

[0065] Compare the formant information in the spectra between the original dry voice signal and the new dry voice signal, and construct the weight coefficients for each time-frequency point of the new dry voice signal according to the formant comparison result; use each weight coefficient to weight the new dry voice signal at the corresponding time-frequency point to correct and obtain the target dry voice signal that conforms to the formant distribution of the original dry voice signal and does not change the pitch characteristics of the new dry voice signal.

[0066] In actual situations, the human voice or dry voice mainly determines the voice characteristics of the singer, that is, the timbre, through the resonance and sound absorption of the vocal tract. The formant information of the audio signal, such as the resonant frequency and spectral envelope, can directly reflect or affect the content of the audio signal and the timbre characteristics of the sound source. Taking the female vocalization audio "ah a" with a sampling rate of 32 kHz and a duration of about 2 s as an example, the relative position of its formants in the spectrogram is described as Figure 6 shown in (a) below, and the audio power and its formant envelope curve plotted based on frame analysis are shown as Figure 6 shown in (b) below. It can be seen that the formant curve ( Figure 6 the thinner curve above (b) below) reflects the actual current signal power spectrum ( Figure 6The envelope characteristics of the lower dot trace curve shown in (b) of the Chinese figure, that is, the distribution of formants can completely reflect the spectral envelope characteristic curve (or simply referred to as the envelope curve); of course, from Figure 6 It can be seen that the distribution characteristics between the envelope curve and the formants can be expressed mutually. Therefore, to maintain the timbre characteristics of the output signal after pitch shifting, that is, to retain the (short-time) spectral envelope of the input signal.

[0067] Precisely because of this, the comparison result between the formant information of the audio signal (or called the sound signal) before and after pitch shifting can be used to construct the mask weight, that is, the masking parameter, at each time-frequency point. This mask weight can be regarded as a weight coefficient with a value between 0 and 1. Then, based on this mask weight sequence in the time-frequency domain, the frequency-converted signal, that is, the new dry sound signal, can be weighted and output to obtain the original timbre signal with natural preservation. Among them, it can be understood that performing a weighted operation on the new dry sound signal at each time-frequency point is equivalent to correcting or restoring the formant curve (the object to be changed) of the new dry sound signal to the same distribution as the formant position (the reference object) of the original dry sound signal, and the frequency of the sound signal is not changed during this process, that is, no pitch shifting occurs. Or it can be understood that by correcting the formants to compensate for the envelope curve of the frequency-converted signal, so that the frequency-converted signal (that is, the signal after pitch shifting) can also be adaptively named or close to the envelope curve expression of the original dry sound signal, so as to achieve the traditional pitch shifting while retaining the timbre characteristics of the singer itself for the pitch correction effect.

[0068] The target dry sound signal obtained by the above processing steps 32 to 33 can be expressed as y voc (n).

[0069] II. Perform pitch shifting on the accompaniment signal (please refer to step 34 for details):

[0070] 34. Perform speed change without pitch shifting and resampling on the original accompaniment signal to change the pitch characteristics of the original accompaniment signal to obtain a new accompaniment signal.

[0071] Since the accompaniment track x acc (n) often contains a percussion part and a harmonic part, and the two can be represented separately by That is, there is Therefore, in some specific examples, based on the Harmonic-Percussion-Separation (HPS) framework, the following operations can be carried out to specifically implement step 34:

[0072] Separate the original percussion signal and the original harmonic signal from the original accompaniment signal, and perform speed change without pitch shifting and resampling on both the original percussion signal and the original harmonic signal; mix the newly processed percussion signal and harmonic signal to obtain a new accompaniment signal.

[0073] Among them, specifically, the percussion signal and the harmonic signal can be distinguished and separated through the median matrix equivalent to the scanning window. For the original percussion signal The pitch change can be realized by using the waveform similarity superposition WSOLA algorithm (one of the TSM algorithms) + the resampling resample scheme. The new percussion signal obtained by such processing can be expressed as For the original harmonic signal The pitch change can be realized by using the phase vocoder phasevocoder + resample scheme. The new harmonic signal obtained by such processing can be expressed as The new accompaniment signal obtained after mixing with pitch change can be expressed as

[0074] It should be mentioned that the accompaniment signal is often the mixed audio of multiple instrument sounds with rich components. This also results in the lack of the timbre characteristics of the instruments in the accompaniment signal itself, that is, there is no concept of timbre in the accompaniment signal. Therefore, it is not necessary to perform timbre preservation or timbre restoration processing on the accompaniment signal like the dry voice signal (i.e., the operation in step 33). However, to avoid the distortion of the human voice caused by simple pitch change, it is still necessary to perform timbre preservation processing on the dry voice signal to restore the vocal tract timbre characteristics of the singer itself.

[0075] 35. Mix the target dry voice signal and the new accompaniment signal to obtain the target song.

[0076] After performing the above-mentioned steps 32 to 34 on the original dry voice signal and the original accompaniment signal, the target dry voice signal and the new accompaniment signal obtained by the processing can be mixed to obtain the target song with pitch change and dry voice timbre preservation.

[0077] In some specific examples, step 35 may specifically include the following operation steps:

[0078] According to the signal power corresponding to each time-frequency point in the accompaniment signal before and after pitch change, calculate the accompaniment power spectrum parameters before and after pitch change respectively; according to the signal power corresponding to each time-frequency point in the dry voice signal before and after timbre correction, calculate the dry voice power spectrum parameters before and after timbre correction respectively; based on the preset proportional relationship value or the preset logarithmic relationship difference between the accompaniment power spectrum parameters before and after pitch change, determine the accompaniment superposition weight; based on the preset proportional relationship value or the preset logarithmic relationship difference between the dry voice power spectrum parameters before and after timbre correction, determine the dry voice superposition weight; use the dry voice superposition weight and the accompaniment superposition weight to perform weighted mixing on the target dry voice signal and the new accompaniment signal.

[0079] Exemplarily, the calculation representation of the power spectrum parameters can be as follows:

[0080] Among them, u can be selected as "voc" or "acc", which respectively correspond to the separated dry voice signal and the separated accompaniment signal. L represents the length or the number of samples of the audio sequence (the number of samples is the digital representation of the song time length), and the value range of n is [0, L - 1]. The target dry voice signal obtained by pitch shifting and the new accompaniment signal can be superimposed with weights according to the original ratio and remixed into the pitch-shifted song, that is, the target song.

[0081] The superimposed weight of the mixing can be expressed as a proportional relationship value used to reflect the energy ratio:

[0082] Among them, u can be selected as "voc" or "acc", which respectively correspond to the separated dry voice signal and the separated accompaniment signal; x and y respectively represent the signals before and after pitch shifting. It should be noted that the dry voice signal after timbre correction or pitch shifting mentioned in step 35 refers to the target dry voice signal obtained through the combined processing of steps 32 and 33.

[0083] The superimposed weight of the mixing can also be expressed as the difference of the logarithmic relationship:

[0084] Among them, dB represents the decibel value of the audio signal.

[0085] As above, the finally obtained target song can be expressed as y(n) = w voc ·y voc (n) + w acc ·y acc (n).

[0086] The operation contents of the above steps 34 and step 32 in the face of the need for pitch up or down are similar, and the operation contents of steps 31 to 33 and steps 21 to 23 are similar. Specifically, they will not be elaborated here; the execution order of step 34 and step 32 or 33 is not limited and can be executed simultaneously.

[0087] In contrast, the existing mainstream pitch adjustment schemes will cause obvious timbre distortion effects in the vocal or dry voice signal part after the pitch of the song deviates by more than 1 semitone, such as the voice-changing effect like that of a Minion, deviating from the original timbre characteristics of the singer's vocal tract; while the pitch shifting strategy provided by the embodiments of the present application can ensure that the pitch shifting of the song is within 7 semitones (that is, within the comfortable pitch shifting range of up and down 7 keys), and at the same time has a high fidelity of the original singing voice, that is, the timbre preservation effect, and the maximum pitch shifting range that can be covered can reach one octave up and down.

[0088] Please refer to Figure 7, the electronic device 700 according to the second aspect of the embodiments of the present application may include one or more central processing units (CPUs) 701 and a memory 705, and one or more applications or data are stored in the memory 705.

[0089] Among them, the memory 705 may be volatile storage or persistent storage. The program stored in the memory 705 may include one or more modules, and each module may include a series of instruction operations on the electronic device. Further, the central processing unit 701 may be configured to communicate with the memory 705 and execute a series of instruction operations in the memory 705 on the electronic device 700.

[0090] The electronic device 700 may further include one or more power supplies 702, one or more wired or wireless network interfaces 703, one or more input / output interfaces 704, and / or one or more operating systems, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0091] The central processing unit 701 may perform the operations executed by the foregoing first aspect or any specific method embodiment of the first aspect, and details are not described herein again.

[0092] A computer-readable storage medium provided by the present application includes instructions, and when the instructions run on a computer, the computer is caused to execute the method described in the foregoing first aspect or any specific implementation manner of the first aspect.

[0093] A computer program product provided by the present application includes instructions or a computer program, and when the computer program product runs on a computer, the computer is caused to execute the method described in the foregoing first aspect or any specific implementation manner of the first aspect.

[0094] It can be understood that in various embodiments of the present application, the sequence numbers of the steps do not mean the order of execution, and the execution order of each step should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0095] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system (if any) and device can refer to the corresponding processes in the foregoing method embodiments, and details are not described herein again.

[0096] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system or device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.

[0097] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0098] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0099] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product (computer program product) is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a business server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks or optical discs and other various media that can store program codes.

Claims

1. A method for processing a song, characterized in that, comprising: obtaining an original dry voice signal in the original song; performing pitch-shifting without changing the pitch and resampling on the original dry voice signal to change the pitch characteristics of the original dry voice signal to obtain a new dry voice signal; adjusting the formant distribution of the new dry voice signal based on the formant information of the original dry voice signal to obtain a target dry voice signal with corrected timbre; The correcting the timbre of the new dry voice signal obtained by pitch-shifting based on the formant information of the original dry voice signal includes: comparing the formant information of the spectra between the original dry voice signal and the new dry voice signal, and constructing weight coefficients for each time-frequency point of the new dry voice signal according to the formant comparison result; using each of the weight coefficients to weight the new dry voice signal at the corresponding time-frequency point to correct and obtain a target dry voice signal that conforms to the formant distribution of the original dry voice signal and does not change the pitch characteristics of the new dry voice signal; if the original song includes an accompaniment signal and a dry voice signal, separating the original accompaniment signal and the original dry voice signal from the original song; performing pitch-shifting without changing the pitch and resampling on the original accompaniment signal to change the pitch characteristics of the original accompaniment signal to obtain a new accompaniment signal; mixing the target dry voice signal and the new accompaniment signal to obtain a target song.

2. The song processing method according to claim 1, characterized in that, if a target dry voice signal with a higher pitch relative to the original dry voice signal is to be obtained, the process of performing pitch-shifting without changing the pitch and resampling on the original dry voice signal includes: using a signal frame expansion algorithm to stretch the frame segment of the original dry voice signal to obtain a pitch-shifted dry voice signal that is decelerated but does not change the pitch relative to the original dry voice signal; downsampling the time-frequency point signals in the pitch-shifted dry voice signal to obtain a target dry voice signal that has the same frame segment length as the original dry voice signal but has a higher pitch.

3. The song processing method according to claim 1, characterized in that, if a target dry voice signal with a lower pitch relative to the original dry voice signal is to be obtained, the process of performing pitch-shifting without changing the pitch and resampling on the original dry voice signal includes: upsampling the time-frequency point signals in the original dry voice signal to obtain a pitch-shifted dry voice signal that is longer than the frame segment of the original dry voice signal but has a lower pitch; using a signal frame compression algorithm to shorten the frame segment of the pitch-shifted dry voice signal to obtain a target dry voice signal that has the same frame segment length as the original dry voice signal but has a lower pitch.

4. The song processing method according to claim 1, characterized in that, the process of performing pitch-shifting without changing the pitch and resampling on the original accompaniment signal includes: separating the original percussion signal and the original harmonic signal from the original accompaniment signal, and performing pitch-shifting without changing the pitch and resampling on both the original percussion signal and the original harmonic signal; mixing the new percussion signal and the new harmonic signal obtained by the final processing to obtain a new accompaniment signal.

5. The song processing method according to claim 1, characterized in that, the mixing of the target dry voice signal and the new accompaniment signal includes: According to the signal power corresponding to each time-frequency point in the accompaniment signal before and after pitch shifting, calculate the accompaniment power spectrum parameters before and after pitch shifting respectively; according to the signal power corresponding to each time-frequency point in the dry voice signal before and after timbre correction, calculate the dry voice power spectrum parameters before and after timbre correction respectively; Based on the preset proportional relationship value or the preset logarithmic relationship difference between the accompaniment power spectrum parameters before and after pitch shifting, determine the accompaniment superposition weight; based on the preset proportional relationship value or the preset logarithmic relationship difference between the dry voice power spectrum parameters before and after timbre correction, determine the dry voice superposition weight; Use the dry voice superposition weight and the accompaniment superposition weight to perform weighted mixing on the target dry voice signal and the new accompaniment signal.

6. The song processing method according to claim 1, wherein, Use the WSOLA, phasevocoder or ESOLA signal frame algorithm to perform the pitch-invariant speed change processing on the signal.

7. An electronic device, wherein, comprising: a central processing unit, a memory and an input / output interface; The memory is a transient storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute the instruction operations in the memory to execute the method according to any one of claims 1 to 6.

8. A computer-readable storage medium, wherein, comprising instructions, when the instructions run on a computer, cause the computer to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Audio processing method and device and readable storage medium

    CN113113033A

  • Audio processing method and computer device

    CN114220409A