Audio synthesis methods, computer equipment, readable storage media, and program products
By splicing and pitch-shifting audio signals and adjusting wave-blocking characteristic parameters, the problem of existing audio synthesis methods relying on sound source sampling data is solved, thereby improving the flexibility of audio synthesis.
Patent Information
- Application Number
- CN202411710816.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-11-27
AI Technical Summary
Existing audio synthesis methods rely on massive amounts of audio source sampling data, resulting in low flexibility in audio synthesis.
By obtaining the target audio signal and melody feature information of the target timbre, splicing and pitch shifting processing are performed, and the envelope feature parameters are adjusted to generate audio that meets the target melody.
It can flexibly synthesize target audio that meets any duration and pitch requirements without relying on massive audio source sampling data, thus improving the flexibility of audio synthesis.
Smart Images

Figure CN119763589B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio technology, and in particular to an audio synthesis method, computer device, readable storage medium, and program product. Background Art
[0002] With the development of audio technology, audio synthesis has been widely applied in various fields such as modern music composition, film scoring, and voice synthesis for smart devices. Audio synthesis technology primarily uses computers to simulate and generate sound, thereby meeting the audio needs of different scenarios. Especially in music production, audio synthesis can provide music creators with a rich selection of timbres without relying on the performance of real instruments. Therefore, audio synthesis technology can greatly expand creative possibilities, reduce production costs, and improve work efficiency.
[0003] Currently, audio synthesis methods mainly include software-based audio source rendering and sampler-based rendering. Software-based audio source rendering uses software-based audio synthesizers to synthesize audio, employing digital signal processing techniques such as frequency modulation synthesis, wavetable synthesis, and subtractive synthesis. Sampler-based rendering uses audio synthesizers based on recorded samples to generate sound by pre-playing recorded audio samples and then adjusting the samples through pitch, time stretching, filtering, and other processing. However, all of these current audio synthesis methods rely on massive amounts of audio source sampling data to complete the synthesis, resulting in very low flexibility in audio synthesis. Summary of the Invention
[0004] Therefore, it is necessary to provide an audio synthesis method, computer device, readable storage medium, and program product that can improve the flexibility of audio synthesis in response to the above-mentioned technical problems.
[0005] In a first aspect, this application provides an audio synthesis method, including:
[0006] Acquire the target audio signal of the target timbre, and acquire the melody feature information of the target melody, wherein the melody feature information includes melody duration and melody pitch;
[0007] Based on the duration of the melody and the duration of the target audio signal, several target audio signals are spliced together to obtain a spliced timbre audio signal, and the spliced timbre audio signal is then subjected to pitch shifting processing based on the melody pitch to obtain an adjusted timbre audio signal;
[0008] Based on the duration of the melody, the duration of at least one wave-sealing characteristic parameter of the target audio signal is adjusted to obtain the adjusted wave-sealing characteristic parameter; the wave-sealing characteristic parameter of the target audio signal is used to reflect the change of the amplitude of the target audio signal within the duration of the wave-sealing characteristic parameter;
[0009] Based on the adjusted timbre audio signal and the adjusted wave sealing characteristic parameters, a target audio is generated; the target audio is a sound audio that possesses the target timbre and conforms to the target melody.
[0010] In one embodiment, the step of splicing together several target audio signals to obtain a spliced audio signal includes:
[0011] Several target audio signals are spliced together end to end, with the two target audio signals overlapping to obtain a spliced audio signal with the same melody duration as the target melody.
[0012] For the overlapping portion in the spliced audio signal, the end portion of the overlapping portion belonging to the target audio signal is faded out, and the beginning portion of the overlapping portion belonging to the target audio signal is faded in, to obtain the spliced audio signal.
[0013] In one embodiment, the step of performing pitch shifting processing on the spliced timbre audio signal according to the melody pitch to obtain an adjusted timbre audio signal includes:
[0014] The resampling coefficients are determined based on the pitch difference between the melody pitch and the pitch of the spliced timbre audio signal;
[0015] The spliced timbre audio signal is resampled according to the resampling coefficient to obtain the adjusted timbre audio signal, so that the pitch of the adjusted timbre audio signal is consistent with the pitch of the melody, and the playback rate of the adjusted timbre audio signal is consistent with that of the spliced timbre audio signal.
[0016] In one embodiment, acquiring the target audio signal for the target timbre includes:
[0017] The original audio signal of the target timbre and multiple wave-sealing feature parameters corresponding to the original audio signal are obtained; the wave-sealing feature parameters of the original audio signal are used to reflect the change of the amplitude of the original audio signal during the duration of the wave-sealing feature parameters;
[0018] Based on the ratio between the amplitude change and the duration corresponding to each wave sealing feature parameter, wave sealing feature parameters whose ratio is less than a preset ratio threshold are determined from the plurality of wave sealing feature parameters as target wave sealing feature parameters;
[0019] The target audio signal with the target timbre is generated based on the audio signal segment corresponding to the target wave sealing feature parameter in the original audio signal.
[0020] In one embodiment, generating the target audio signal with the target timbre based on the audio signal segment corresponding to the target wave blocking feature parameter in the original audio signal includes:
[0021] If the average amplitude of the audio signal segment is less than a preset amplitude threshold, the audio signal segment is amplified to obtain the target audio signal with the target timbre.
[0022] If the amplitude fluctuation frequency of the audio signal segment is greater than a preset frequency threshold, the audio signal segment is compressed to obtain the target audio signal with the target timbre.
[0023] In one embodiment, adjusting the duration of at least one wave-sealing feature parameter of the target audio signal according to the duration of the melody to obtain the adjusted wave-sealing feature parameter includes:
[0024] Determine the sum of the durations of the at least one wave-blocking characteristic parameter;
[0025] The duration of the at least one wave-sealing feature parameter is scaled according to the duration ratio between the sum of the melody duration and the duration, so that the sum of the durations of the at least one wave-sealing feature parameter after scaling is equal to the melody duration, thus obtaining the adjusted wave-sealing feature parameter.
[0026] In one embodiment, generating the target audio based on the adjusted timbre audio signal and the adjusted waveband feature parameters includes:
[0027] The amplitude of the adjusted timbre audio signal is multiplied point-by-point at each time point with the amplitude included in the adjusted wave sealing characteristic parameters to generate the target audio.
[0028] Secondly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.
[0029] Thirdly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.
[0030] Fourthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0031] The aforementioned audio synthesis method, apparatus, computer equipment, computer-readable storage medium, and computer program product acquire a target audio signal with a target timbre and acquire melodic feature information of a target melody, wherein the melodic feature information includes melody duration and melody pitch; based on the melody duration and the duration of the target audio signal, several target audio signals are spliced to obtain a spliced timbre audio signal, and the spliced timbre audio signal is pitch-shifted according to the melody pitch to obtain an adjusted timbre audio signal; based on the melody duration, the duration of at least one wave-sealing feature parameter of the target audio signal is adjusted to obtain an adjusted wave-sealing feature parameter, wherein the wave-sealing feature parameter of the target audio signal is used to reflect the change in amplitude of the target audio signal within the duration of the wave-sealing feature parameter; based on the adjusted timbre audio signal and the adjusted wave-sealing feature parameter, a target audio is generated, wherein the target audio is a sound audio with a target timbre and conforming to a target melody. By splicing and pitch-shifting the target audio signal, the target audio signal can be flexibly adjusted according to the melodic characteristics of the target melody. At the same time, the duration of at least one wave-sealing characteristic parameter of the target audio signal can be flexibly adjusted according to the duration of the melody. Finally, based on the adjusted timbre and the adjusted wave-sealing characteristic parameter, a target audio with the target timbre and conforming to the target melody can be synthesized. It does not rely on a large amount of audio source sampling data, but only requires a small amount of target audio signal with the target timbre to synthesize target audio that meets the requirements of any duration and pitch, thus improving the flexibility of audio synthesis. Attached Figure Description
[0032] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0033] Figure 1 This is a diagram illustrating the application environment of an audio synthesis method in one embodiment.
[0034] Figure 2 This is a flowchart illustrating an audio synthesis method in one embodiment;
[0035] Figure 3 This is a schematic diagram of an envelope model in one embodiment;
[0036] Figure 4 This is a schematic diagram of an envelope in another embodiment;
[0037] Figure 5 This is a flowchart illustrating another audio synthesis method in one embodiment;
[0038] Figure 6 This is a structural block diagram of an audio synthesis device in one embodiment;
[0039] Figure 7 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0040] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0041] The audio synthesis method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on a cloud or other network server. Terminal 102 acquires the target audio signal of the target timbre and the melody feature information of the target melody, including melody duration and melody pitch. Based on the melody duration and the duration of the target audio signal, terminal 102 splices several target audio signals to obtain a spliced timbre audio signal, and performs pitch shifting processing on the spliced timbre audio signal according to the melody pitch to obtain an adjusted timbre audio signal. Based on the melody duration, terminal 102 adjusts the duration of at least one wave-sealing feature parameter of the target audio signal to obtain an adjusted wave-sealing feature parameter. The wave-sealing feature parameter of the target audio signal reflects the change in amplitude of the target audio signal within the duration of the wave-sealing feature parameter. Terminal 102 generates the target audio based on the adjusted timbre audio signal and the adjusted wave-sealing feature parameter. The target audio is a sound audio with the target timbre and conforming to the target melody. The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle systems, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. The server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0042] In an exemplary embodiment, Figure 2As shown, an audio synthesis method is provided, which can be applied to... Figure 1 Taking terminal 102 as an example, the explanation includes:
[0043] Step S202: Obtain the target audio signal of the target timbre and the melodic feature information of the target melody.
[0044] Among them, the melody feature information includes melody duration and melody pitch.
[0045] The target timbre can refer to the timbre of any sound-producing body, such as the timbre of any musical instrument like a violin or piano, the timbre of any animal like a dog or cat, or even the timbre of a particular person. The target timbre is the specific timbre that is expected to be simulated or synthesized during the audio synthesis process.
[0046] The target audio signal for the desired timbre can be an audio signal synthesized using a synthesizer or an actual recorded audio signal. For example, it can be an audio signal for a violin timbre directly synthesized using a synthesizer, or it can be an audio signal obtained by recording during a violin performance.
[0047] The embodiments of this application can be applied to synthesizing high-quality target audio in scenarios where sound sources are scarce. Therefore, in specific implementations, the duration of the target audio signal with the target timbre can be a relatively short duration such as 1 second, 2 seconds, or 5 seconds.
[0048] Taking the timbre of a musical instrument as an example, the target audio signal can be the audio signal generated by the instrument playing any note for several seconds. In other words, the target audio signal can include only the sound of a single tone emitted by any sound-producing body.
[0049] The target melody can refer to the expected melodic structure in the audio to be synthesized. Melodic feature information can refer to detailed descriptions of the melody, including pitch (frequency of notes), melody tone, melody duration, duration (duration of each note), rhythm (time interval between notes), and other information.
[0050] Step S204: Based on the duration of the melody and the duration of the target audio signal, several target audio signals are spliced together to obtain a spliced timbre audio signal. The spliced timbre audio signal is then pitch-shifted according to the melody pitch to obtain an adjusted timbre audio signal.
[0051] In practice, the terminal can splice several target audio signals according to the melody duration to obtain a spliced timbre audio signal, ensuring that the duration of the spliced timbre audio signal matches the duration of the target melody. Optionally, when splicing several target audio signals, the signals can be non-overlapping or overlapped at the beginning and end of the splice. For example, if the target melody lasts 15 seconds and the target audio signal lasts 5 seconds, then at least three target audio signals are needed for splicing.
[0052] In specific implementation, the terminal performs pitch shifting processing on the spliced timbre audio signal according to the melody pitch. This can be achieved by changing the frequency of the spliced timbre audio signal to adjust its pitch, resulting in an adjusted timbre audio signal whose pitch matches the melody pitch of the target melody. Optionally, pitch shifting processing can include time-domain methods, frequency-domain methods, and parametric methods. Time-domain methods can change the pitch using resampling techniques without altering the playback speed; frequency-domain methods can analyze the frequency components of the audio using Fourier transform, adjust the pitch, and then perform an inverse transform.
[0053] Step S206: Based on the melody duration, adjust the duration of at least one wave-sealing characteristic parameter of the target audio signal to obtain the adjusted wave-sealing characteristic parameter.
[0054] Among them, the wave sealing characteristic parameter of the target audio signal is used to reflect the change of the amplitude of the target audio signal during the duration of the wave sealing characteristic parameter.
[0055] A wave seal refers to the amplitude profile of an audio signal, consisting of a series of peaks and troughs. It can be viewed as an envelope, representing the outer contour of the audio signal's amplitude as it changes over time. Wave seal characteristic parameters are parameters that quantify the wave seal's properties, and may include amplitude, duration, etc.
[0056] For example, a corresponding envelope model (ADSR model) can be constructed for a specific target timbre. Different timbres and different fundamental frequencies have different harmonic patterns in the frequency domain, which is the basic basis for human hearing to distinguish different musical instruments. When synthesizing sound, if the fundamental frequency and harmonic patterns of the instrument in the frequency domain, as well as the curve changes of the envelope model in the time domain, can be reproduced, then audio simulating the timbre of that instrument can be synthesized.
[0057] For the convenience of those skilled in the art, Figure 3An exemplary schematic diagram of an envelope model is provided. The envelope model is used to describe the characteristics of the amplitude variation of an audio signal over time, and can include the four stages of sound production: attack (A), decay (D), sustain (S), and release (R). When a musical instrument produces sound, its characteristics change continuously over time, containing both periodic and non-periodic components. In the initial stage, energy typically increases, i.e. Figure 3 In the A phase of the sound production, the sound gradually intensifies, containing a large number of non-periodic components distributed across the entire frequency range. After the attack (A) phase, the sound stabilizes (D phase) and reaches a stable phase (S phase) with more or less a periodic pattern. The S phase accounts for the majority of the duration, during which the audio energy may remain constant or gradually decrease. The final phase of the instrument's sound production is the R phase, where the sound gradually fades away. Envelope models can provide a highly consistent fit to the amplitude envelope of the instruments that produce the pitch.
[0058] The wave-sealing feature parameters of the target audio signal can include feature parameters of any articulation stage in the envelope model of the target audio signal. Since the amplitude of the target audio signal varies differently within the duration of the wave-sealing feature parameters in different articulation stages, the wave-sealing feature parameters corresponding to different articulation stages are different. For example, the wave-sealing feature parameters of the target audio signal can include the amplitude and duration of any articulation stage in the envelope model of the target audio signal. For instance, at least one wave-sealing feature parameter of the target audio signal can include the amplitude and duration of at least one of the following: stage A in the ADSR model, stage D in the ADSR model, stage S in the ADSR model, and stage R in the ADSR model.
[0059] As an example, the envelope model of the target audio signal can be determined by calculating the upper envelope, lower envelope, and amplitude envelope of the target audio signal using a sliding window on the waveform and finding extrema within each sliding window. Here, the upper envelope represents the upper boundary of the target audio signal; the lower envelope represents the lower boundary of the target audio signal; and the amplitude envelope is the curve showing the change in amplitude of the target audio signal, representing the maximum value of the target audio signal's amplitude.
[0060] Specifically, each sampling point of the target audio signal can be traversed, and based on a sliding window of a preset window length, the maximum value of the amplitude, the maximum value of the upper envelope, and the minimum value of the lower envelope within the current sliding window can be calculated, thereby obtaining three types of envelopes, namely the amplitude envelope, the upper envelope, and the lower envelope.
[0061] Alternatively, the four articulation stages of the envelope model can be represented by w and f(x), where w represents the proportion of the duration of that articulation stage to the total duration, and f(x) can be used to represent the functional fitting relationship of that articulation stage. For example, the amplitude recorded in the envelope model ranges from 0 to 1, or from -1 to 1.
[0062] In a specific implementation, adjusting the duration of at least one wave-sealing feature parameter of the target audio signal to obtain an adjusted wave-sealing feature parameter can be achieved. This can include adjusting the duration of each wave-sealing feature parameter of the target audio signal and using all the adjusted wave-sealing feature parameters as a whole as the adjusted wave-sealing feature parameter; or, it can include adjusting the duration of one of the wave-sealing feature parameters of the target audio signal and using that individually adjusted wave-sealing feature parameter as the adjusted wave-sealing feature parameter. The duration of the adjusted wave-sealing feature parameter is equal to the melody duration of the target melody.
[0063] Optionally, the durations of each wave-sealing feature parameter can be adjusted proportionally so that the sum of the durations of each wave-sealing feature parameter after adjustment is equal to the duration of the target melody; or, the duration of one of the wave-sealing feature parameters can be adjusted so that the duration of that wave-sealing feature parameter after adjustment is equal to the duration of the target melody.
[0064] The adjusted wave-sealing characteristic parameters can match the duration requirements of the melody and reflect the natural changes in the sound, ensuring that the sound not only has pitch and timbre, but also has a realistic dynamic performance.
[0065] Step S208: Generate the target audio based on the adjusted timbre audio signal and the adjusted wave seal characteristic parameters.
[0066] The target audio is a sound audio that has the target timbre and matches the target melody.
[0067] The adjusted timbre audio signal provides frequency information (pitch, timbre), while the adjusted wave-sealing characteristic parameters provide amplitude variation information (dynamic characteristics). The combination of these two ensures that the generated target audio not only possesses correct timbre and pitch but also realistic and natural dynamic fluctuations. In practice, the terminal generates the target audio based on the adjusted timbre audio signal and the adjusted wave-sealing characteristic parameters. This can be achieved by multiplying the adjusted timbre audio signal and the adjusted wave-sealing characteristic parameters point by point, ensuring that the amplitude of the adjusted timbre audio signal is adjusted by the adjusted wave-sealing characteristic parameters at each time point, thereby generating target audio with natural dynamic performance.
[0068] Based on a small amount of audio source sampling data, this application proposes a lightweight timbre synthesis method that combines pitch transfer and time scaling through splicing and pitch shifting. The method only requires the sound source of the target timbre to emit or play a single note. By collecting this note, the target timbre can be used to render target audio of different durations with a natural listening experience.
[0069] In the aforementioned audio synthesis method, a target audio signal with a target timbre and melodic feature information of a target melody are acquired, wherein the melodic feature information includes melody duration and melody pitch; based on the melody duration and the duration of the target audio signal, several target audio signals are spliced together to obtain a spliced timbre audio signal, and the spliced timbre audio signal is pitch-shifted according to the melody pitch to obtain an adjusted timbre audio signal; based on the melody duration, the duration of at least one wave-sealing feature parameter of the target audio signal is adjusted to obtain an adjusted wave-sealing feature parameter, wherein the wave-sealing feature parameter of the target audio signal is used to reflect the change in amplitude of the target audio signal within the duration of the wave-sealing feature parameter; based on the adjusted timbre audio signal and the adjusted wave-sealing feature parameter, a target audio is generated, wherein the target audio is a sound audio with the target timbre and conforming to the target melody. By splicing and pitch-shifting the target audio signal, the target audio signal can be flexibly adjusted according to the melodic characteristics of the target melody. At the same time, the duration of at least one wave-sealing characteristic parameter of the target audio signal can be flexibly adjusted according to the duration of the melody. Finally, based on the adjusted timbre and the adjusted wave-sealing characteristic parameter, a target audio with the target timbre and conforming to the target melody can be synthesized. It does not rely on a large amount of audio source sampling data, but only requires a small amount of target audio signal with the target timbre to synthesize target audio that meets the requirements of any duration and pitch, thus improving the flexibility of audio synthesis.
[0070] In another embodiment, splicing several target audio signals to obtain a spliced timbre audio signal includes: splicing several target audio signals end to end, with the two spliced target audio signals partially overlapping, to obtain a spliced audio signal with the same melody duration as the target melody; for the overlapping part in the spliced audio signal, fading out the end part of the overlapping part that belongs to the target audio signal, and fading in the beginning part of the overlapping part that belongs to the target audio signal, to obtain the spliced timbre audio signal.
[0071] The terminal splices several target audio signals end to end, and the spliced two target audio signals overlap to a certain extent at the beginning and end.
[0072] Assuming the duration of the target melody is longer than the duration of a single target audio signal, multiple target audio signals can be spliced together end to end. Then, the beginning and end parts can be faded in and out to create a smooth transition, extending the duration to the duration of the target melody. This ensures that the duration of the generated spliced audio signal is consistent with the duration of the target melody, and the sound quality is natural, avoiding the audio disjointedness caused by direct splicing.
[0073] For overlapping portions generated during the splicing of beginning and end audio signals, a fade-in / fade-out transition process can be applied to these overlapping portions. Specifically, the ending portion of the overlapping portion belonging to the target audio signal is faded out, with the amplitude gradually decreasing until it approaches zero, ensuring a smooth audio transition when the ending portion of any target audio signal transitions to the next target audio signal. The beginning portion of the overlapping portion belonging to the target audio signal is faded in, with the amplitude gradually increasing from zero to a normal level, ensuring a smoother and more natural transition when the beginning portion of any target audio signal connects to the previous target audio signal.
[0074] The technical solution of this embodiment generates a spliced audio signal with a smooth transition by splicing several target audio signals end to end and fading in and out of overlapping parts. This ensures a natural connection between the target audio signals, avoids abrupt volume changes, and ensures that the audio duration of the final spliced audio signal is consistent with the duration of the target melody. This allows the target audio signals to adaptively change their audio length, ensuring that the final synthesized target audio matches the melodic structure of the target melody. Furthermore, the fading in and out processing of the overlapping parts gives the final synthesized target audio a more natural auditory effect.
[0075] In another embodiment, the spliced timbre audio signal is pitch-shifted according to the melody pitch to obtain an adjusted timbre audio signal, including: determining a resampling coefficient based on the pitch difference between the melody pitch and the pitch of the spliced timbre audio signal; and resampling the spliced timbre audio signal according to the resampling coefficient to obtain an adjusted timbre audio signal, so that the pitch of the adjusted timbre audio signal is consistent with the melody pitch, and the playback rate of the adjusted timbre audio signal is consistent with that of the spliced timbre audio signal.
[0076] The resampling coefficient is determined based on the pitch difference between the melody pitch and the pitch of the spliced timbre audio signal. For example, the resampling coefficient... ,in, It is the semitone difference between two pitches.
[0077] Resampling is a technique that alters the sampling rate of an audio signal. By upsampling or downsampling the audio signal, its pitch can be changed without affecting playback speed. During resampling, the density of sampling points in the audio signal changes, which directly affects the signal's frequency (i.e., pitch).
[0078] The resampling factor is a scaling factor used during resampling, determining the magnitude of the change in the audio signal's sampling rate. In the specific implementation, it is assumed that the resampling factor is... Where P is the upsampling coefficient and Q is the downsampling coefficient. The terminal resamples the superimposed audio based on the resampling coefficients. During the upsampling process, insertions are made between adjacent points of the original signal. This sampling point will cause the pitch period to become P times the original, and the spectrum will be compressed to the original value. The duration becomes P times the original. Therefore, the base frequency will become P times the original. The pitch will drop back to its original level. The audio signal's speech rate will be slowed down to its original speed. Times. Similarly, during the downsampling process, every Extracting from each point will cause the pitch period length to change from the original value. The spectrum expands to Q times its original size, while the duration is reduced to the original size. Therefore, the fundamental frequency will increase to Q times its original value, the pitch will increase to Q times its original value, and the speech rate of the audio signal will increase to Q times its original value. Combining the above two processes, the speech rate of the audio signal can be made to return to its original value. The output signal y(n) is obtained by multiplying the values by 1, and then y(n) is processed by... Double sampling is used. This allows you to obtain speech with a normal speaking speed, but with the pitch changed to the original level. The final output speech z(n) is multiplied by a factor of 1. After resampling by a factor of 1, the playback rate can remain unchanged, but the speech rate and pitch of the signal will revert to their original values. times.
[0079] The technical solution of this embodiment resamples the spliced timbre audio signal according to the resampling coefficient. While adjusting the pitch of the spliced timbre audio signal to match the pitch of the target melody, the playback speed remains unchanged. Thus, the resampling process achieves pitch variation without speed variation, thereby improving the naturalness, flexibility and efficiency of audio synthesis and ensuring accurate matching between the generated audio and the expected melody.
[0080] In another embodiment, acquiring the target audio signal of the target timbre includes: acquiring the original audio signal of the target timbre and multiple wave-sealing feature parameters corresponding to the original audio signal; the wave-sealing feature parameters of the original audio signal are used to reflect the change in amplitude of the original audio signal within the duration of the wave-sealing feature parameters; based on the ratio between the amplitude change and the duration corresponding to each wave-sealing feature parameter, determining the wave-sealing feature parameters whose ratio is less than a preset ratio threshold as target wave-sealing feature parameters from the multiple wave-sealing feature parameters; and generating the target audio signal of the target timbre based on the audio signal segment corresponding to the target wave-sealing feature parameter in the original audio signal.
[0081] Optionally, the multiple wave-sealing feature parameters corresponding to the original audio signal can include feature parameters of each articulation stage in the envelope model of the original audio signal, such as the feature parameters of each articulation stage, including onset (A), decay (D), sustain (S), and release (R). These multiple wave-sealing feature parameters reflect the variation pattern of the amplitude of the original audio signal over different time periods. For example, the amplitude may change drastically in one articulation stage, while remaining stable in another. Therefore, the ratio between the amplitude change and duration corresponding to each wave-sealing feature parameter can reflect the drasticness of the amplitude change; a large ratio indicates drastic amplitude change and significant sound fluctuations; a small ratio indicates stable amplitude and a relatively smooth sound. Optionally, the ratio between the amplitude change and duration corresponding to the wave-sealing feature parameter can refer to the ratio between the difference between the maximum and minimum amplitude values and the duration.
[0082] If the ratio is less than a preset ratio threshold, it indicates that the amplitude change reflected by the wave-sealing feature parameter is relatively gradual, meaning the sound is relatively stable. Therefore, wave-sealing feature parameters with ratios less than the preset ratio threshold are used as target wave-sealing feature parameters, and the audio signal segment corresponding to the target wave-sealing feature parameter in the original audio signal is used as the target audio signal for the target timbre. Optionally, since the S-phase of the envelope model is generally a phase with stable amplitude, the wave-sealing feature parameter corresponding to the S-phase is usually a target wave-sealing feature parameter with a ratio less than the preset ratio threshold. Therefore, the audio signal segment corresponding to the S-phase in the original audio signal is used as the target audio signal for the target timbre.
[0083] For the convenience of those skilled in the art, Figure 4 An exemplary schematic diagram of an envelope is provided. It can be seen that... Figure 4In the envelope of the piano sound, stage A has a large amplitude variation range and a short duration, resulting in a large ratio between amplitude variation and duration. Therefore, the wave sealing characteristic parameters of stage A are not suitable as target wave sealing characteristic parameters. In contrast, stage S has a medium amplitude variation range and a long duration, resulting in a smaller ratio between amplitude variation and duration. Therefore, the wave sealing characteristic parameters of stage S are suitable as target wave sealing characteristic parameters. Thus, the audio signal segment corresponding to stage S can be extracted from the envelope of the piano sound as the target audio signal for the target timbre; that is, the audio signal segment corresponding to the target wave sealing characteristic parameters can be used as the target audio signal for the target timbre.
[0084] For example, in the envelope of a violin sound, the duration of phase A is relatively long. If the ratio between the amplitude variation range of phase A and the duration of phase A is less than a preset ratio threshold, then the wave-sealing characteristic parameter of phase A can also be used as the target wave-sealing characteristic parameter. Similarly, phase S has a relatively long duration and a smaller amplitude variation range, so the wave-sealing characteristic parameter of phase S can also be used as the target wave-sealing characteristic parameter. Therefore, audio signal segments corresponding to phases S and A can be extracted from the violin's envelope as the target audio signal for the target timbre; that is, the audio signal segment corresponding to the target wave-sealing characteristic parameter is used as the target audio signal for the target timbre.
[0085] The technical solution of this embodiment determines the audio signal segment corresponding to the target wave-sealing feature parameters from the original timbre audio, which serves as the target audio signal for the target timbre and is used in subsequent audio processing steps (splicing and pitch shifting). This ensures that the target audio signal processed subsequently is sound-stable and exhibits the expected dynamic characteristics. The audio signal segment corresponding to the target wave-sealing feature parameters has a relatively stable amplitude and sound, thus providing a stable audio foundation during audio synthesis and contributing to the generation of smoother and more coherent audio effects.
[0086] In another embodiment, generating a target audio signal for the target timbre based on the audio signal segment corresponding to the target wave blocking feature parameter in the original audio signal includes: if the average amplitude of the audio signal segment is less than a preset amplitude threshold, amplifying the audio signal segment to obtain the target audio signal for the target timbre; if the amplitude fluctuation frequency of the audio signal segment is greater than a preset frequency threshold, compressing the audio signal segment to obtain the target audio signal for the target timbre.
[0087] If the average amplitude is lower than the amplitude threshold, it indicates that the average amplitude of the audio signal segment is not large enough and needs to be amplified. Amplification processing can increase the amplitude of the entire audio signal segment so that the average amplitude is higher than the preset amplitude threshold or as close as possible to the maximum allowable amplitude value.
[0088] Amplitude fluctuation frequency refers to the frequency of amplitude change over a short period of time. A high amplitude fluctuation frequency in an audio signal segment indicates rapid and significant amplitude fluctuations, resulting in noticeable sound instability. Therefore, if the amplitude fluctuation frequency exceeds a preset frequency threshold, it signifies rapid and large fluctuations in the audio signal segment, requiring compression. Compression reduces the amplitude of high-frequency fluctuations in the audio signal segment, making the audio signal smoother, reducing high-frequency amplitude fluctuations, and preventing audio instability.
[0089] The technical solution of this embodiment, through amplification and compression processing, fully amplifies the amplitude of the audio signal segment and controls the dynamic range, so that the target audio signal of the target timbre is a full-amplitude audio signal. That is, the amplitude of the target audio signal makes full use of the available dynamic range and can approach the maximum allowable value, ensuring the loudness enhancement and quality optimization of the target audio signal, and can have a smooth dynamic performance, avoiding sound instability, which helps to improve the quality, clarity and stability of the subsequently generated target audio, and has a high-quality auditory effect.
[0090] In another embodiment, adjusting the duration of at least one wave-sealing feature parameter of the target audio signal according to the melody duration to obtain the adjusted wave-sealing feature parameter includes: determining the sum of the durations of at least one wave-sealing feature parameter; scaling the duration of at least one wave-sealing feature parameter according to the duration ratio between the melody duration and the sum of the durations, so that the sum of the scaled durations of at least one wave-sealing feature parameter is equal to the melody duration, thereby obtaining the adjusted wave-sealing feature parameter.
[0091] Optionally, the sum of the durations included in at least one wave-sealing feature parameter can refer to the length of time from start to finish of the ADSR model, i.e., the time taken for the complete amplitude change process of the target audio signal. The sum of the durations included in at least one wave-sealing feature parameter can be equal to the sum of the durations of each phase of the envelope model (such as onset, decay, hold, and release).
[0092] For example, the sum of the durations of at least one wave-blocking feature parameter can be equal to the sum of the durations of the four phases in the ADSR model of the target audio signal: the duration of phase A, the duration of phase D, the duration of phase S, and the duration of phase R. Assuming the duration of phase A is 1 second, phase D is 2 seconds, phase S is 5 seconds, and phase R is 2 seconds, then the sum of the durations of at least one wave-blocking feature parameter of the target audio signal is equal to 1+2+5+2=10 seconds. If the duration of the target melody is 20 seconds, then the duration ratio between the melody duration and the sum of the durations is 2. Therefore, the durations of each wave-blocking feature parameter can be scaled to twice their original value, so that the sum of the durations of each wave-blocking feature parameter after scaling is equal to 2+4+10+4=20 seconds. That is, the duration of the adjusted wave-blocking feature parameter is equal to 20 seconds, which is equal to the duration of the target melody.
[0093] Scaling the duration of at least one wave-sealing characteristic parameter according to the duration ratio can be achieved by scaling the duration of each wave-sealing characteristic parameter according to the same ratio. This can keep the amplitude dynamic characteristics of at least one wave-sealing characteristic parameter unchanged, and only extend or compress it in time.
[0094] At least one scaled wave-sealing feature parameter is used together as the adjusted wave-sealing feature parameter. The sum of the durations of the scaled at least one wave-sealing feature parameter is equal to the melody duration. That is, the duration of the adjusted wave-sealing feature parameter is equal to the melody duration, ensuring that the amplitude dynamics of at least one wave-sealing feature parameter can be accurately mapped to the time frame of the target melody.
[0095] In another embodiment, generating target audio based on the adjusted timbre audio signal and the adjusted wave seal feature parameters includes: multiplying the amplitude of the adjusted timbre audio signal with the amplitude included in the adjusted wave seal feature parameters at each time point to generate target audio.
[0096] Since the amplitude included in the adjusted wave-sealing feature parameters can refer to the amplitude of any vocalization stage of the envelope model, and the amplitude range is usually between 0 and 1, or between -1 and 1, multiplying the amplitude of the adjusted timbre audio signal with the amplitude included in the adjusted wave-sealing feature parameters at each time point is equivalent to modulating the amplitude of the adjusted timbre audio signal through the amplitude included in the adjusted wave-sealing feature parameters. This ensures that the amplitude of the adjusted timbre audio signal is adjusted by the amplitude included in the adjusted wave-sealing feature parameters at each time point, ensuring that the generated target audio not only has the correct pitch and duration, but also has natural amplitude dynamic changes, making the generated target audio sound more layered and realistic.
[0097] The technical solution of this embodiment improves the quality of the generated target audio by multiplying the amplitude of the adjusted timbre audio signal with the amplitude included in the adjusted wave-sealing characteristic parameters at each time point. This allows the target timbre to be used to render target audio of different durations and pitches with a natural listening experience.
[0098] In another embodiment, such as Figure 5 As shown, an audio synthesis method is provided, which can be applied to... Figure 1 Taking terminal 102 as an example, the explanation includes the following steps:
[0099] S502 acquires the target audio signal of the target timbre and the melodic feature information of the target melody.
[0100] Melodic feature information includes melody duration and melody pitch.
[0101] In one embodiment, the original audio signal of the target timbre and multiple wave-sealing feature parameters corresponding to the original audio signal are obtained; the wave-sealing feature parameters of the original audio signal are used to reflect the change of the amplitude of the original audio signal during the duration of the wave-sealing feature parameters; based on the ratio between the amplitude change and the duration corresponding to each wave-sealing feature parameter, wave-sealing feature parameters with a ratio less than a preset ratio threshold are determined from the multiple wave-sealing feature parameters as target wave-sealing feature parameters; and the target audio signal of the target timbre is generated based on the audio signal segment corresponding to the target wave-sealing feature parameter in the original audio signal.
[0102] In one embodiment, a target audio signal with a target timbre is generated based on the audio signal segment corresponding to the target wave blocking feature parameter in the original audio signal, including: if the average amplitude of the audio signal segment is less than a preset amplitude threshold, the audio signal segment is amplified to obtain the target audio signal with the target timbre; if the amplitude fluctuation frequency of the audio signal segment is greater than a preset frequency threshold, the audio signal segment is compressed to obtain the target audio signal with the target timbre.
[0103] S504, based on the duration of the melody and the duration of the target audio signal, several target audio signals are spliced together end to end, with the two spliced target audio signals partially overlapping, to obtain a spliced audio signal with the same duration as the target melody.
[0104] S506, for the overlapping parts in the spliced audio signal, fade out the end part of the overlapping part that belongs to the target audio signal, and fade in the beginning part of the overlapping part that belongs to the target audio signal, to obtain the spliced audio signal.
[0105] S508 determines the resampling coefficients based on the pitch difference between the melody pitch and the pitch of the spliced timbre audio signal.
[0106] S510 resamples the spliced timbre audio signal according to the resampling coefficient to obtain the adjusted timbre audio signal, so that the pitch of the adjusted timbre audio signal is consistent with the melody pitch, and the playback rate of the adjusted timbre audio signal is consistent with that of the spliced timbre audio signal.
[0107] S512, determine the sum of durations of at least one wave-sealing characteristic parameter.
[0108] The wave sealing characteristic parameter of the target audio signal is used to reflect the change in the amplitude of the target audio signal during the duration of the wave sealing characteristic parameter.
[0109] S514, according to the duration ratio between the sum of the melody duration and the duration of the duration, the duration of at least one wave-sealing feature parameter is scaled so that the sum of the durations of the scaled wave-sealing feature parameter is equal to the melody duration, thus obtaining the adjusted wave-sealing feature parameter.
[0110] S516 multiplies the amplitude of the adjusted timbre audio signal with the amplitude included in the adjusted wave seal characteristic parameters at each time point to generate the target audio.
[0111] The target audio is a sound audio that has the target timbre and matches the target melody.
[0112] It should be noted that the specific limitations of the above steps can be found in the specific limitations of an audio synthesis method described above.
[0113] In practical applications, the embodiments of this application can be well applied to the construction of high-quality audio source files in scenarios where audio sources are scarce. This construction can generate high-quality audio synthesis data in batches, and can be used as a data augmentation scheme or the construction of the dataset itself to support tasks such as timbre transfer and audio domain generation. Moreover, it can realize the conversion from symbolic musical representation to audio content, and therefore can serve as a downstream generation interface for music-related tasks based on symbolic representation. It can also generate high-quality audio with low computational cost. At the same time, it can realize the rapid sampling of a sound, for example, for the preservation or restoration of some endangered ethnic musical instruments that need protection, or for some interesting applications (ambient sounds, birdsong, etc.).
[0114] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0115] Based on the same inventive concept, this application also provides an audio synthesis apparatus for implementing the audio synthesis method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more audio synthesis apparatus embodiments provided below can be found in the limitations of the audio synthesis method described above, and will not be repeated here.
[0116] In an exemplary embodiment, Figure 6 As shown, an audio synthesis apparatus is provided, comprising:
[0117] The acquisition module 610 is used to acquire the target audio signal of the target timbre and the melody feature information of the target melody, wherein the melody feature information includes the melody duration and the melody pitch.
[0118] The audio adjustment module 620 is used to splice several target audio signals to obtain a spliced timbre audio signal according to the duration of the melody and the duration of the target audio signal, and to perform pitch shifting processing on the spliced timbre audio signal according to the pitch of the melody to obtain an adjusted timbre audio signal.
[0119] The parameter adjustment module 630 is used to adjust the duration of at least one wave-sealing characteristic parameter of the target audio signal according to the duration of the melody to obtain the adjusted wave-sealing characteristic parameter; the wave-sealing characteristic parameter of the target audio signal is used to reflect the change of the amplitude of the target audio signal within the duration of the wave-sealing characteristic parameter.
[0120] The generation module 640 is used to generate target audio based on the adjusted timbre audio signal and the adjusted wave sealing characteristic parameters; the target audio is a sound audio that has the target timbre and conforms to the target melody.
[0121] In one embodiment, the audio adjustment module 620 is specifically used to splice several target audio signals end to end, with the spliced two target audio signals partially overlapping, to obtain a spliced audio signal with the same melody duration as the target melody; for the overlapping part in the spliced audio signal, the end part of the overlapping part belonging to the target audio signal is faded out, and the beginning part of the overlapping part belonging to the target audio signal is faded in, to obtain the spliced timbre audio signal.
[0122] In one embodiment, the audio adjustment module 620 is specifically used to determine a resampling coefficient based on the pitch difference between the melody pitch and the pitch of the spliced timbre audio signal; and to resample the spliced timbre audio signal according to the resampling coefficient to obtain the adjusted timbre audio signal, so that the pitch of the adjusted timbre audio signal is consistent with the melody pitch, and the playback rate of the adjusted timbre audio signal is consistent with that of the spliced timbre audio signal.
[0123] In one embodiment, the acquisition module 610 is specifically used to acquire the original audio signal of the target timbre and a plurality of wave-sealing feature parameters corresponding to the original audio signal; the wave-sealing feature parameters of the original audio signal are used to reflect the change of the amplitude of the original audio signal within the duration of the wave-sealing feature parameters; based on the ratio between the amplitude change and the duration corresponding to each wave-sealing feature parameter, wave-sealing feature parameters whose ratio is less than a preset ratio threshold are determined from the plurality of wave-sealing feature parameters as target wave-sealing feature parameters; and the target audio signal of the target timbre is generated based on the audio signal segment corresponding to the target wave-sealing feature parameter in the original audio signal.
[0124] In one embodiment, the acquisition module 610 is specifically used to generate a target audio signal of the target timbre based on the audio signal segment corresponding to the target wave blocking feature parameter in the original audio signal, including: if the average amplitude of the audio signal segment is less than a preset amplitude threshold, amplifying the audio signal segment to obtain the target audio signal of the target timbre; if the amplitude fluctuation frequency of the audio signal segment is greater than a preset frequency threshold, compressing the audio signal segment to obtain the target audio signal of the target timbre.
[0125] In one embodiment, the parameter adjustment module 630 is specifically used to determine the sum of the durations of at least one wave-sealing feature parameter; and to scale the duration of at least one wave-sealing feature parameter according to the duration ratio between the melody duration and the sum of the durations, so that the sum of the durations of at least one wave-sealing feature parameter after scaling is equal to the melody duration, thereby obtaining the adjusted wave-sealing feature parameter.
[0126] In one embodiment, the generation module 640 is specifically used to multiply the amplitude of the adjusted timbre audio signal with the amplitude included in the adjusted wave seal feature parameters at each time point to generate the target audio.
[0127] Each module in the aforementioned audio synthesis device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0128] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 7 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements an audio synthesis method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0129] Those skilled in the art will understand that Figure 7The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0130] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0131] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0132] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0133] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0134] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0135] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0136] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. An audio synthesis method, characterized in that, The method includes: Acquire the target audio signal of the target timbre, and acquire the melody feature information of the target melody, wherein the melody feature information includes melody duration and melody pitch; Based on the duration of the melody and the duration of the target audio signal, several target audio signals are spliced together to obtain a spliced timbre audio signal, and the spliced timbre audio signal is then subjected to pitch shifting processing based on the melody pitch to obtain an adjusted timbre audio signal; Based on the duration of the melody, the duration of at least one wave-sealing characteristic parameter of the target audio signal is adjusted to obtain the adjusted wave-sealing characteristic parameter; the wave-sealing characteristic parameter of the target audio signal is used to reflect the change of the amplitude of the target audio signal within the duration of the wave-sealing characteristic parameter; Based on the adjusted timbre audio signal and the adjusted wave sealing characteristic parameters, a target audio is generated; the target audio is a sound audio that possesses the target timbre and conforms to the target melody.
2. The method according to claim 1, characterized in that, The process of splicing together several target audio signals to obtain a spliced audio signal includes: Several target audio signals are spliced together end to end, with the two target audio signals overlapping to obtain a spliced audio signal with the same melody duration as the target melody. For the overlapping portion in the spliced audio signal, the end portion of the overlapping portion belonging to the target audio signal is faded out, and the beginning portion of the overlapping portion belonging to the target audio signal is faded in, to obtain the spliced audio signal.
3. The method according to claim 1, characterized in that, The step of performing pitch shifting processing on the spliced timbre audio signal according to the melody pitch to obtain the adjusted timbre audio signal includes: The resampling coefficients are determined based on the pitch difference between the melody pitch and the pitch of the spliced timbre audio signal; The spliced timbre audio signal is resampled according to the resampling coefficient to obtain the adjusted timbre audio signal, so that the pitch of the adjusted timbre audio signal is consistent with the pitch of the melody, and the playback rate of the adjusted timbre audio signal is consistent with that of the spliced timbre audio signal.
4. The method according to claim 1, characterized in that, The target audio signal for acquiring the target timbre includes: The original audio signal of the target timbre and multiple wave-sealing feature parameters corresponding to the original audio signal are obtained; the wave-sealing feature parameters of the original audio signal are used to reflect the change of the amplitude of the original audio signal during the duration of the wave-sealing feature parameters; Based on the ratio between the amplitude change and the duration corresponding to each wave sealing feature parameter, wave sealing feature parameters whose ratio is less than a preset ratio threshold are determined from the plurality of wave sealing feature parameters as target wave sealing feature parameters; The target audio signal with the target timbre is generated based on the audio signal segment corresponding to the target wave sealing feature parameter in the original audio signal.
5. The method according to claim 4, characterized in that, The step of generating the target audio signal with the target timbre based on the audio signal segment corresponding to the target wave blocking feature parameter in the original audio signal includes: If the average amplitude of the audio signal segment is less than a preset amplitude threshold, the audio signal segment is amplified to obtain the target audio signal with the target timbre. If the amplitude fluctuation frequency of the audio signal segment is greater than a preset frequency threshold, the audio signal segment is compressed to obtain the target audio signal with the target timbre.
6. The method according to claim 1, characterized in that, The step of adjusting the duration of at least one wave-sealing feature parameter of the target audio signal according to the duration of the melody to obtain the adjusted wave-sealing feature parameter includes: Determine the sum of the durations of the at least one wave-blocking characteristic parameter; The duration of the at least one wave-sealing feature parameter is scaled according to the duration ratio between the sum of the melody duration and the duration, so that the sum of the durations of the at least one wave-sealing feature parameter after scaling is equal to the melody duration, thus obtaining the adjusted wave-sealing feature parameter.
7. The method according to claim 1, characterized in that, The step of generating the target audio based on the adjusted timbre audio signal and the adjusted wave-blocking characteristic parameters includes: The amplitude of the adjusted timbre audio signal is multiplied point-by-point at each time point with the amplitude included in the adjusted wave sealing characteristic parameters to generate the target audio.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for realizing audio pitch shifting
CN101847404A
Device and method for manipulating an audio signal
CN102365681A