Tone synthesis method, device, equipment and medium
By adaptively adjusting pitch offset, distortion intensity, and formant, and combining low-frequency oscillation signals, the problem of unnatural timbre in existing acoustic effects algorithms has been solved, achieving personalized and natural timbre synthesis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-03-10
AI Technical Summary
Existing acoustic effects algorithms rely on fixed parameters, resulting in unnatural timbre synthesis effects. They cannot adapt to sound materials with different characteristics, leading to timbre imbalance and loss of detail. In particular, the timbre appears stiff or lacks personality in scenarios such as speech synthesis and virtual singers.
By extracting key acoustic parameters from the audio to be processed, the pitch shift, distortion intensity, and bandpass formant are adaptively adjusted, and combined with low-frequency oscillation signals and distortion shaping modules, specific timbre effects are dynamically generated.
It achieves personalized and natural timbre synthesis, adapts to different sound source characteristics, generates diverse timbre effects, and maintains consistency and balance in sound quality.
Smart Images

Figure CN121640983A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing, and in particular to a method, apparatus, device and medium for timbre synthesis. Background Technology
[0002] Most current mainstream acoustic effects algorithms, such as equalization, compression, reverberation, and various vocoder modulation techniques, still rely on pre-set fixed parameters for audio processing. While these methods are simple to implement, computationally efficient, and easy to deploy in production lines, their inherent static nature also brings significant limitations.
[0003] Applying the same set of processing parameters to sound materials with different characteristics often produces significantly different auditory effects. This one-size-fits-all approach not only leads to an imbalance in timbre in the final output but also easily causes loss of detail in key frequency bands. Especially in scenarios such as speech synthesis, virtual singers, and real-time voice changing, where the naturalness and expressiveness of timbre are highly demanding, fixed-parameter algorithms result in synthesized timbres that sound stiff, flat, or lack personality, leading to poor overall sound quality and auditory realism.
[0004] Therefore, how to improve the timbre synthesis effect is a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0005] The purpose of this application is to provide a method, apparatus, device, and medium for timbre synthesis that can improve the timbre synthesis effect.
[0006] Firstly, a method for synthesizing timbre is provided, including:
[0007] Get the audio to be processed and the target timbre type;
[0008] Extract the key acoustic parameters of the audio to be processed;
[0009] Based on the key acoustic parameters and the target timbre type, determine the pitch shift, distortion intensity, and bandpass formant.
[0010] The pitch of the audio to be processed is adjusted according to the pitch offset to obtain a first signal; distortion shaping is performed based on the first signal and the distortion intensity to obtain a second signal; resonance enhancement is performed based on the second signal and the bandpass formant to obtain a synthesized audio.
[0011] In a preferred embodiment, this application may be further configured such that the key acoustic parameters include: average loudness, spectral centroid, and average pitch;
[0012] Based on the key acoustic parameters and the target timbre type, determine the pitch shift, distortion intensity, and bandpass formant, including:
[0013] Based on the average pitch, determine the pitch offset corresponding to the target timbre type;
[0014] Based on the key acoustic parameters, determine the acoustic performance score; determine the reference frequency and preset frequency corresponding to the target timbre type; determine the mechanical intensity based on the spectral centroid, reference frequency, and average loudness; determine the distortion intensity based on the mechanical intensity and key acoustic parameters.
[0015] The center frequency is determined based on the preset frequency and the spectral centroid; the bandpass resonance peak is determined based on the center frequency and the quality factor.
[0016] In a preferred embodiment, this application may be further configured to: after adjusting the pitch of the audio to be processed according to the pitch offset to obtain the first signal, further comprising:
[0017] Determine whether the target timbre type is a mechanical timbre;
[0018] If so, the first signal is modulated by a low-frequency oscillation signal to obtain the modulated first signal;
[0019] Accordingly, distortion shaping is performed based on the first signal and the distortion intensity to obtain the second signal, including:
[0020] The second signal is obtained by distortion shaping based on the modulated first signal and the distortion intensity.
[0021] In a preferred embodiment, this application may be further configured such that: when the first signal is modulated by a low-frequency oscillation signal, at least amplitude modulation is performed; and when the input audio is singing audio, frequency modulation is also performed.
[0022] In a preferred embodiment, this application may be further configured such that the modulation parameters of the low-frequency oscillation signal include modulation frequency and modulation depth, wherein the modulation frequency is determined based on the spectral centroid and the acoustic performance score;
[0023] The amplitude modulation is based on mechanical intensity modulation.
[0024] In a preferred embodiment, this application can be further configured to: determine an acoustic performance score based on the key acoustic parameters, including:
[0025] Based on the key acoustic parameters, an acoustic performance score is determined using a fully connected network.
[0026] In a preferred embodiment, this application may be further configured to include:
[0027] Determine whether the acoustic performance score is the median value;
[0028] If so, then timbre synthesis will not be performed;
[0029] If not, then the pitch of the audio to be processed is adjusted according to the pitch offset to obtain a first signal; distortion shaping is performed based on the first signal and the distortion intensity to obtain a second signal; and resonance enhancement is performed based on the second signal and the bandpass formant to obtain a synthesized audio.
[0030] Secondly, a timbre synthesis device is provided, comprising:
[0031] The acquisition module is used to acquire the audio to be processed and the target timbre type;
[0032] The extraction module is used to extract key acoustic parameters of the audio to be processed;
[0033] The determination module is used to determine the pitch shift, distortion intensity, and bandpass formant based on the key acoustic parameters and the target timbre type.
[0034] The synthesis module is used to adjust the pitch of the audio to be processed according to the pitch offset to obtain a first signal; perform distortion shaping based on the first signal and distortion intensity to obtain a second signal; and perform resonance enhancement based on the second signal and bandpass formant to obtain synthesized audio.
[0035] Thirdly, an electronic device is provided, the electronic device including a memory and a processor, the memory storing a computer program, the processor executing the timbre synthesis method according to any one of the first aspects when running the computer program.
[0036] Fourthly, a computer-readable storage medium is provided, wherein at least one piece of program code is stored therein, the program code being loaded and executed by a processor to implement the timbre synthesis method as described in any of the first aspects.
[0037] Fifthly, a computer program product is provided, including a computer program or instructions that, when executed by a processor, implement the timbre synthesis method as described in any of the first aspects.
[0038] In summary, the timbre synthesis method provided in this application has the following beneficial technical effects:
[0039] The process involves acquiring the audio to be processed and the target timbre type; extracting key acoustic parameters from the audio to be processed, which accurately reflect the original characteristics of the audio; then, based on the key acoustic parameters of the audio to be processed and the target timbre type, adaptively determining three key elements: pitch shift, distortion intensity, and bandpass formant; and finally, adjusting the audio to be processed using pitch shift, distortion intensity, and bandpass formant to obtain synthesized audio. This process can adaptively synthesize audio with specific timbre characteristics based on the characteristics of the audio to be processed, meeting diverse timbre requirements and improving the accuracy of timbre synthesis.
[0040] In addition, this application also provides a timbre synthesis device, equipment, and medium, all of which have the aforementioned beneficial technical effects. Attached Figure Description
[0041] To more clearly illustrate the technical solutions of the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 This is a schematic diagram of a timbre synthesis method provided in an embodiment of this application;
[0043] Figure 2 This is a schematic diagram of the overall structure of a timbre synthesis provided in an embodiment of this application;
[0044] Figure 3 This is a schematic diagram of the structure of a tone synthesis device provided in an embodiment of this application;
[0045] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0046] This specific embodiment is merely an explanation of this application and is not intended to limit it. After reading this specification, those skilled in the art can make modifications to this embodiment without contributing any inventive step, but such modifications are protected by patent law as long as they are within the scope of this application.
[0047] It should be noted that, in the optional embodiments of this application, the data related to object information, when applied to specific products or technologies, requires the permission or consent of the object. Furthermore, the collection, use, and processing of this data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. In other words, if the embodiments of this application involve data related to an object, it must be obtained with the permission and consent of the object, the permission and consent of relevant departments, and in accordance with the relevant laws, regulations, and standards of the country and region. If the embodiments involve personal information, the acquisition of all personal information requires the consent of the individual. If sensitive information is involved, the separate consent of the information subject is required. The embodiments also need to be implemented with the permission and consent of the object.
[0048] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0049] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.
[0050] Most current mainstream acoustic effects algorithms are based on fixed parameters, including traditional methods such as linear filtering, distortion functions, oscillation modulation, and resonance enhancement. While these algorithms are simple to implement, they have significant limitations. First, they lack adaptability to differences in timbre; the same processing parameters will produce completely different effects on male, female, or children's voices. For example, when the parameters are too high, high-pitched female voices are prone to harsh popping sounds, while low-pitched male voices may sound muffled and lack brightness.
[0051] Secondly, the inconsistency between sound brightness and dynamic response is another key issue. Fixed filtering and equalization strategies cannot automatically adjust according to the energy distribution and spectral structure of the sound, making some timbres sound thin in the high-frequency range and overly thick in the low-frequency range. Especially in complex multi-segment audio or mixed vocals, static parameters struggle to maintain consistency and naturalness in the output.
[0052] Furthermore, traditional distortion and modulation algorithms often exhibit clipping or popping issues at high energy inputs, while displaying indistinct characteristics and a lack of mechanical feel at low energy inputs. The result is an unbalanced tone, severe loss of detail, and a lack of intelligent adaptation.
[0053] Traditional special effects vocal processing often relies on fixed filtering, equalization, and distortion parameters, lacking the ability to adapt to differences in the input sound source. Fixed parameters often lead to excessive distortion, insufficient brightness, or loss of dynamics under different genders, languages, and vocal tones. This invention, however, performs real-time feature analysis on the vocal signal and dynamically adjusts the timbre parameters based on the analysis results, achieving personalized, controllable, and highly natural "special timbre" generation. It is particularly suitable for multimedia scenarios such as virtual singers, digital human singing, film and television dubbing, and game character sound effects.
[0054] This application extracts acoustic features from audio through a feature analysis module, generates special timbre intensity parameters, and adaptively controls modules such as pitch shift, modulation mixing, distortion shaping, and resonance enhancement based on these parameters. The algorithm dynamically superimposes low-frequency oscillations and nonlinear distortion into the audio to achieve a balance between brightness and saturation, ultimately outputting a human voice with a special texture and timbre. This application introduces a real-time acoustic feature detection and parameter adaptive adjustment mechanism, enabling the system to automatically determine pitch, loudness, and spectral structure based on the characteristics of the input sound source during audio generation or processing, thereby dynamically adjusting the mechanical feel, metallic texture, and brightness levels. In this way, the system can not only adapt to the timbre of different sound sources but also achieve intelligent generation and control of special timbres.
[0055] Specifically, embodiments of this application provide a timbre synthesis method, such as... Figure 1 As shown, the method provided in this application embodiment can be executed by an electronic device, which is a server. This server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal device can be a smartphone, tablet, laptop, desktop computer, etc., but is not limited to these. The terminal device and electronic device can be directly or indirectly connected via wired or wireless communication. This application embodiment does not impose any limitations on this connection. The method includes:
[0056] S101. Obtain the audio to be processed and the target timbre type;
[0057] The audio to be processed can be singing, speech, or other types of audio. The target timbre type indicates the user's expectation that the audio to be processed will be converted into a specific timbre category, such as mechanical timbre, children's voice, or broadcast voice.
[0058] S102. Extract the key acoustic parameters of the audio to be processed;
[0059] When an electronic device receives an input audio segment (the audio to be processed), the feature analysis module first performs basic acoustic analysis, including frame segmentation, feature extraction, and energy statistics. It automatically calculates key acoustic parameters such as average loudness, spectral centroid, and average pitch to reflect the sound's energy distribution, brightness characteristics, and acoustic structure. These features collectively constitute a multi-dimensional characterization of the overall sound performance. Average loudness refers to the average measure of sound intensity in the audio to be processed; spectral centroid refers to the center of gravity of the audio signal's spectral energy distribution, indicating the degree of concentration of energy in the audio spectrum; and average pitch refers to the average level of sound pitch throughout the entire audio to be processed, calculated using the `analyze_audio` function.
[0060] S103. Determine the pitch shift, distortion intensity, and bandpass formant based on key acoustic parameters and target timbre type.
[0061] In this embodiment, the pitch, distortion, and frequency band characteristics of the audio are precisely adjusted based on key acoustic parameters and the target timbre type to achieve the desired timbre effect. By inputting the characteristics of the audio itself and determining a suitable pitch offset, the overall pitch of the audio can be changed to meet the requirements of the target timbre type; determining a reasonable distortion intensity can add specific distortion effects to the audio; and determining accurate bandpass formants can highlight the characteristics of the audio in a specific frequency band.
[0062] One possible implementation of this application embodiment includes key acoustic parameters such as: average loudness, spectral centroid, and average pitch.
[0063] S103 determines the pitch shift, distortion intensity, and bandpass formant based on key acoustic parameters and the target timbre type, including:
[0064] S1031. Determine the pitch offset corresponding to the target timbre type based on the average pitch.
[0065] The average pitch of the input audio is calculated using the analyze_audio function.
[0066] After determining the target timbre type, the corresponding pitch offset is determined by using the offsets of multiple timbre types and multiple pitch categories pre-stored in the electronic device.
[0067] Specifically, taking the mechanical sound category as an example, the preset correspondence is shown in Table 1:
[0068] Table 1. Correspondence between categories of mechanical sounds
[0069]
[0070] Based on the average pitch, the vocal range is roughly divided into two pitch categories: "high-mid" and "low-mid". Different semitone offsets are assigned to each category. After determining the average pitch, the pitch category is determined, and then the pitch offset is determined.
[0071] S1032. Determine the acoustic performance score based on key acoustic parameters;
[0072] To achieve a unified description at the algorithm level, the system integrates these features into a new evaluation metric, the Acoustic Expression Index (AEI).
[0073] AEI (Adaptive Energy Index) is a comprehensive quantitative indicator of sound characteristics, used to describe the overall performance of audio in terms of energy, brightness, and dynamic range. It reflects the "expressiveness" or "characteristic intensity" of the sound. A high AEI usually indicates concentrated audio energy, strong brightness, and active dynamics; while a low AEI indicates a relatively soft, dark sound or a smaller dynamic range. The introduction of AEI allows algorithms to adaptively adjust based on the inherent characteristics of the audio without human intervention. After obtaining the AEI, the system automatically generates key adaptive control parameters: modulation frequency and distortion intensity.
[0074] The corresponding pitch offset, modulation frequency, distortion intensity, and resonance ratio are determined. This process is the core of the algorithm's self-learning, enabling it to automatically generate special effects configurations based on the characteristics of the sound source.
[0075] These four parameters collectively determine the spatiality, saturation, brightness, and metallicity of the timbre. They are not independent but rather form a dynamic balance.
[0076] (1) When the AEI is high, the system will automatically reduce the distortion intensity and increase the modulation frequency to generate a smoother, more transparent, and electronically textured tone.
[0077] (2) When the AEI is low, the system enhances the distortion and resonance components, making the output sound thicker and more penetrating, presenting a mechanical or metallic effect;
[0078] (3) When AEI is at the middle value of 0.5, all parameters remain balanced, thus generating a natural, delicate and spatially layered timbre.
[0079] For example, taking mechanical timbre as an example, a high AEI indicates high brightness (large Centroid, e.g., >3500); low loudness (small RMS, e.g., <0.1); and high treble (large PitchMean, e.g., >250Hz). A low AEI indicates low brightness (small Centroid, e.g., <2000); high loudness (large RMS, e.g., >0.4); and high bass (small PitchMean, e.g., <150Hz). A higher AEI corresponds to a sharp, weak, or dynamically active sound, requiring more refined electronic modulation and less distortion. A lower AEI corresponds to a full, loud, or dynamically stable sound, requiring stronger distortion and resonance to enhance the mechanical feel.
[0080] In one feasible approach, after determining the key acoustic parameters, they are sent to the user client; then, in response to the user's input rating on the user client, an acoustic performance score is obtained.
[0081] In another feasible approach, an acoustic performance score is determined using a fully connected network based on key acoustic parameters. This fully connected network is an existing neural network structure used to determine the acoustic performance score. In this embodiment, a large amount of training data is acquired, including training key acoustic parameters and corresponding labels, where the labels represent the acoustic performance score. This large amount of training data is then used to train the fully connected network, enabling it to learn the relationship between the key acoustic parameters and the acoustic performance score. This results in a trained fully connected network used to determine the acoustic performance score corresponding to the actual key acoustic parameters.
[0082] Specifically, during training, key acoustic parameters are input into the fully connected network. The fully connected network calculates the predicted acoustic performance score based on the current weights and biases, then compares it with the true label, calculates the loss function (such as the mean squared error loss function), and adjusts the weights and biases of the fully connected network through backpropagation to make the predicted value gradually approach the true value. After multiple iterations of training, the fully connected network learns the relationship between the key acoustic parameters and the acoustic performance score, resulting in a well-trained fully connected network. It is understood that in this embodiment, the acoustic performance score ranges from 0 to 1.
[0083] S1033. Determine the reference frequency and preset frequency corresponding to the target timbre type;
[0084] Different timbre types correspond to different hyperparameters. In this embodiment, after determining the target timbre type, a reference frequency and a preset frequency can be determined. The reference frequency is used to determine the mechanical intensity, and the preset frequency is used to determine the center frequency.
[0085] S1034. Determine the mechanical intensity based on the spectral centroid, reference frequency, and average loudness;
[0086] The mechanical intensity is determined by using mech_intensity=np.clip((centroid / 3000)+(0.5-rms),0.3,1.2), where mech_intensity represents the mechanical intensity, centroid represents the spectral centroid, 3000 Hz is the reference frequency, and rms is the average loudness.
[0087] S1035. Determine the distortion intensity based on mechanical intensity and key acoustic parameters;
[0088] The distortion_strength is calculated as 1.2 + (mech_intensity * AEI * 0.8), where distortion_strength is the distortion intensity.
[0089] In this embodiment, the distortion intensity calculation method employs the following approach: The required degree of mechanization is calculated based on the luminance and loudness of the sound. Luminance contribution: The system first calculates the ratio of the spectral centroid to a reference frequency (e.g., 3000 Hz for mechanical timbre). Loudness contribution: The system subtracts the average loudness (rms) from a constant (0.5). The reference intensity and adaptive gain are then added together to obtain the final distortion intensity.
[0090] Understandably, during distortion processing, the distortion intensity can also be combined with the resonance component to form the input parameters of the distortion shaping module for processing. The resonance component is a weighted average of 0.3 * res, where res is a component in AEI.
[0091] The Acoustic Performance Index (AEI) is output by a fully connected neural network. This fully connected neural network takes average loudness, spectral centroid, and average pitch as inputs, and after passing through multiple layers of nonlinear mapping, outputs a continuous value between 0 and 1 through a sigmoid activation function to characterize the overall performance of the audio.
[0092] res is the luminance component generated by the luminance-specific branch in the FCN network. This component reflects the strength of the high-frequency energy of the sound and participates in the final calculation of AEI.
[0093] Input the average loudness, spectral centroid, and average pitch. Obtain the intermediate feature vector F (32-dimensional) through Dense(64, ReLU) and Dense(32, ReLU). Then obtain res through Dense(8, ReLU) and Dense(1, Sigmoid). Then, AEI_input=concat([F,res]), 33-dimensional, and output AEI through Dense(16, ReLU) and Dense(1, Sigmoid).
[0094] S1036. Determine the center frequency based on the preset frequency and spectral centroid; determine the bandpass resonance peak based on the center frequency and quality factor.
[0095] In this embodiment, the resonance region (i.e., the center frequency freq) is adaptively determined based on the spectral centroid. The higher the spectral centroid (the brighter the sound), the formant peak should shift to a higher frequency to match and enhance the high-frequency part of the sound, maintaining the consistency of the target timbre type.
[0096] For example, taking a mechanical tone as an example, the center frequency freq is set to 1500 + centroid / 8. Here, 1500 is the preset frequency. For example, when the centroid is 3000Hz, the freq is 1875Hz. In order to ensure that the resonance area adapts to the "brightness" of the input audio, the mid-high frequency range (above 1500Hz) is the key area for expressing the metallic texture and spatial luster.
[0097] The bandwidth / quality factor (q) is fixed at 3.0 in the code, where the q value determines the sharpness of the formants. q=3.0 is a value of medium sharpness, providing a noticeable formant without being overly sharp and causing harshness or feedback, offering a stable mechanical texture. This quality factor also corresponds to the target timbre type.
[0098] S104. Adjust the pitch of the audio to be processed according to the pitch offset to obtain a first signal; perform distortion shaping based on the first signal and distortion intensity to obtain a second signal; perform resonance enhancement based on the second signal and bandpass formant to obtain a synthesized audio.
[0099] During the execution of the algorithm, the pitch shift module adjusts the frequency structure of the input signal to make the timbre more stable and layered overall.
[0100] The determined semitone offset (pitch offset) is applied to the `librosa.effects.pitch_shift` function, using techniques such as phase vocoders to achieve high-quality pitch changes while preserving the temporal characteristics of the speech to the greatest extent possible, thus adjusting the pitch of the entire audio signal. The code includes: `y_pitch = librosa.effects.pitch_shift(y=y, sr=sr, n_steps=pitch_shift_steps)`. If the target timbre is a mechanical timbre, this algorithm achieves personalized processing of different vocal input signals: slightly lowering high frequencies and significantly lowering low frequencies, ensuring that the final mechanical timbre maintains a consistent and stable thickness.
[0101] Furthermore, for mechanical timbres, in addition to pitch adjustment, low-frequency modulation processing can also be performed. Pitch adjustment is used to balance the thickness of the vocal timbre and the center of gravity of the frequency band; low-frequency modulation superimposes periodic fluctuations on the original signal, which can generate various styles such as mechanical, science fiction, fantasy, and spatial, depending on the application scenario.
[0102] Based on the distortion intensity obtained from AEI, the nonlinear intensity of the distortion function is controlled to increase saturation and brightness in low-energy audio, while moderately suppressing gain in high-energy audio to prevent popping or clipping. Through this adaptive mechanism, the algorithm can maintain the balance and stability of the output sound under various input conditions, achieving continuous adjustment from slight overload to deep saturation. Subsequently, a controllable bandpass formant is added in the mid-high frequency range (determined according to the target timbre type; for mechanical timbre, it corresponds to 1500). For mechanical timbre, this can make the sound brighter, more metallic, or have a spatial sheen, allowing the output timbre to achieve a more distinctive and personalized expression while maintaining the characteristics of the original sound source. Appropriate filtering can also be applied to achieve high-fidelity, low-noise special effects output.
[0103] Furthermore, at the end of the entire processing flow, standardization and dynamic range adjustment are performed to ensure that the output audio achieves the optimal balance between loudness and clarity. At this point, the adaptively adjusted sound possesses highly customized special effects attributes. Through different combinations of parameter configurations, this algorithm can not only output a mechanical, metallic "mecha sound" style, but also be extended to generate various timbre types such as electronic pop, retro synthesis, vocal enhancement, or ambient sound design.
[0104] It is understood that the proposed solution is not limited to the generation of "mechanical sounds," but rather lies in its general acoustic expression model. This model achieves a complete closed-loop mechanism of "feature-driven—parameter generation—timbre synthesis" by analyzing sound source features, quantifying expressiveness, and adaptively mapping them to multiple sound effect control parameters.
[0105] Furthermore, after all processing is complete, amplitude normalization and dynamic range standardization are performed to ensure consistent loudness and balanced sound quality of the output signal on the playback device.
[0106] As can be seen, in this embodiment, the audio to be processed and the target timbre type are obtained; key acoustic parameters of the audio to be processed are extracted, which can accurately reflect the original characteristics of the audio. Then, based on the key acoustic parameters of the audio to be processed and the target timbre type, the three key elements of pitch shift, distortion intensity, and bandpass formant are adaptively determined; and then, the audio to be processed is adjusted by adjusting the pitch shift, distortion intensity, and bandpass formant to obtain synthesized audio. Based on the characteristics of the audio to be processed itself, audio with specific timbre characteristics can be adaptively synthesized to meet diverse timbre requirements and improve the accuracy of timbre synthesis.
[0107] One possible implementation of this application embodiment involves adjusting the pitch of the audio to be processed based on the pitch offset to obtain a first signal, and then further includes: determining whether the target timbre type is a mechanical timbre; if so, modulating the first signal with a low-frequency oscillation signal to obtain a modulated first signal.
[0108] In this embodiment, if the sound is mechanical, a low-frequency oscillation signal needs to be introduced. The intensity and frequency of this low-frequency oscillation signal can be adaptively adjusted according to the characteristics of the input audio. Introducing a low-frequency oscillation signal can superimpose periodic disturbances into the original sound, simulating a sense of motion or spatial vibration, such as mechanical resonance or electronic humming.
[0109] In one possible implementation of this application, when the first signal is modulated by a low-frequency oscillation signal, at least amplitude modulation is performed; when the input audio is singing audio, frequency modulation is also performed.
[0110] In this embodiment, after determining the low-frequency oscillation signal, amplitude modulation is performed by default; if the input audio is vocal audio, frequency modulation is also performed. During amplitude modulation, the low-frequency oscillation signal is multiplied by a first signal; the multiplication result and the first signal are then weighted and mixed, and amplitude modulation is performed to obtain the amplitude-adjusted signal. The low-frequency oscillation signal is multiplied by the first signal to obtain the frequency-modulated signal. The order of these two steps is not limited in this embodiment.
[0111] Specifically, the modulation parameters of low-frequency oscillation signals include modulation frequency and modulation depth. The modulation frequency is determined based on the spectral centroid and acoustic performance score; amplitude modulation is based on mechanical intensity.
[0112] In this embodiment, the modulation frequency mod_freq=20+(centroid*AEI / 1000). The higher the brightness (spectral centroid) of the sound, the higher the LFO frequency, in order to match the high-frequency electronic timbre.
[0113] The modulation depth / intensity is determined by the mechanical intensity (mech_intensity). The higher the intensity, the more pronounced the modulation effect. The mixing formula, 0.4 * mech_intensity * (y_pitch * mod_signal), determines the mixing weight of the LFO signal, where y_pitch is the signal to be modulated and mod_signal is the LFO signal.
[0114] This technique involves introducing low-frequency oscillation signals through a modulation mixing module to slightly and periodically perturb the amplitude or frequency of the audio, thereby creating a subtle sense of "motion" or "spatial vibration" on the original sound. This modulation can not only simulate the hum of mechanical resonance but also be used to shape the timbre of various styles, such as electronic, dreamy, and retro.
[0115] Accordingly, distortion shaping is performed based on the first signal and the distortion intensity to obtain the second signal, including: distortion shaping is performed based on the modulated first signal and the distortion intensity to obtain the second signal.
[0116] One possible implementation of this application further includes: determining whether the acoustic performance score is an intermediate value; if so, no timbre synthesis is performed; if not, pitch adjustment is performed on the audio to be processed based on the pitch offset to obtain a first signal; distortion shaping is performed based on the first signal and distortion intensity to obtain a second signal; resonance enhancement is performed based on the second signal and bandpass formant to obtain synthesized audio. The intermediate value can be 0.5. When the AEI is at the intermediate value of 0.5, all parameters remain balanced, thereby generating a natural, delicate, and spatially layered timbre performance without the need for adjustment.
[0117] Based on any of the above embodiments, the algorithm system proposed in this application is a general acoustic effects algorithm framework (such as...). Figure 2 This system (as shown) can adaptively generate various timbre effects based on the acoustic characteristics of the input audio. Its overall structure consists of five core modules: feature analysis module, pitch shift module, modulation mixing module, distortion shaping module, and resonance enhancement and filtering module. The following explanation uses a mechanical timbre as an example.
[0118] First, the feature analysis module is responsible for calculating the key acoustic parameters of the input audio, including average loudness, spectral centroid, and average pitch. These parameters comprehensively characterize the sound energy distribution, spectral center, and vocal features, providing fundamental data support for subsequent adaptive adjustments of special effects.
[0119] Secondly, the pitch offset module dynamically adjusts the pitch based on the analysis results to adapt to the characteristics of different types of sound sources. It can automatically generate a semitone offset based on the pitch mean to thicken the high-frequency timbre or enhance the low-frequency sound, thereby achieving ideal auditory layering and spatial depth in different scenarios.
[0120] Third, the modulation and mixing module modulates the pitch using a low-frequency oscillation signal (LFO). This module can flexibly simulate various periodic physical characteristics, such as mechanical resonance, electronic hum, water ripples, or electromagnetic pulse effects, and has a highly configurable style adaptability.
[0121] Fourth, the distortion shaping module dynamically compresses or expands signal energy through nonlinear transformation. It can be used to achieve the "overload" and "saturation" texture in electronic sound effects, and can also be used as a dynamic range controller to maintain a stable balance of sound under different input intensities.
[0122] Finally, the resonance enhancement and filtering module intelligently adds bandpass formants in the mid-to-high frequency region (above 1500), which can generate different timbre characteristics according to the target effect, such as metallic texture, electronic spatiality, or glass-like high-frequency brightness, enabling the algorithm to have multi-scenario application capabilities from natural sound to synthesized sound.
[0123] This invention relates to the fields of audio signal processing and artificial intelligence acoustic synthesis, and particularly to a special timbre singing synthesis algorithm based on adaptive feature adjustment. This algorithm can automatically adjust pitch shift, mechanical modulation frequency, distortion intensity, and resonance enhancement parameters according to the acoustic characteristics of the input human voice, such as pitch, loudness, and spectral centroid, thereby achieving adaptive synthesis of different timbres. This allows the output singing voice to achieve auditory effects with a "metallic," "mechanical," or "special vocal style" while maintaining the original timbre characteristics.
[0124] In this embodiment, an adaptive timbre control mechanism is first introduced. Unlike traditional fixed-parameter algorithms, this system can automatically adjust core parameters based on the spectrum, energy, and pitch characteristics of the input sound, achieving closed-loop control of "input perception—dynamic parameter adjustment—timbre output," thereby maintaining sound quality balance and consistency among different sound sources.
[0125] Secondly, this invention implements a multi-parameter collaborative modulation mechanism. Modules such as pitch, modulation, distortion, and resonance no longer operate independently, but are coupled together through a unified mechanical intensity parameter. The synergistic effect of multi-dimensional parameters makes the generated timbre more three-dimensional and spatially expansive, with a more natural mechanical texture, avoiding the harsh or excessive auditory impact of traditional special effects algorithms.
[0126] Finally, the algorithm achieves voice consistency protection while converting timbre. By preserving the fundamental frequency structure and superimposing mechanical features on top, the system effectively avoids semantic ambiguity and sibilance enhancement, ensuring that the output has both a distinctive timbre and clearly expresses the original speech content.
[0127] The following describes a timbre synthesis device provided in an embodiment of this application. The device described below can be referred to in correspondence with the method described above. The device in this embodiment is installed in an electronic device. Figure 3 , Figure 3 This is a structural block diagram of an apparatus according to one embodiment of this application, comprising:
[0128] The acquisition module 210 is used to acquire the audio to be processed and the target timbre type;
[0129] Extraction module 220 is used to extract key acoustic parameters of the audio to be processed;
[0130] The determination module 230 is used to determine the pitch shift, distortion intensity, and bandpass formant based on key acoustic parameters and target timbre type.
[0131] The synthesis module 240 is used to adjust the pitch of the audio to be processed according to the pitch offset to obtain a first signal; perform distortion shaping based on the first signal and distortion intensity to obtain a second signal; and perform resonance enhancement based on the second signal and bandpass formant to obtain the synthesized audio.
[0132] In one feasible approach, key acoustic parameters include: mean loudness, spectral centroid, and mean pitch;
[0133] Determine module 230, used for:
[0134] Determine the pitch offset corresponding to the target timbre type based on the average pitch;
[0135] Based on key acoustic parameters, determine the acoustic performance score; determine the reference frequency and preset frequency corresponding to the target timbre type; determine the mechanical intensity based on the spectral centroid, reference frequency, and average loudness; determine the distortion intensity based on the mechanical intensity and key acoustic parameters.
[0136] The center frequency is determined based on the preset frequency and the spectral centroid; the bandpass resonance peak is determined based on the center frequency and the quality factor.
[0137] One possible approach also includes:
[0138] The mechanical timbre determination module is used to determine whether the target timbre type is a mechanical timbre;
[0139] A modulation module is used to modulate the first signal with a low-frequency oscillation signal to obtain a modulated first signal if the condition is met.
[0140] Accordingly, the synthesis module 240 performs distortion shaping based on the first signal and the distortion intensity to obtain a second signal, which is used for:
[0141] The second signal is obtained by distortion shaping based on the modulated first signal and the distortion intensity.
[0142] In one feasible approach, when the first signal is modulated by a low-frequency oscillation signal, at least amplitude modulation is performed; when the input audio is singing audio, frequency modulation is also performed.
[0143] In one feasible approach, the modulation parameters of the low-frequency oscillation signal include modulation frequency and modulation depth, wherein the modulation frequency is determined based on the spectral centroid and acoustic performance score.
[0144] Amplitude modulation is based on mechanical intensity modulation.
[0145] In one feasible manner, the determining module 230 determines an acoustic performance score based on key acoustic parameters, for the purpose of:
[0146] Based on key acoustic parameters, an acoustic performance score is determined using a fully connected network.
[0147] One possible approach also includes:
[0148] The intermediate value determination module is used to determine whether the acoustic performance score is an intermediate value; if so, no timbre synthesis is performed; if not, the pitch of the audio to be processed is adjusted according to the pitch offset to obtain a first signal; distortion shaping is performed based on the first signal and the distortion intensity to obtain a second signal; resonance enhancement is performed based on the second signal and the bandpass formant to obtain the synthesized audio.
[0149] Figure 4 A structural diagram of an electronic device provided in an embodiment of the present invention, such as... Figure 4 As shown, the electronic device includes: a memory 60 for storing a computer program; and a processor 61 for executing the computer program to implement the steps of the method as described in the above embodiments.
[0150] The electronic devices provided in this embodiment may include, but are not limited to, smartphones, tablets, laptops, or desktop computers.
[0151] The processor 61 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 61 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The processor 61 may also include a main processor and a coprocessor. The main processor, also known as the Central Processing Unit (CPU), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 61 may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 61 may also include an Artificial Intelligence (AI) processor, which handles computational operations related to machine learning.
[0152] The memory 60 may include one or more computer-readable storage media, which may be non-transitory. The memory 60 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In this embodiment, the memory 60 is used to store at least the following computer program 601, which, after being loaded and executed by the processor 61, is capable of implementing the relevant steps of the timbre synthesis method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 60 may also include an operating system 602 and data 603, etc., and the storage method may be temporary storage or permanent storage. The operating system 602 may include Windows, Unix, Linux, etc.
[0153] In some embodiments, the electronic device may further include a display screen 62, an input / output interface 63, a communication interface 64, a power supply 65, and a communication bus 66.
[0154] Those skilled in the art will understand that Figure 4 The structures shown do not constitute a limitation on electronic devices and may include more or fewer components than those shown.
[0155] It is understood that if the methods in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the current technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods in the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, magnetic disks, or optical disks, and other media capable of storing program code.
[0156] Based on this, embodiments of the present invention also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the timbre synthesis method described above.
[0157] Based on this, embodiments of the present invention also provide a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the above-described method. It should be understood that although the steps in the flowcharts of the accompanying drawings are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.
[0158] The above are only some embodiments of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A timbre synthesis method characterized by, The method comprises the following steps: obtaining a to-be-processed audio and a target timbre type; extracting key acoustic parameters of the to-be-processed audio; determining a pitch offset, a distortion intensity and a band-pass formant according to the key acoustic parameters and the target timbre type; performing pitch adjustment on the to-be-processed audio according to the pitch offset to obtain a first signal; performing distortion shaping on the first signal and the distortion intensity to obtain a second signal; and performing resonance enhancement on the second signal and the band-pass formant to obtain a synthesized audio.
2. The timbre synthesis method according to claim 1, characterized in that, The key acoustic parameters comprise average loudness, spectral centroid and average pitch; determining the pitch offset, the distortion intensity and the band-pass formant according to the key acoustic parameters and the target timbre type comprises: determining the pitch offset corresponding to the target timbre type according to the average pitch; determining an acoustic performance degree score according to the key acoustic parameters; determining a reference frequency and a preset frequency corresponding to the target timbre type; determining mechanical intensity according to the spectral centroid, the reference frequency and the average loudness; and determining the distortion intensity according to the mechanical intensity and the key acoustic parameters; determining a center frequency according to the preset frequency and the spectral centroid; and determining the band-pass formant according to the center frequency and a quality factor.
3. The timbre synthesis method according to claim 2, characterized in that, After performing pitch adjustment on the to-be-processed audio according to the pitch offset to obtain the first signal, the method further comprises the following steps: determining whether the target timbre type is mechanical timbre; if yes, modulating the first signal by a low-frequency oscillation signal to obtain a modulated first signal; Correspondingly, the method of performing distortion shaping on the first signal and the distortion intensity to obtain the second signal comprises: performing distortion shaping on the modulated first signal and the distortion intensity to obtain the second signal.
4. The timbre synthesis method according to claim 3, characterized in that, When modulating the first signal by the low-frequency oscillation signal, at least amplitude modulation is performed; when the input audio is a song audio, frequency modulation is further performed.
5. The timbre synthesis method of claim 4, wherein, The modulation parameters of the low-frequency oscillation signal comprise a modulation frequency and a modulation depth, wherein the modulation frequency is determined according to the spectral centroid and the acoustic performance degree score; The amplitude modulation is performed based on the mechanical intensity.
6. The timbre synthesis method of claim 2, wherein, The method of determining the acoustic performance degree score according to the key acoustic parameters comprises: determining the acoustic performance degree score by using a full-connection network according to the key acoustic parameters.
7. The timbre synthesis method of claim 2, wherein, The method further comprises the following steps: determining whether the acoustic performance degree score is an intermediate value; if yes, not performing timbre synthesis; if not, performing the following steps: performing pitch adjustment on the to-be-processed audio according to the pitch offset to obtain the first signal; performing distortion shaping on the first signal and the distortion intensity to obtain the second signal; and performing resonance enhancement on the second signal and the band-pass formant to obtain the synthesized audio.
8. A timbre synthesizing apparatus characterized by comprising: The method comprises the following steps: an obtaining module, configured to obtain a to-be-processed audio and a target timbre type; an extracting module, configured to extract key acoustic parameters of the to-be-processed audio; a determining module, configured to determine a pitch offset, a distortion intensity and a band-pass formant according to the key acoustic parameters and the target timbre type; a synthesizing module, configured to perform pitch adjustment on the to-be-processed audio according to the pitch offset to obtain a first signal; perform distortion shaping on the first signal and the distortion intensity to obtain a second signal; and perform resonance enhancement on the second signal and the band-pass formant to obtain a synthesized audio.
9. An electronic device, comprising: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor executes the method according to any one of claims 1 to 7 when running the computer program.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores at least one program code, the program code is loaded and executed by the processor to realize the method according to any one of claims 1 to 7.