Audio processing method and device, equipment, storage medium and program product

By configuring carrier and modulation channels and using a vocoder plugin to synthesize audio of target timbre and sound effects, the problem of low efficiency and monotonous effects in generating bio-textured trembling sounds in existing technologies is solved, achieving efficient and low-cost audio generation and enhancing the immersive experience of game audio.

CN121815153APending Publication Date: 2026-04-07NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently generate vibrating sounds with a biological feel. On-site recording is inefficient, commercial sound effects libraries offer limited creative freedom, and sound distortion plugins are cumbersome to operate and produce monotonous results, leading to high game development costs and reduced immersion.

Method used

By configuring the carrier channel and modulation channel, the audio of the target timbre and target sound effect is aligned and synthesized using the vocoder plugin to generate a target audio that blends the target timbre and target sound effect.

Benefits of technology

It can quickly generate audio with target sound effects and timbres, reducing the time and cost of traditional recording, overcoming the problem of existing plugins generating monotonous or mechanical effects, and enhancing the lifelike texture and immersiveness of audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121815153A_ABST
    Figure CN121815153A_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of audio processing, and provides an audio processing method, apparatus and device, a storage medium and a program product, the method comprising: configuring at least one carrier channel and at least one modulation channel configured with a vocoder plug-in, the vocoder plug-in comprising a carrier sub-channel and a modulation sub-channel; inputting a first audio which comes from a carrier channel and carries a target tone to a carrier sub-channel, and inputting a second audio which carries a target sound effect to a modulation sub-channel; the vocoder plug-in is used for aligning and synthesizing the first audio and the second audio so as to generate the target audio fusing the target tone and the target sound effect, the first audio bearing the basic tone and the second audio bearing the specific sound effect are synthesized so as to generate the audio with the target sound effect and the target tone; the time and economic cost of traditional recording are reduced, and the problem that an existing plug-in is single in generation effect or mechanical is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of audio technology, specifically to audio processing methods, apparatus, devices, storage media, and program products. Background Technology

[0002] In audio design, the biological texture of monster cries primarily stems from simulating the vibrational characteristics of bodily tissue. When unstable airflow passes through a monster's throat, vocal cords, and mouth, it creates air turbulence, which in turn stimulates muscles or soft tissue to produce rhythmically spaced trembling sounds. These trembling sounds, through variations in intensity and rhythm, enhance the biological characteristics and emotional expressiveness of the monster cries, making them a key element in improving player immersion and game entertainment. Currently, the main technical approaches to obtaining such sounds include: recording biological or human voices on location and then post-processing them; using commercial sound effect libraries; synthesizing them in digital audio workstations using sound distortion plugins; or simulating trembling effects using vibrato effects.

[0003] However, the methods for obtaining such sounds in related technologies all have obvious limitations: on-site recording is subject to harsh shooting conditions, performance ability, and animal cooperation, resulting in low efficiency and difficulty in adapting to the fast-paced demands of game development; while commercial sound effect libraries are easy to obtain, they offer low creative freedom and easily lead to homogenization of sound materials among products, affecting the uniqueness of the work; sound morphing plugins have cumbersome operation processes, and the generated results are singular and unstable, failing to meet the requirements of video games for a large number of diverse sound samples; and the tremolo sound produced by the processing method based on vibrato effects is too mechanical and stiff, with obvious traces of artificial processing, greatly reducing the user's immersion. Summary of the Invention

[0004] This disclosure provides an audio processing method, apparatus, device, storage medium, and program product to address the problem of how to improve the efficiency and quality of acquiring audio with tremor effects.

[0005] In a first aspect, this disclosure provides an audio processing method, the method comprising: Configure at least one carrier channel and at least one modulation channel configured with a vocoder plugin, wherein the vocoder plugin includes a carrier sub-channel and a modulation sub-channel; The first audio signal, which carries the target timbre, is input from the carrier channel to the carrier sub-channel, and the second audio signal, which carries the target sound effect, is input to the modulation sub-channel. The vocoder plugin is used to align and synthesize the first audio and the second audio to generate a target audio that blends the target timbre and the target sound effect.

[0006] Secondly, this disclosure provides an audio processing apparatus, the apparatus comprising: A configuration module is used to configure at least one carrier channel and at least one modulation channel configured with a vocoder plugin, wherein the vocoder plugin includes a carrier sub-channel and a modulation sub-channel; The input module is used to input a first audio signal from the carrier channel carrying the target timbre to the carrier sub-channel, and to input a second audio signal carrying the target sound effect to the modulation sub-channel; The synthesis module is used to align and synthesize the first audio and the second audio using the vocoder plugin to generate a target audio that blends the target timbre and the target sound effect.

[0007] Thirdly, this disclosure provides a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the audio processing method of the first aspect or any corresponding embodiment described above.

[0008] Fourthly, this disclosure provides a computer-readable storage medium storing computer instructions for causing a computer to perform the audio processing method of the first aspect or any corresponding embodiment described above.

[0009] Fifthly, this disclosure provides a computer program product, including computer instructions for causing a computer to execute the audio processing method of the first aspect or any corresponding embodiment described above.

[0010] The audio processing method disclosed herein includes: configuring at least one carrier channel and at least one modulation channel configured with a vocoder plugin, wherein the vocoder plugin includes a carrier sub-channel and a modulation sub-channel; inputting a first audio from the carrier channel carrying a target timbre to the carrier sub-channel, and inputting a second audio carrying a target sound effect to the modulation sub-channel; aligning and synthesizing the first and second audio using the vocoder plugin to generate a target audio that blends the target timbre and the target sound effect. This disclosure aligns and synthesizes the first audio carrying a basic timbre and the second audio carrying a specific sound effect to quickly generate audio with the target sound effect and the target timbre, which reduces the time and economic cost of traditional recording and overcomes the problem of existing plugins generating single or mechanical effects. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the specific embodiments of this disclosure or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram illustrating an application scenario according to an embodiment of this disclosure; Figure 2 This is a schematic flowchart of a first audio processing method according to an embodiment of the present disclosure; Figure 3 This is a schematic diagram of the first operation of the audio processing device in the audio processing method according to an embodiment of the present disclosure; Figure 4 This is a second operational schematic diagram of the audio processing device in the audio processing method according to an embodiment of the present disclosure; Figure 5 This is a schematic diagram of a third operation of the audio processing device in the audio processing method according to an embodiment of the present disclosure; Figure 6 This is a schematic diagram of a second flow of an audio processing method according to an embodiment of the present disclosure; Figure 7 This is a schematic diagram of a third audio processing method according to an embodiment of the present disclosure; Figure 8 This is a schematic diagram of a fourth operation of the audio processing device in the audio processing method according to an embodiment of the present disclosure; Figure 9 This is a fifth operational schematic diagram of the audio processing device in the audio processing method according to an embodiment of the present disclosure; Figure 10 This is a schematic diagram of the fourth flow of the audio processing method according to an embodiment of the present disclosure; Figure 11 This is a fifth flowchart illustrating an audio processing method according to an embodiment of the present disclosure; Figure 12 This is a structural block diagram of an audio processing apparatus according to an embodiment of the present disclosure; Figure 13 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present disclosure. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0014] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0015] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this disclosure, "a plurality of" means two or more, unless otherwise expressly specified.

[0016] Before providing a detailed description of the embodiments of this disclosure, some of the nouns and terms involved in the embodiments of this disclosure will be explained.

[0017] (1) The bio-texture of monster cries: This refers to the sound design that combines physical logic and physiological characteristics to make the cries of fictional monsters present a sense of realism and persuasiveness that conforms to the laws of natural biology. Bio-texture is not simply imitating existing animals, but rather making logical deductions based on the fictional form, physiological structure and social behavior of the monster, so that the sound has organic characteristics that conform to its setting. For example, if the monster is set to have a huge chest cavity and bony resonating cavity, its cry may incorporate low-frequency chest resonance and hard friction sounds, imitating the respiratory resonance of large mammals and the friction sound of the shell of crustaceans; if it is set to be an aquatic creature, the sound may contain wet bubble sounds or high-frequency water pressure vibrations, as if it is a physical deformation produced when it is transmitted through water. Bio-texture allows the audience to subconsciously recognize the unity between sound and image, that is, the rise and fall of the monster's cry should correspond to the monster's breathing rhythm, the tension of the roar should reflect its muscle burst state, and even the vibrato or pauses under different emotions should respond to its possible nervous system reactions.

[0018] (2) Tremor: refers to a rhythmic, fluctuating acoustic phenomenon formed by the complex interaction between airflow and biological tissues. When a monster shouts, the unstable airflow passing through its throat and vocal cavity causes high-frequency vibrations in its vocal cords, mucous membranes, or muscles. This vibration is not a uniform, continuous sound, but rather a periodic fluctuation superimposed on the main tone. Its essence is the pulsed release of airflow energy due to changes in resistance when passing through narrow or elastic channels. For example, when the airflow breaks through the temporarily closed vocal cord folds, a dense group of sound pressure pulses is formed, which sounds like a trembling sound similar to convulsions; while the repeated adhesion and separation of the oral mucosa under the impact of airflow generates a periodic vibrato similar to a wet, slapping sound. The interval pattern of this tremor sound often reflects the physiological characteristics of the monster. For example, a slower rhythm may correspond to fatigued respiratory muscle control, while rapid tremors may indicate nerve abnormalities or organ lesions.

[0019] (3) Audio: refers to the range of mechanical wave frequencies that the human auditory system can perceive and its related applications. Its physical essence is sound waves formed by the propagation of object vibrations in media such as air. At the physics level, audio usually refers to the frequency range of 20 Hz to 20,000 Hz. This spectrum covers the basic ability of the human ear to perceive the pitch of sound—low-frequency booming, mid-frequency human voice dialogue, and even high-frequency cicada chirping all exist in this acoustic universe. Specifically, sound wave vibrations are converted into electrical signals by a microphone, then encoded into digital audio files, and finally the electrical signals are retranslated into audible sound waves by a loudspeaker.

[0020] (4) Audio processing equipment: refers to the Digital Audio Workstation (DAW) in the field of audio processing technology. It is essentially a comprehensive digital audio studio that integrates recording, editing, mixing and mastering. First, when sound is recorded through a microphone or instrument, the DAW converts the continuous analog sound wave signal into discrete digital signals and records them as audio files. At the same time, the DAW itself or through plugins can also generate MIDI data. MIDI data is not the audio itself, but is used to record performance information such as "which note, how loud, and how long" to drive virtual instruments to produce sound. The DAW arranges audio blocks or MIDI notes on a timeline-based multitrack, mixes them into a complete and coherent audio stream, and can render and export the audio stream as a standard audio file.

[0021] (5) Carrier Channel and Modulation Channel: The carrier channel and modulation channel are the core components of DAW's FM synthesizer plugin. The waveform generated by the carrier channel itself (usually a sine wave) determines the fundamental frequency and pitch of the output audio, but its sound itself is relatively simple. The modulation channel, on the other hand, outputs a specific modulation waveform. The modulation waveform cannot be directly heard by the user but is used to continuously modulate the frequency of the waveform in the carrier channel. For example, when a high-frequency sine wave is used as a modulation signal to drive the carrier channel, it will cause the carrier frequency to fluctuate rapidly and periodically. This modulation depth determines the intensity of the timbre change, while the modulation frequency determines how many new harmonic components are generated. Therefore, in DAW's FM synthesizer plugin, users generate the desired synthesized sound by precisely configuring the hierarchical relationship and parameters between the carrier channel and one or more modulation channels.

[0022] (6) Carrier signal: The basic waveform that carries audio information during signal transmission. It is usually a sine wave of a specific frequency. When the carrier signal is not modulated, the amount of effective information carried by the carrier signal is very small. "Modulation" is essentially to accurately correspond to changes in the information signal by changing one or more basic properties of the carrier signal (such as its amplitude, frequency or phase).

[0023] (7) Modulation signal: A modulation signal is a raw signal that carries audio information to be transmitted or is used to change the characteristics of other signals. It adds its own data, sound, or other information to a base waveform (i.e., carrier signal) that is easier to transmit or process through a specific method. Modulation signals add audio information to carrier signals by changing parameters such as amplitude, frequency, or phase of the carrier signal for long-distance transmission. For example, radio stations modulate audio signals onto high-frequency radio waves for transmission. Modulation signals can also be periodic waveforms generated by low-frequency oscillators, which create a vibrato effect by regularly changing the volume of another audio signal.

[0024] (8) Timbre: refers to the distinguishable sound wave characteristics of different sound-producing bodies (such as musical instruments or human voices) at the same pitch and loudness, mainly determined by the superposition of the fundamental tone and overtones. Its physical essence depends on the combination of overtones produced when the sound-producing body vibrates. The overtone frequencies are integer multiples of the fundamental tone and have low intensity. The dynamic distribution of frequency and intensity in the spectrum constitutes the waveform characteristics of the sound wave. For example, the flute sounds pure because it is dominated by the fundamental frequency, while the oboe sounds sharp and bright because of its rich higher harmonics.

[0025] (9) Sound effects are sound elements created or processed by humans to construct a context, convey information or evoke emotions through auditory channels. Sound effects focus on simulating reality or creating psychological cues. Their scope covers three main levels: First, onomatopoeia, which simulates real sound sources through physical props, such as footsteps, creaking doors and windows, etc. Second, design sound effects, which are often surreal sounds that do not exist in reality, such as the roar of a science fiction spaceship engine, the buzzing of magical energy, the roar of a monster, etc. These sound effects are often created by synthesizers or by layering, transforming and modulating multiple original recordings. Third, ambient sounds, which serve as the acoustic background of a scene, such as the chirping of insects in the forest, the flow of traffic in the city, etc.

[0026] (10) Spectrum analysis: refers to the process of decomposing a complex sound into its internal constituent frequencies. It is characterized by converting the audio signal from the time domain (amplitude changes over time) to the frequency domain (intensity distribution of different frequency components) to obtain the harmonic structure and energy characteristics of the audio.

[0027] (11) Frequency band: refers to the frequency range of sound, usually measured in Hertz, and is divided into ultra-low frequency, low frequency, mid frequency, mid-high frequency, high frequency and extremely high frequency. Each frequency band has a different effect on the characteristics and listening experience of sound. Ultra-low frequency is 20Hz-40Hz, low frequency is 40Hz-200Hz, mid frequency is 500Hz-2kHz, mid-high frequency is 2kHz-4kHz, high frequency is 5kHz-10kHz, and extremely high frequency is 10kHz-20kHz.

[0028] (12) Amplitude: The amplitude of the pressure change generated by the vibration of a sound wave, measured in decibels. The size of the audio amplitude determines the sound intensity, that is, the "loudness" or "volume" that the listener can perceive. The larger the amplitude, the higher the energy carried by the sound, and the louder and more impactful it sounds; the smaller the amplitude, the weaker and softer the sound.

[0029] (13) Vocoder plugin: This refers to an audio processing tool that uses the spectral characteristics of one sound to shape the spectrum of another sound. When working, the vocoder needs to receive two input signals simultaneously: one is the carrier, which is usually a synthesizer tone rich in harmonics, responsible for providing the pitch and basic texture of the final effect; the other is the modulator, which is usually a signal containing clear speech information, responsible for providing the spectral profile. The vocoder performs real-time spectral analysis on the modulator signal to detect its amplitude changes in each frequency band. Then, it applies the dynamic spectrum obtained from this analysis to the carrier signal, and uses filters to control the gain of the carrier in each frequency band, forcing the harmonics of the synthesizer to imitate the shape of the spectrum in the modulator.

[0030] (14) Filter: A filter is a tool that can selectively process the frequency content of a sound signal, allowing a specific range of frequency components to pass through while suppressing or attenuating other unwanted frequencies. Filters filter frequencies by setting one or more cutoff frequencies and slopes. For example, a high-pass filter blocks muffled low-frequency rumbles while allowing bright high-frequency details to pass through, and is often used to eliminate breathing noise in microphone recordings; a low-pass filter filters out harsh high-frequency hissing while retaining warm and full mid-low frequencies. In addition, a band-pass filter retains only a narrow frequency band of sound, while a notch filter precisely removes a specific frequency.

[0031] (15) Instantaneous amplitude energy: used to describe the acoustic intensity or physical energy of a sound signal at a specific instant. The instantaneous amplitude energy of a certain audio signal sample point reflects the signal level intensity at that sampling moment.

[0032] (16) Modulation signal spectral envelope data: refers to the trend of the frequency distribution of the modulation signal over time. Specifically, the modulation signal is the source signal carrying audio change information. The spectrum is used to characterize which frequency components constitute the modulation signal at any given time and their respective intensities. The envelope, on the time axis, generates a smooth, generalized contour line for the overall intensity of these constantly changing frequency components, characterizing the overall fluctuation state of the low-frequency, mid-frequency, and high-frequency energy regions of the modulation signal at different times. For example, when a human voice produces the vowel "Ah", its spectral envelope will form a prominent energy peak in the low-mid frequency region; while when it turns into the sound "See", the energy peak will shift to the high-frequency region.

[0033] As one optional application scenario of this disclosure embodiment, such as Figure 1 As shown, the application system may include at least one terminal device and at least one server. Figure 1 The example shows that the application system includes a computer 101, a mobile terminal 102, and a server 103, and the terminal devices such as the computer 101 and the mobile terminal 102 are connected to the server 103 through the network 110.

[0034] Specifically, the terminal device can be a smartphone, tablet, laptop, PDA, desktop computer, game console, smart TV, smart wearable device, in-vehicle terminal, VR (Virtual Reality) device, AR (Augmented Reality) device, etc. Server 103 can be a standalone physical server, a server cluster, a distributed system, or a cloud server providing cloud services. Network 110 can be a wired or wireless network, examples of which include, but are not limited to, the Internet, corporate intranet, local area network, wide area network, mobile communication network, and combinations thereof.

[0035] It should be noted that, Figure 1 This is merely an example of an application scenario and does not limit the scope of protection of this disclosure.

[0036] It should be understood that the actions described relative to the terminal device can be performed by an application on the terminal device, or by the application in conjunction with its server (e.g., a server).

[0037] In audio design, the biological texture of monster cries primarily stems from simulating the vibrational characteristics of bodily tissue. When unstable airflow passes through a monster's throat, vocal cords, and mouth, it creates air turbulence, which in turn stimulates muscles or soft tissue to produce rhythmically spaced trembling sounds. These trembling sounds, through variations in intensity and rhythm, enhance the biological characteristics and emotional expressiveness of the monster cries, making them a key element in improving player immersion and game entertainment. Currently, the main technical approaches to obtaining such sounds include: recording biological or human voices on location and then post-processing them; using commercial sound effect libraries; synthesizing them in digital audio workstations using sound distortion plugins; or simulating trembling effects using vibrato effects.

[0038] However, the methods for obtaining such sounds in related technologies all have obvious limitations: on-site recording is subject to harsh shooting conditions, performance ability, and animal cooperation, resulting in low efficiency and difficulty in adapting to the fast-paced demands of game development; while commercial sound effect libraries are easy to obtain, they offer low creative freedom and easily lead to homogenization of sound materials among products, affecting the uniqueness of the work; sound morphing plugins have cumbersome operation processes, and the generated results are singular and unstable, failing to meet the requirements of video games for a large number of diverse sound samples; and the tremolo sound produced by the processing method based on vibrato effects is too mechanical and stiff, with obvious traces of artificial processing, greatly reducing the user's immersion.

[0039] Based on this, embodiments of this disclosure provide an audio processing method, apparatus, device, storage medium, and program product. By configuring at least one carrier channel and at least one modulation channel configured with a vocoder plugin, the vocoder plugin includes a carrier sub-channel and a modulation sub-channel. A first audio signal carrying a target timbre from the carrier channel is input to the carrier sub-channel, and a second audio signal carrying a target sound effect is input to the modulation sub-channel. The vocoder plugin is used to align and synthesize the first and second audio signals to generate a target audio signal that blends the target timbre and target sound effect. This disclosure aligns and synthesizes the first audio signal carrying a basic timbre with the second audio signal carrying a specific sound effect to quickly generate audio with both the target sound effect and the target timbre, reducing the time and economic costs of traditional recording and overcoming the problems of existing plugins generating monotonous or mechanical effects.

[0040] According to an embodiment of this disclosure, an audio processing method embodiment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0041] This embodiment provides an audio processing method. Figure 2 This is a flowchart of an audio processing method according to an embodiment of the present disclosure, such as... Figure 2 As shown, the process includes the following steps: Step S201: Configure at least one carrier channel and at least one modulation channel configured with a vocoder plugin, wherein the vocoder plugin includes a carrier sub-channel and a modulation sub-channel.

[0042] In one embodiment of this disclosure, at least one audio processing unit can be configured in an audio processing device. Each audio processing unit includes a carrier channel and a modulation channel configured with a vocoder plugin. The audio processing device refers to a Digital Audio Workstation (DAW), and the audio processing unit refers to a functional module within the DAW. A DAW contains multiple parallel or serial audio processing units, each including a carrier channel and a modulation channel, which together complete complex audio processing tasks.

[0043] The carrier channel and modulation channel are core components of the FM synthesizer plugin in a digital audio workstation. The FM synthesizer plugin is a virtual musical instrument based on frequency modulation technology in a digital audio workstation. It generates complex and varied timbres by interfering with waveforms of different audio frequencies. The core architecture of the FM synthesizer plugin consists of multiple operators, each of which is essentially a simple, independently configurable oscillator. Users connect these operators into a signal network using specific algorithms. When a high-frequency sine wave acts as a modulator on the carrier, it forces the carrier's frequency to fluctuate periodically. The modulation depth determines the harmonic density, and the modulation frequency determines the harmonic interval.

[0044] In this embodiment, the carrier channel is the audio stream path carrying the target timbre. Specifically, the carrier channel routes the audio signal carrying the target timbre to the carrier sub-channels included in the vocoder. The audio signal waveform (such as a sine wave or square wave) output by the carrier channel is used to determine the reference pitch and basic texture of the sound. For example, a carrier channel that generates a stable 440Hz sine wave provides the basis for the "A4" pitch of the final synthesized target audio.

[0045] In this embodiment, the modulation channel is the audio stream path carrying the target sound effect. Specifically, the modulation channel routes the audio signal carrying the target sound effect to the modulation sub-channel contained in the vocoder. The audio signal transmitted by the modulation channel is used for parameterizing control data or secondary audio streams of the audio signal in the carrier channel. For example, a modulation channel that generates a 5Hz triangular wave can be used to periodically change the volume of the carrier channel, thereby creating a vibrato effect.

[0046] A vocoder plugin is a component loaded within the modulation channel of an audio processing device, which uses the spectral profile of one signal to shape the spectrum of another. The carrier subchannel and the modulation subchannel are two independent input interfaces within the vocoder plugin. In embodiments of this disclosure, the modulation subchannel is the audio stream subpath in the vocoder used to receive audio signals from the modulation channel. For example, when a human voice is input, the vocoder will analyze its spectral envelope (i.e., the frequency characteristics of vowel and consonant variations) in real time. In embodiments of this disclosure, the carrier subchannel is the audio stream subpath in the vocoder used to receive audio signals from the carrier channel. For example, when a signal composed of white noise or complex synthesizer timbres is input.

[0047] In one embodiment, at least one audio processing unit is configured in the DAW, and a clear signal flow is established within or between it. For example, the user creates a new audio processing unit in the DAW and loads a virtual instrument with an FM synthesis architecture within that unit, thereby configuring the included carrier channel and modulation channel. A vocoder plugin is inserted into the plugin chain of the modulation channel, and a signal source is specified for each input interface of the vocoder plugin. For example, on the vocoder plugin interface, input connections are established between the carrier sub-channels and the carrier channel, while the modulation sub-channel is configured to receive audio signals from the modulation channel.

[0048] Preferred, Figure 3 This is a schematic diagram illustrating the first operation of the audio processing device in the audio processing method of this disclosure embodiment, referring to... Figure 3 Taking Reaper v6.83 as an example, the DAW software uses its built-in vocoder plugin ReaVocode (Cockos). The modulation track is the track with the ReaVocode plugin mounted, which is used to receive the modulation signal, that is, the second audio.

[0049] Wet is the proportion of the sound after vocoder processing, used to determine the intensity of the final effect. Dry is the proportion of the raw, unprocessed carrier signal. In practice, it needs to be set to -inf to indicate that the carrier signal is completely off, thus avoiding mixing synthesizer timbre with effects. Mod Dry is the proportion of the raw, unprocessed modulation signal. In practice, it needs to be set to -inf to indicate that the modulation signal is completely off, avoiding mixing clear vocals into the output. In other words, during setup and operation, both Dry and Mod Dry should be set to -inf, retaining only the Wet signal to obtain a pure, completely synthesized new timbre from the vocoder.

[0050] The parameter Bands characterizes the number of filter banks within the vocoder. A vocoder can divide the audio spectrum into multiple bands. The more bands, the finer the spectrum division, resulting in higher analysis and processing accuracy, and a clearer, more natural timbre. Fewer bands result in a coarser sound, with a more mechanical or telephone-like tone. Ideally, Bands should be set to 32.

[0051] Stereo is the on / off switch for stereo mode. When enabled, the vocoder performs independent spectral analysis and processing on the left and right channel signals, thus preserving or creating stereo width. This option is disabled if the input is a mono signal.

[0052] Figure 4 This is a second operational schematic diagram of the audio processing device in the audio processing method according to an embodiment of the present disclosure, referring to... Figure 4 Create a second audio track as a carrier track, send its audio channels 1 and 2 to the modulation track's audio channels 3 and 4, and uncheck the "Master send channels from / to" option so that the audio material of this track will not be played on the main channel.

[0053] Figure 5 This is a third operational schematic diagram of the audio processing device in the audio processing method according to an embodiment of the present disclosure, referring to... Figure 5 Click Figure 4 In the "Param" field, open the "Plug-in pin connector" window of the vocoder in the modulation track. Check if the input and output channels are correct. Half of the input channel is the left / right input of the vocoder's detector, and 3 / 4 of the input channel is the left / right input of the vocoder's carrier. Half of the output channel is the left / right output of the vocoder.

[0054] Step S202: Input the first audio from the carrier channel, which carries the target timbre, to the carrier sub-channel, and input the second audio, which carries the target sound effect, to the modulation sub-channel.

[0055] The first audio signal refers to the audio signal output from the carrier channel that carries the target timbre. The second audio signal refers to an audio signal that carries the target sound effect and is independent of the first audio signal.

[0056] In one embodiment, at least one audio processing unit is created in the DAW, and each audio processing unit has a pre-defined carrier channel and a modulation channel. A vocoder plugin is configured on the modulation channel. The carrier channel is connected to a carrier sub-channel, so that after inputting a first audio signal containing the target timbre into the carrier channel, the first audio signal corresponding to the first audio can be directly routed to the carrier sub-channel of the vocoder plugin. Simultaneously, the second audio signal corresponding to the second audio signal containing the target sound effect, input to the modulation channel, is routed to the modulation sub-channel of the vocoder. Specifically, through the vocoder plugin's user interface, the user can specify an audio signal from a certain channel to a specific input sub-channel of the vocoder plugin via a drop-down menu, drag-and-drop operation, or path selection. For example, the user selects and assigns the output bus of the carrier channel "CarrierTrack" from the routing options to the "Carrier" input (carrier sub-channel) of the vocoder plugin in the DAW; simultaneously, the output of the modulation channel "Modulator Track" is assigned to the "Modulator" input (modulation sub-channel) of the same vocoder plugin.

[0057] During operation, the vocoder plugin analyzes the instantaneous amplitude energy and spectral envelope data of the second audio signal through the modulation sub-channel, such as vowel features like "Ah" and "Oh", to obtain the spectral envelope data of the modulation signal. It then applies the instantaneous amplitude energy and spectral envelope data obtained from this analysis to the second audio signal of the carrier sub-channel.

[0058] Step S203: Use a vocoder plugin to align and synthesize the first and second audio to generate a target audio that blends the target timbre and target sound effects.

[0059] In one embodiment, the vocoder plugin performs real-time spectral analysis on the second audio signal of the input modulation sub-channel using a filter bank to extract its instantaneous amplitude energy and time-varying spectral envelope data. For example, it analyzes in real time to determine if the input second audio signal has energy peaks at 500Hz, 1500Hz, etc. at a certain moment.

[0060] When applying instantaneous amplitude energy and spectral envelope data to the second audio signal of the carrier sub-channel, the vocoder plugin adjusts the first audio signal and the second audio signal to be synchronized in time, thereby ensuring that the spectral features extracted from the second audio signal can be applied to the corresponding time point of the first audio signal in real time without delay, so that the generated target audio simultaneously contains the target timbre of the first audio and the target sound effect of the second audio.

[0061] The audio processing method provided in this embodiment configures at least one carrier channel and at least one modulation channel configured with a vocoder plugin. The vocoder plugin includes a carrier sub-channel and a modulation sub-channel. A first audio signal carrying the target timbre from the carrier channel is input to the carrier sub-channel, and a second audio signal carrying the target sound effect is input to the modulation sub-channel. The vocoder plugin is used to align and synthesize the first and second audio signals to generate a target audio signal that blends the target timbre and the target sound effect. This disclosure aligns and synthesizes the first audio signal carrying the basic timbre with the second audio signal carrying the specific sound effect to quickly generate audio with the target sound effect and the target timbre. This reduces the time and economic cost of traditional recording and overcomes the problem of existing plugins generating monotonous or mechanical effects.

[0062] This embodiment provides an audio processing method. Figure 6 This is a flowchart of an audio processing method according to an embodiment of the present disclosure, such as... Figure 6 As shown, the process includes the following steps: Step S601: Configure at least one carrier channel and at least one modulation channel configured with a vocoder plugin. The vocoder plugin includes a carrier sub-channel and a modulation sub-channel. See details below. Figure 2 Step S201 of the illustrated embodiment will not be described again here.

[0063] Step S602: Input the first audio from the carrier channel, which carries the target timbre, to the carrier sub-channel, and input the second audio, which carries the target sound effect, to the modulation sub-channel.

[0064] Specifically, step S602 includes: Step S6021: Obtain the first audio from the carrier channel. The first audio is the audio carrying the target timbre.

[0065] Step S6022: Input the first audio signal to the carrier sub-channel.

[0066] The first audio refers to the audio input from the outside to the carrier channel that carries the target timbre, and the target timbre refers to the timbre that the user expects the final output audio to have.

[0067] Since a connection has been established between the carrier channel and the carrier sub-channel, after the first audio containing the target timbre is input to the carrier channel, the first audio can be directly specified and transmitted to the carrier sub-channel of the vocoder plug-in. The carrier sub-channel is a specific input interface inside the vocoder plug-in for receiving audio signals to be shaped and processed. The carrier sub-channel itself does not produce sound but is used to process externally input audio. For example, on the user interface of the vocoder plug-in, the user sets the output channel of the carrier channel as the input source of the carrier sub-channel in the vocoder through a drop-down menu or routing button. Specifically, on the interface of the vocoder plug-in, find the carrier sub-channel input selector marked with "Carrier" or "载波" and select the FM synthesizer track from the list as its signal source.

[0068] Step S603: Input the second audio carrying the target sound effect into the modulation sub-channel.

[0069] Specifically, the above step S603 includes: Step S6031: Obtain the second audio from the modulation channel. The second audio is the audio carrying the target sound effect.

[0070] Step S6032: Input the second audio into the modulation sub-channel.

[0071] Modulation channel: Refers to the path inside the audio processing device for carrying control signals or characteristic signals, used to provide dynamically changing control information. For example, in an FM synthesizer, the output path of an operator that generates a low-frequency triangular wave (such as 2 Hz) to periodically change the pitch of the carrier is the modulation channel.

[0072] The second audio refers to the audio that is externally input to the modulation channel and carries the target sound effect. The target sound effect refers to the sound effect that the user expects the final output target audio to have, which is the contour with undulations, impact, and spectral changes in the audio. For example, the second audio can be the tremor sound when a monster roars.

[0073] Since the vocoder plug-in is set on the modulation channel, there is a natural connection between the modulation channel and the modulation sub-channel. To avoid inputting the second audio in the modulation channel into the carrier sub-channel, the output channel of the modulation channel is set to only include the modulation sub-channel. The modulation sub-channel will analyze in real-time the spectral characteristics of the second audio signal corresponding to the second audio input from the modulation channel and generate corresponding control data. Specifically, on the panel of the vocoder plug-in, the user finds a modulation sub-channel input selector marked with "Modulator" or "分析端", and then selects the modulation channel from the drop-down menu as its input.

[0074] The audio processing method provided in this embodiment achieves the fusion of timbre base and dynamic features by accurately routing the first audio carrying the target timbre and the second audio carrying the target sound effect to the carrier sub-channel and modulation sub-channel of the vocoder, respectively. This breaks through the limitation of traditional vocoders that only process human voices and can perform real-time deep synthesis of any timbre and any sound effect.

[0075] This embodiment provides an audio processing method. Figure 7 This is a flowchart of an audio processing method according to an embodiment of the present disclosure, such as... Figure 7 As shown, the process includes the following steps: Step S701: Configure at least one carrier channel and at least one modulation channel configured with a vocoder plugin. The vocoder plugin includes a carrier sub-channel and a modulation sub-channel. See details below. Figure 2 Step S201 of the illustrated embodiment will not be described again here.

[0076] Step S702: Input the first audio signal from the carrier channel, carrying the target timbre, to the carrier sub-channel, and input the second audio signal carrying the target sound effect to the modulation sub-channel. For details, please refer to [link to relevant documentation]. Figure 2 Step S202 of the illustrated embodiment will not be described again here.

[0077] Step S703: Use a vocoder plugin to align and synthesize the first and second audio to generate a target audio that blends the target timbre and target sound effects.

[0078] Specifically, step S703 includes: Step S7031: The carrier signal corresponding to the first audio and the modulation signal corresponding to the second audio are respectively input into a filter group with the same center frequency distribution, and the filter group includes at least one bandpass filter.

[0079] The carrier signal refers to the audio signal corresponding to the first audio level, which is the foundation for carrying the target timbre. The modulation signal refers to the audio signal corresponding to the second audio level, which is the control source for carrying the target sound effect characteristics.

[0080] A filter bank is a set of parallel bandpass filters that work together. Each filter in the filter bank is responsible for processing a specific, adjacent frequency band.

[0081] The center frequency distribution refers to the arrangement of the center frequencies allowed by each bandpass filter in a filter bank across the entire audio spectrum. The same center frequency distribution indicates that two filter banks, one for processing the carrier signal and the other for processing the modulation signal, are structurally mirror-symmetrical. For example, if the first filter bank has bandpass filters with center frequencies of 100Hz, 500Hz, 2000Hz, etc., along the modulation signal path, then the second filter bank will also have identical bandpass filters with center frequencies of 100Hz, 500Hz, 2000Hz, etc., along the carrier signal path. A bandpass filter is a filter that allows only signal components within a specific frequency range to pass through, while significantly attenuating frequency components outside that range. For example, a bandpass filter with a center frequency of 1kHz and a bandwidth of 100Hz primarily allows frequency components between 950Hz and 1050Hz to pass through. Specifically, the carrier signal and the modulation signal are input synchronously but independently into two structurally identical filter banks.

[0082] For example, suppose there is a filter bank consisting of 32 bandpass filters, with the center frequencies of the filter bank distributed according to a specific pattern between 80Hz and 12kHz. The carrier signal carrying the target timbre and the modulation signal carrying the target sound effect are input into two filter banks with identical distributions of 32 center frequencies. The first filter bank decomposes the modulation signal into 32 independent narrowband signals. For example, the filter with a center frequency of 250Hz outputs the energy changes in speech related to that frequency band, while the filter with a center frequency of 4kHz outputs the energy of hissing sounds in speech. Simultaneously, the second filter bank also synchronously decomposes the carrier signal into 32 corresponding narrowband signals.

[0083] Furthermore, the amplitude dynamics of the modulated signal at the output of filter N will be used to specifically control the gain of the carrier signal at the output of filter N. This one-to-one correspondence ensures that the spectral profile of the modulated signal can be appended to the spectrum of the carrier signal with extremely high fidelity and time synchronization.

[0084] For example, when importing audio into the modulation track, this audio material needs to contain the sound dynamics expected in the final effect. The signal required in the modulation track can be the audio material, the MIDI signal driving the sound synthesizer, or the audio signal exported by the sound synthesizer.

[0085] Insert a sound synthesizer plugin before the vocoder on the modulation track, drive the synthesizer with a MIDI signal to obtain the desired sound effect, insert an EQ effect after the vocoder to adjust the frequency of the synthesizer sound, and then modulate the synthesizer parameters through the track envelope to make the sound effect change dynamically over time.

[0086] Figure 8 This is a schematic diagram illustrating the fourth operation of the audio processing device in the audio processing method according to an embodiment of the present disclosure. Figure 8 The signal diagrams of the four channels shown are those of the MIDI signal. The audio file generated by the sound synthesizer is exported and placed into the modulation track "Modulator Track" as a new audio material, thereby obtaining the modulation signal that carries the target sound effect.

[0087] Figure 9 This is a fifth operational schematic diagram of the audio processing device in the audio processing method according to an embodiment of the present disclosure. Figure 9 The signal diagrams of the four channels to the right of the "Modulator Track" are the signal diagrams of the modulation signal, and the signal diagrams of the four channels to the right of the "Carrier Track" are the signal diagrams of the carrier signal. The first audio is imported into the "Carrier Track". The first audio determines the timbre of the newly generated target audio, that is, this timbre is used to play the sound effects dynamics and part of the sound content in the modulation track.

[0088] Step S7032: Perform real-time spectrum analysis on the modulated signal using a filter bank to extract the instantaneous amplitude energy of the modulated signal in the corresponding frequency band of each bandpass filter.

[0089] Real-time spectrum analysis refers to the continuous and non-delayed frequency component analysis of a modulated signal to respond instantly to changes in the signal.

[0090] A bandpass filter's corresponding frequency band refers to a specific frequency range defined by a single bandpass filter in a filter bank. The entire filter bank covers the entire spectrum of the audio signal, while each filter is responsible for a narrow segment of the spectrum. For example, a bandpass filter with a center frequency of 1kHz and a bandwidth of 100Hz corresponds to a frequency band of 950Hz to 1050Hz.

[0091] Instantaneous amplitude energy refers to the instantaneous value of the acoustic intensity or power of an audio signal in a specific frequency band at a specific point in time. Instantaneous amplitude energy is usually proportional to the square or absolute value of the amplitude of the signal in that frequency band at that moment. For example, in a scenario where the modulated signal is a human voice, at a certain instant when the vowel "Ah" is spoken, a filter corresponding to the 500Hz frequency band may output a relatively high instantaneous amplitude energy value (e.g., 0.8).

[0092] The modulation signal is separated, measured, and output to determine its intensity at different frequency bands. For example, for the output of each bandpass filter, the instantaneous amplitude energy representing the intensity variation of the narrowband signal is extracted using an envelope follower.

[0093] Step S7033: Obtain the time-varying modulation signal spectrum envelope data based on the instantaneous amplitude energy. The modulation signal spectrum envelope data is used to align and synthesize with the carrier signal after passing through the filter bank.

[0094] Modulated signal spectral envelope data is a continuous data stream or sequence that records the evolution of the spectral characteristics of the modulated signal. For example, when the audio changes from "Ah" to "Oh", the energy of the 500Hz band decreases while the energy of the 800Hz band increases, thus revealing a continuous change over time.

[0095] The spectral envelope data of a modulated signal is the spectral energy profile formed by the instantaneous amplitude energy values ​​of all filter channels at any given time point. The envelope line formed by the spectral envelope data of the modulated signal depicts the overall energy distribution of the modulated signal at different frequencies. For example, at a certain time point, the spectral envelope data of the modulated signal shows a profile with significant energy peaks at 200Hz and 1500Hz, thus revealing the vowel characteristics of the modulated signal at that time.

[0096] The dispersed instantaneous amplitude energy values, representing the independent energy of each frequency band, obtained from the aforementioned steps, are aggregated to generate a unified data structure capable of characterizing the overall spectral morphology. For example, in each processing cycle of a digital audio workstation (e.g., each audio buffer), N instantaneous amplitude energy values ​​from all N bandpass filter channels are collected and arranged in frequency order to form a spectral frame for the current moment. For instance, the filter bank continuously extracts the instantaneous amplitude energy of each frequency band and combines the instantaneous amplitude energy of all channels at a preset time resolution (e.g., 44,100 times per second) to obtain spectral envelope data.

[0097] Step S7034: The carrier signal is divided into several carrier sub-signals of different frequency bands by a filter bank.

[0098] Different frequency bands refer to specific frequency ranges that do not overlap or overlap only slightly, as defined by the center frequency and bandwidth of each bandpass filter in the filter bank. For example, one frequency band might be 80Hz to 120Hz, and another might be 3800Hz to 4200Hz.

[0099] By utilizing the frequency selectivity of filter banks, the input wideband carrier signal is decomposed and separated into multiple parallel, narrowband signals with limited frequency ranges, i.e., carrier sub-signals. Each carrier sub-signal contains only the harmonics or noise content of the original carrier signal that belongs to that narrowband frequency range.

[0100] Step S7035: The modulation signal spectrum envelope data is used as a control signal to perform amplitude modulation on several carrier sub-signals of different frequency bands respectively.

[0101] In one embodiment, the real-time acquired modulation signal spectrum envelope data is converted into multiple parallel control voltage signals, which are strictly synchronized with the carrier sub-signals generated by filter bank segmentation in both time and frequency band.

[0102] Each carrier sub-signal in a frequency band is connected to an independent voltage-controlled amplifier, and its corresponding control signal comes from the spectral envelope data of the same frequency band. When the modulating signal energy of a certain frequency band increases, the control voltage of that frequency band increases accordingly, driving the voltage-controlled amplifier to increase the gain of the corresponding carrier sub-signal; conversely, when the modulating signal energy decreases, the control voltage decreases, causing the carrier sub-signal gain to decrease.

[0103] Reference Figure 10 In some optional implementations, step S4035 includes: Step a1: For each frequency band's carrier sub-signal, obtain the corresponding modulation signal energy value for that frequency band.

[0104] Step a2: Modulate the amplitude of the carrier sub-signal according to the energy value of the modulation signal.

[0105] First, a real-time frequency band correspondence is established. Specifically, for the Nth frequency band carrier sub-signal generated by filter bank segmentation (such as a narrowband signal with a center frequency of 1kHz), the modulation signal energy value belonging to the same Nth frequency band is synchronously acquired from the spectrum analysis module. This modulation signal energy value represents the instantaneous amplitude intensity of the modulation signal near 1kHz at this moment. The modulation signal and the carrier sub-signal are strictly aligned in timing and frequency band, forming an independent control channel.

[0106] During amplitude modulation, the energy value of the modulation signal in each frequency band is converted into a control voltage, which directly drives the voltage-controlled amplifier of the carrier sub-signal in that band. When the energy of the modulation signal in a specific frequency band increases, such as when the energy of the 500Hz band increases during the human voice's vowel "a", the amplitude of the corresponding carrier sub-signal increases accordingly; when the energy decreases, such as when the energy of the 500Hz band decreases during the consonant "s", the amplitude of the carrier sub-signal decreases accordingly. The amplitude modulation process is executed in parallel across all frequency bands, ensuring that the amplitude changes of each frequency component of the carrier signal completely reproduce the spectral envelope characteristics of the modulation signal.

[0107] By precisely controlling the amplitude of each sub-band, the dynamic spectral characteristics of the modulated signal can be imparted to the carrier signal, ultimately outputting a fused audio that maintains both the harmonic structure of the carrier signal and the dynamic changes of the modulated signal.

[0108] Step S7036: The carrier sub-signals after amplitude modulation are superimposed and synthesized to generate the target audio signal. The target audio signal carries the harmonic characteristics of the carrier signal and the spectral morphology characteristics of the modulation signal.

[0109] The amplitude-modulated carrier sub-signals of each frequency band are re-integrated into the hybrid bus for linear superposition. Each carrier sub-signal retains the harmonic components of the original input carrier signal in the corresponding frequency band, but its amplitude envelope has been reshaped by the spectral characteristics of the modulated signal. During the synthesis process, the carrier sub-signals of all frequency bands are recombined according to their frequency distribution to form a complete frequency domain reconstruction.

[0110] Reference Figure 11 In some optional implementations, step S4036 includes: Step b1: Select the start time point and the end time point on the amplitude-modulated carrier sub-signal, and superimpose and synthesize the amplitude-modulated carrier sub-signal according to the start time point and the end time point to generate the target audio signal.

[0111] Step b2: Generate target audio by fusing target timbre and target sound effects based on the target audio signal.

[0112] In one embodiment, the frequency band carrier sub-signals, after amplitude modulation, are defined in the time domain. Specifically, based on preset audio processing requirements, a start time point and an end time point are determined on the time axis. These two time points define the effective range of signal synthesis. In practice, the start time point typically corresponds to the moment when the characteristics of the modulated signal begin to appear, such as the speech start point of the first audio signal, while the end time point corresponds to the moment when the characteristics end. Within this time-domain window, the frequency band carrier sub-signals are linearly superimposed, with each sub-signal retaining its modulated amplitude characteristics, to generate a target audio signal with complete spectral characteristics.

[0113] The target audio signal is digitally encoded and format-converted. Using the rendering engine of a digital audio workstation, the time-domain signal is converted into a standard audio file format. During the conversion process, the target audio signal simultaneously retains the target timbre from the carrier signal and the target sound effects from the modulation signal, achieving a deep and creative fusion of the two audio characteristics. The final generated target audio retains the musicality of the carrier signal while also exhibiting the dynamic effects of the modulation signal.

[0114] The audio processing method provided in this embodiment establishes a precise band-level signal mapping relationship by using a filter bank with the same center frequency distribution to process the carrier signal and the modulation signal in parallel, ensuring the accuracy of spectrum analysis and the synchronization of modulation processing. Secondly, it uses real-time spectral envelope data extracted from the modulation signal as a control signal to precisely modulate the amplitude of the segmented carrier sub-signals, enabling the dynamic characteristics of the carrier signal to fully reproduce the spectral morphology of the modulation signal, thus allowing the two audio signals to be deeply combined while maintaining their respective core characteristics. Finally, through a synthesis mechanism of time-domain segmentation and signal superposition, the generated target audio retains both the harmonic characteristics and timbre of the carrier signal and incorporates the dynamic envelope and spectral characteristics of the modulation signal.

[0115] This embodiment also provides an audio processing apparatus for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0116] This embodiment provides an audio processing device, such as... Figure 12 As shown, it includes: Configuration module 1201 is used to configure at least one carrier channel and at least one modulation channel configured with a vocoder plugin, wherein the vocoder plugin includes a carrier sub-channel and a modulation sub-channel.

[0117] The input module 1202 is used to input the first audio from the carrier channel, which carries the target timbre, to the carrier sub-channel, and to input the second audio carrying the target sound effect to the modulation sub-channel.

[0118] Synthesis module 1203 is used to align and synthesize the first and second audio using a vocoder plugin to generate a target audio that blends the target timbre and target sound effects.

[0119] In some alternative implementations, the input module 1202 includes: The first audio acquisition unit is used to acquire the first audio from the carrier channel, wherein the first audio is the audio carrying the target timbre.

[0120] The first audio input unit is used to input the first audio signal to the carrier sub-channel.

[0121] In some alternative implementations, the input module 1202 further includes: The second audio acquisition unit is used to acquire the second audio from the modulation channel. The second audio is the audio that carries the target sound effect.

[0122] The second audio input unit is used to input the second audio to the modulation sub-channel.

[0123] In some alternative implementations, the synthesis module 1203 includes: An audio input unit is used to input the carrier signal corresponding to the first audio and the modulation signal corresponding to the second audio into a filter bank with the same center frequency distribution, wherein the filter bank includes at least one bandpass filter.

[0124] The spectrum analysis unit is used to perform real-time spectrum analysis on the modulated signal through the filter bank to extract the instantaneous amplitude energy of the modulated signal in the corresponding frequency band of each bandpass filter.

[0125] The envelope data acquisition unit is used to obtain the time-varying spectrum envelope data of the modulated signal based on the instantaneous amplitude energy.

[0126] The signal segmentation unit is used to divide the carrier signal into several carrier sub-signals of different frequency bands through a filter bank.

[0127] The amplitude modulation unit is used to use the spectrum envelope data of the modulation signal as a control signal to perform amplitude modulation on several carrier sub-signals of different frequency bands.

[0128] The signal superposition unit is used to superimpose and synthesize the amplitude-modulated carrier sub-signals to generate the target audio signal, which carries the harmonic characteristics of the carrier signal and the spectral morphology of the modulation signal.

[0129] In some alternative implementations, the amplitude modulation unit includes: The energy acquisition subunit is used to acquire the modulation signal energy value corresponding to the frequency band for each frequency band carrier sub-signal.

[0130] The amplitude modulation subunit is used to modulate the amplitude of the carrier sub-signal according to the energy value of the modulation signal.

[0131] In some alternative implementations, the signal superposition unit includes: The signal superposition subunit is used to select the start time point and the end time point on the amplitude-modulated carrier sub-signal, and to superimpose and synthesize the amplitude-modulated carrier sub-signal according to the start time point and the end time point to generate the target audio signal.

[0132] The audio generation subunit is used to generate target audio that blends the target timbre and target sound effects based on the target audio signal.

[0133] The audio processing apparatus provided in this disclosure can execute the audio processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the various modules and units described above are the same as in the corresponding embodiments described above, and will not be repeated here.

[0134] Figure 13 This is a schematic diagram of the structure of a computer device provided in an embodiment of this disclosure.

[0135] The following is a detailed reference. Figure 13 This diagram illustrates a structural schematic suitable for implementing a computer device according to embodiments of the present disclosure. The computer device may include a processor (e.g., a central processing unit, graphics processor, etc.) 1301, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1302 or a program loaded from memory 1308 into random access memory (RAM) 1303. The RAM 1303 also stores various programs and data required for the operation of the computer device. The processor 1301, ROM 1302, and RAM 1303 are interconnected via a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.

[0136] Typically, the following devices can be connected to I / O interface 1305: input devices 1306 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 1307 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; memory 1308 including, for example, magnetic tape, hard disk, etc.; and communication devices 1309. Communication device 1309 allows the computer device to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 13 Computer equipment with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0137] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 1309, or installed from memory 1308, or installed from ROM 1302. When the computer program is executed by processor 1301, it performs the functions defined in the audio processing method of embodiments of this disclosure.

[0138] Figure 13The computer device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0139] This disclosure also provides a computer-readable storage medium in which the methods described in this disclosure can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium after being downloaded over a network. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium may also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implement the audio processing methods shown in the above embodiments.

[0140] A portion of this disclosure can be applied to computer program products, such as computer program instructions, which, when executed by a computer, can invoke or provide methods and / or technical solutions according to this disclosure through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, and installation package files. Accordingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions; the computer compiling the instructions and then executing the corresponding compiled program; the computer reading and executing the instructions; or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0141] Although embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. An audio processing method, characterized in that, The method includes: Configure at least one carrier channel and at least one modulation channel configured with a vocoder plugin, wherein the vocoder plugin includes a carrier sub-channel and a modulation sub-channel; The first audio signal, which carries the target timbre, is input from the carrier channel to the carrier sub-channel, and the second audio signal, which carries the target sound effect, is input to the modulation sub-channel. The vocoder plugin is used to align and synthesize the first audio and the second audio to generate a target audio that blends the target timbre and the target sound effect.

2. The method according to claim 1, characterized in that, The step of inputting the first audio signal from the carrier channel, which carries the target timbre, to the carrier sub-channel includes: A first audio signal is obtained from the carrier channel, wherein the first audio signal is an audio signal carrying the target timbre; The first audio signal is input to the carrier sub-channel.

3. The method according to claim 1, characterized in that, Inputting the second audio signal carrying the target sound effect into the modulation sub-channel includes: The second audio is obtained from the modulation channel, and the second audio is the audio carrying the target sound effect; The second audio input is sent to the modulation sub-channel.

4. The method according to claim 1, characterized in that, The step of aligning and synthesizing the first audio and the second audio using the vocoder plugin includes: The carrier signal corresponding to the first audio and the modulation signal corresponding to the second audio are respectively input into a filter group with the same center frequency distribution, wherein the filter group includes at least one bandpass filter. The modulation signal is subjected to real-time spectrum analysis using the filter bank to extract the instantaneous amplitude energy of the modulation signal in the corresponding frequency band of each bandpass filter. Based on the instantaneous amplitude energy, the modulation signal spectrum envelope data that varies with time is obtained. The modulation signal spectrum envelope data is used to align and synthesize with the carrier signal that has passed through the filter bank.

5. The method according to claim 4, characterized in that, The step of aligning and synthesizing the first audio and the second audio using the vocoder plugin also includes: The carrier signal is divided into several carrier sub-signals of different frequency bands by the filter bank; The modulation signal spectrum envelope data is used as a control signal to perform amplitude modulation on several carrier sub-signals of different frequency bands respectively; The carrier sub-signals after amplitude modulation are superimposed and synthesized to generate a target audio signal, which carries the harmonic characteristics of the carrier signal and the spectral morphology characteristics of the modulation signal.

6. The method according to claim 5, characterized in that, The step of using the modulated signal spectral envelope data as a control signal to perform amplitude modulation on several carrier sub-signals of different frequency bands includes: For each frequency band's carrier sub-signal, obtain the modulation signal energy value corresponding to that frequency band; The amplitude of the carrier sub-signal is modulated according to the energy value of the modulation signal.

7. The method according to claim 5, characterized in that, The step of superimposing and synthesizing the amplitude-modulated carrier sub-signals to generate the final output audio signal includes: A start time point and an end time point are selected on the amplitude-modulated carrier sub-signal. The amplitude-modulated carrier sub-signal is superimposed and synthesized according to the start time point and the end time point to generate the target audio signal. A target audio signal is generated by fusing the target timbre and the target sound effect based on the target audio signal.

8. An audio processing apparatus, characterized in that, The device includes: A configuration module is used to configure at least one carrier channel and at least one modulation channel configured with a vocoder plugin, wherein the vocoder plugin includes a carrier sub-channel and a modulation sub-channel; The input module is used to input a first audio signal from the carrier channel carrying the target timbre to the carrier sub-channel, and to input a second audio signal carrying the target sound effect to the modulation sub-channel; The synthesis module is used to align and synthesize the first audio and the second audio using the vocoder plugin to generate a target audio that blends the target timbre and the target sound effect.

9. A computer device, characterized in that, include: A memory and a processor are communicatively connected, the memory storing computer instructions, and the processor executing the computer instructions to perform the audio processing method of any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the audio processing method according to any one of claims 1 to 7.

11. A computer program product, characterized in that, Includes computer instructions for causing a computer to perform the audio processing method according to any one of claims 1 to 7.