An audio generation method, apparatus, device, and computer-readable storage medium
By acquiring timbre parameters and combining them with oscillators, filters, and effects to generate audio clips, the problem of timbre monotony in MIDI audio generation methods is solved, achieving timbre plasticity and diversity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
- Filing Date
- 2024-09-30
- Publication Date
- 2026-05-26
AI Technical Summary
Existing MIDI-based audio generation methods offer limited timbre, preventing users from customizing or adjusting the sound.
By acquiring pre-set timbre parameters, combining oscillators, filters, and effects to generate audio streaming, multiple audio segments are generated and superimposed to generate the target audio. The timbre parameters include oscillator control information, filter control information, and effects control information.
It greatly enhances the plasticity of timbre, allowing users to customize timbre parameters, support the expansion of new timbres and the modification of existing timbres, and has the ability to simulate a large number of timbres.
Smart Images

Figure CN119169979B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to an audio generation method, apparatus, device, and computer-readable storage medium. Background Technology
[0002] MIDI (Musical Instrument Digital Interface) is a digital music communication protocol used for communication between electronic musical instruments. The MIDI protocol controls various parameters of musical instruments, such as pitch, volume, timbre, and sound effects, by sending digital commands, thereby enabling music composition, performance, and recording. Using MIDI, a digital music description standard, corresponding audio waveforms can be generated from instructions in MIDI files via MIDI chips. The method of generating audio through rendering using a MIDI chip simulator relies on MIDI format communication, resulting in poor compatibility with other data structures. Furthermore, while MIDI chip rendering uses waveform rendering algorithms to produce specific instrument sounds, users cannot adjust these sounds or add their desired sounds.
[0003] Therefore, current MIDI-based audio generation methods suffer from the technical problem of generating only a single timbre. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide an audio generation method, apparatus, device and computer-readable storage medium, which solves the technical problem of the monotonous timbre in the existing audio generation technology.
[0005] To address the aforementioned technical problems, this invention provides an audio generation method, comprising:
[0006] Acquire multiple pre-set tone parameters; wherein each tone parameter includes oscillator control information, filter control information, and effects control information;
[0007] Based on the timbre parameters, combined with oscillators, filters, and effects, audio streaming is generated to obtain multiple audio segments.
[0008] The target audio is obtained by superimposing multiple audio segments.
[0009] Optionally, the timbre parameters further include waveform information; wherein the waveform information includes at least one of self-made wavetable information, low-frequency oscillator oscillation information, and sequencer information; the waveform information is the timbre parameter required by the oscillator.
[0010] Optionally, the audio streaming generation is performed based on the timbre parameters combined with oscillators, filters, and effects to obtain multiple audio segments, including:
[0011] The oscillator control information is processed using the oscillator to obtain audio sampling point data;
[0012] The audio sampling point data is processed using the filter based on the filter control information to obtain waveform data;
[0013] The waveform data is processed using the effector based on the effector control information to obtain the audio segment.
[0014] Optionally, after superimposing multiple audio segments to obtain the target audio, the method further includes:
[0015] The target audio is input into the audio buffer;
[0016] The duration of the interface parameters is determined based on the downstream interface type; wherein, the downstream interface type is the type of the interface used to export audio.
[0017] The target audio in the audio buffer is spliced into data frames according to the duration of the interface parameters to obtain a long time domain; wherein, the long time domain is unsegmented, smooth audio.
[0018] Optionally, the method further includes:
[0019] Determine whether the duration of each pitch information corresponding to the target audio is greater than the duration of the interface parameter;
[0020] When the duration of the current pitch information is longer than the duration of the interface parameter, the cross-domain information corresponding to the long time domain is stored in an array according to the pitch information, so as to control the end time of the target audio according to the cross-domain information in the array; wherein, the cross-domain information is the pitch information when the durations of two interface parameters in the long time domain are concatenated.
[0021] Optionally, the audio streaming generation is performed based on the timbre parameters combined with oscillators, filters, and effects to obtain multiple audio segments, including:
[0022] The multiple timbre parameters are respectively encapsulated with the oscillator, the filter, and the effect to obtain multiple single-tone vocal processing models; wherein, each single-tone vocal processing model is a model for generating a single timbre;
[0023] The three-stage envelope information is encapsulated into the single-tone vocalization processing model; wherein, the three-stage envelope information is a parameter describing the change of the signal over time;
[0024] Accordingly, the audio streaming generation is performed based on the timbre parameters combined with oscillators, filters, and effects to obtain multiple audio segments, including:
[0025] Based on the timbre parameters, an audio streaming generation is performed using a single-tone vocalization processing model that encapsulates the three-stage envelope information to obtain the multiple audio segments.
[0026] Optionally, during the processing using the oscillator, the method further includes:
[0027] The wavetable index of the oscillator is set by the number of beats per minute, and the generated waveform data is updated based on the wavetable index; wherein, the wavetable index is a file that determines the order of waveform samples.
[0028] Optionally, before generating multiple audio segments by combining the timbre parameters with oscillators, filters, and effects, the method further includes:
[0029] Depending on the specific requirements, some or all of the oscillators can be selected, along with some or all of the filters, as oscillators for audio streaming generation. When there are multiple oscillators, each oscillator supports a different oscillation method, and when there are multiple filters, each filter supports a different filtering algorithm.
[0030] This application also provides an audio generation apparatus, comprising:
[0031] The timbre parameter acquisition module is used to acquire multiple pre-set timbre parameters; wherein each timbre parameter includes oscillator control information, filter control information, and effects control information;
[0032] The audio segment generation module is used to generate multiple audio segments by combining the timbre parameters with an oscillator, filter and effects.
[0033] The audio determination module is used to superimpose multiple audio segments to obtain the target audio.
[0034] This application also provides an audio generation device, including:
[0035] Memory, used to store computer programs;
[0036] A processor for executing the computer program to implement the steps of the audio generation method described above.
[0037] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the audio generation method described above.
[0038] This application also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the audio generation method described above.
[0039] As can be seen, this invention acquires multiple pre-set timbre parameters, including oscillator control information, filter control information, and effects control information; it then uses these timbre parameters in conjunction with the oscillator, filter, and effects to generate multiple audio segments via audio streaming; finally, it superimposes these audio segments to obtain the target audio. The beneficial effects of this invention are: it allows for the design of timbre parameters, enabling audio generation based on these parameters in conjunction with the oscillator, filter, and effects. Because the timbre parameters can be designed, expanding upon existing timbres and modifying new ones only requires updating the timbre parameter data. Furthermore, the invention offers significant potential for timbre versatility, possessing the ability to simulate a large number of timbres.
[0040] In addition, the present invention also provides an audio generation apparatus, device, and computer-readable storage medium, which also have the above-mentioned beneficial effects. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0042] Figure 1 A flowchart of an audio generation method provided in an embodiment of the present invention;
[0043] Figure 2 A flowchart illustrating an audio generation method provided in an embodiment of the present invention;
[0044] Figure 3 This is a schematic diagram of an oscillator processing method provided in an embodiment of the present invention;
[0045] Figure 4 A schematic diagram of a Clamp method provided in an embodiment of the present invention;
[0046] Figure 5 A schematic diagram of a single-tone sound processing module provided in an embodiment of the present invention;
[0047] Figure 6 This is a schematic diagram illustrating long-term processing according to an embodiment of the present invention;
[0048] Figure 7 This is a schematic diagram of the structure of an audio generation device provided in an embodiment of the present invention;
[0049] Figure 8 This is a schematic diagram of the structure of an audio generation device provided in an embodiment of the present invention. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0051] Please refer to Figure 1 , Figure 1 A flowchart illustrating an audio generation method provided in an embodiment of the present invention. The method may include:
[0052] S101, acquires multiple preset tone parameters; each tone parameter includes oscillator control information, filter control information, and effects control information.
[0053] The execution subject of this embodiment is an electronic device, such as a computer; or the electronic device in this embodiment can also be a device composed of oscillators, filters, etc. The tone parameters in this embodiment are set by the user according to their needs. This embodiment does not limit specific tone parameters. For example, the tone parameters in this embodiment may include oscillator control information (harmonic richness, envelope information, etc.), filter control information, effects control information, and waveform modulation information; or the tone parameters in this embodiment may include oscillator control information, filter control information, effects control information, waveform modulation information, and self-made wavetable information; or the tone parameters in this embodiment may include oscillator control information, filter control information, and effects control information. The oscillator control information in this embodiment mainly includes frequency control information, phase control information, amplitude control information, and waveform control information; the filter control information in this embodiment refers to the parameters and instructions used to control the behavior and performance of the filter; the effects control information in this embodiment refers to various parameters used to adjust and control the effects unit to produce specific sound effects; the timbre parameters in this embodiment may also include waveform modulation information, which refers to various parameters and settings used to control and adjust the waveform modulation process in order to produce specific signal characteristics according to application requirements.
[0054] It should be further explained that, to improve the audio rendering effect, the aforementioned timbre parameters may also include waveform information; wherein, the waveform information includes at least one of self-made wavetable information, low-frequency oscillator oscillation information, and sequencer information; the waveform information represents the timbre parameters required by the oscillator. In this embodiment, the self-made wavetable information, low-frequency oscillator oscillation information, and sequencer information are set according to user needs. In this embodiment, the self-made wavetable information refers to digitized sound data sampled from real musical instrument timbres, stored in a wavetable file for use in audio playback synthesis. In this embodiment, the low-frequency oscillator oscillation information refers to low-frequency signals or waveforms generated by electronic circuitry. In this embodiment, the sequencer information is information for creating arpeggio effects. In this embodiment, the waveform information can be input to the oscillator so that the oscillator processes the waveform information.
[0055] S102 generates multiple audio segments by combining timbre parameters with oscillators, filters, and effects.
[0056] In this embodiment, the oscillator is the core of the audio signal generation, producing a corresponding waveform based on pitch information. In complex synthesis systems, different types of oscillators (such as sine waves, square waves, sawtooth waves, etc.) may be used to enrich the sound's timbre. The filter in this embodiment is used to shape the audio signal's spectrum, removing or enhancing specific frequency components to change the timbre. For example, a low-pass filter can remove high-frequency components, making the sound smoother; a high-pass filter can remove low-frequency components, making the sound clearer. The effects unit in this embodiment is used to add specific sound effects, such as reverb, delay, and distortion, to increase the sound's spatiality and dynamic range. Streaming generation in this embodiment refers to the audio being continuously generated according to a preset time period. The generated audio stream needs to be cut or segmented, i.e., the start and end points of each segment are determined based on duration information and other musical parameters. In this embodiment, oscillator control information in the timbre parameters is input to the oscillator, filter control information to the filter, and effects control information to the effects unit. After the code angle is input, the parameters are directly defined. This embodiment can combine vocal information, including pitch and duration information, during the generation process. In this embodiment, the pitch information is typically directly related to the oscillator frequency; the higher the frequency, the higher the generated pitch. In digital music synthesis, this is usually achieved by changing the oscillator frequency to achieve different pitches. The duration information in this embodiment determines the duration of the note, which in audio synthesis is typically influenced by controlling the envelope or modulation depth to affect the oscillator's output time. The duration information determines when the oscillator starts and stops generating signals of a specific frequency. This solution achieves real-time rendering output through streaming generation and can be deployed on the edge. Unlike MIDI, this solution does not require maintaining a continuous audio stream; resource allocation only occurs when a call signal occurs, and the process can be freely terminated.
[0057] It should be further explained that, in order to improve the accuracy of processing using oscillators, filters, and effects, the above-mentioned audio streaming generation based on sound information and timbre parameters, combined with oscillators, filters, and effects, to obtain multiple audio segments, may include: processing oscillator control information using an oscillator to obtain audio sampling point data; processing the audio sampling point data using a filter based on filter control information to obtain waveform data; and processing the waveform data using an effects unit based on effects unit control information to obtain audio segments. In this embodiment, the oscillator is an electronic component or device that converts DC power into AC power with a certain frequency. In this embodiment, the filter is a device or circuit that processes signals. In this embodiment, the effects unit is an electronic instrument specifically designed to generate various sound effects, used to change the waveform of the original sound, modulate or delay the phase of the sound wave, enhance the harmonic components of the sound wave, etc., to produce various special sound effects. This embodiment does not limit the number of oscillators. For example, the oscillator in this embodiment can be a single oscillator; or the oscillator in this embodiment can be multiple oscillators. After inputting oscillator control information (mainly harmonic richness, envelope information, etc.) into the oscillator, audio sampling point data is obtained, which is the original waveform data. In this embodiment, the filter can modify the filter control information in the audio sampling point data by changing different frequency bands to obtain waveform data. In this embodiment, waveform data refers to digital information representing sound waves or other vibration waveforms, typically composed of a series of sampling points, each containing information such as amplitude (and sometimes phase). In this embodiment, processing the waveform data using effects can include effects chains such as distortion, delay, reverb, and flanger. This embodiment improves the accuracy of processing using oscillators, filters, and effects by providing a method for continuous processing using oscillators, filters, and effects.
[0058] It needs further explanation that, to ensure the regular changes in the attack and end of each note, audio streaming is performed based on timbre parameters combined with oscillators, filters, and effects to obtain multiple audio segments. This can include: encapsulating multiple timbre parameters with oscillators, filters, and effects respectively to obtain multiple single-tone vocal processing models; where each single-tone vocal processing model generates a single timbre; encapsulating three-stage envelope information into the single-tone vocal processing model; where the three-stage envelope information is a parameter describing the signal's change over time; correspondingly, audio streaming is performed based on timbre parameters combined with oscillators, filters, and effects to obtain multiple audio segments. This can include: using the single-tone vocal processing model encapsulated with three-stage envelope information to perform audio streaming based on timbre parameters to obtain multiple audio segments. In this embodiment, using timbre parameters, oscillators, filters, and effects as single-tone vocal processing models means that each timbre parameter, oscillator, filter, and effect is treated as a whole to generate the audio model. In this embodiment, the three-stage envelope acts on the entire audio. The three-stage envelope state in this embodiment is a commonly used method in audio processing and music production to describe and control the four stages of sound: attack, decay, sustain, and release, often abbreviated as the ADSR model. It can be understood that when the sound processing module starts receiving parameters, it automatically enters a three-stage envelope state process, which primarily simulates the attack of the sound. When the sound processing module ends, it automatically enters the decay process. In this embodiment, the three-stage envelope state is processed by the oscillator and filter, and the data after the three-stage envelope is then fed to the effects unit for further processing.
[0059] It should be further explained that, to ensure the continuity of the basic waveform, the processing using the oscillator may further include: setting the oscillator's wavetable index by beats per minute (BPM), and updating the generated waveform data based on the wavetable index; wherein, the wavetable index is a file that determines the order of waveform samples. In this embodiment, the wavetable index is used for each frame-level generation, serving as a parameter for the oscillator at the very beginning of generation. In this embodiment, beats per minute (BPM) is used to describe the tempo or rhythm of the music. It can be understood that a wavetable is a file containing multiple waveform samples, each corresponding to a specific index value. Based on the current wavetable index, the corresponding waveform sample is read from the wavetable and used to generate the audio signal. As the music rhythm changes, the wavetable index is dynamically updated, resulting in different waveform data. This embodiment ensures the continuity of the basic waveform by setting and updating the oscillator wavetable index.
[0060] It should be further explained that, in order to improve the timbre rendering effect, before generating multiple audio segments by combining oscillators, filters, and effects based on timbre parameters, the following steps are also included:
[0061] Depending on different needs, some or all oscillators and some or all filters can be selected as oscillators for audio streaming generation. When multiple oscillators are used, each supports a different oscillation method; when multiple filters are used, each supports a different filtering algorithm. Understandably, to ensure the capability and diversity of generated timbres, this method uses three oscillator modules for waveform generation, allowing the selection of some or all oscillators based on different requirements. Each module supports multiple different oscillation methods, including analog oscillators (generating basic sine waves, square waves, etc.), frequency-modulated oscillators (controlling waveform information through modulation signals), and wavetable oscillators (accepting customized waveform wavetables), as well as digital signal processing algorithms. The algorithm includes built-in wavetable files of basic waveforms for basic waveform overlay, and allows for waveform deformation and modulation based on information during oscillation. In this embodiment, three filters can be used for waveform correction, each supporting different filtering algorithms. Filters can modify the target audio by altering different frequency bands. This solution uses a biquad filter to achieve basic high-pass, low-pass, and band-pass functions. Simultaneously, this method integrates unique filtering techniques such as comb filtering and formant filtering to quickly create distinctive timbres. Like oscillators, the filter parameters are all controllable.
[0062] S103, superimpose multiple audio segments to obtain the target audio.
[0063] In this embodiment, overlaying multiple audio segments refers to accurately aligning all audio segments in time according to their pitch to ensure they belong to the same audio. After obtaining the target audio, this embodiment can input the target audio into an audio buffer. The audio data stored in the audio buffer can be accessed and transmitted through various downstream interfaces (such as I2S (Inter-IC Sound, a serial bus interface standard), USB (Universal Serial Bus, a widely used serial bus standard), and HDMI (High-Definition Multimedia Interface, a digital interface supporting high-resolution video transmission) to meet the needs of different devices and applications. These interfaces should have high data transfer rates and low latency to ensure the real-time performance and accuracy of the audio data.
[0064] It should be further explained that, in order to achieve streaming generation, after superimposing multiple audio segments to obtain the target audio, the process may also include: inputting the target audio into an audio buffer; determining the duration of the interface parameters according to the downstream interface type; wherein, the downstream interface type is the type of the interface used to export the audio; and splicing the target audio in the audio buffer into data frames according to the duration of the interface parameters to obtain the long time domain; wherein, the long time domain is the unsegmented, smooth audio.
[0065] This embodiment requires inputting the target audio into an audio buffer. In this embodiment, the long time domain refers to a duration without any segmentation; it represents uninterrupted, smooth audio that the user is unaware of, as it is output frame-by-frame and perceived as smooth audio. This embodiment can determine the interface parameter duration based on the downstream interface type. For example, if the current downstream interface type is a short time domain of 5 seconds, (5 seconds x 441000 sample points) divided by 1024 (the number of sample points per frame) corresponds to an interface parameter duration of 2^14. Common audio export interface types in this embodiment include: PCM (Pulse Code Modulation): an uncompressed audio format commonly used for transmitting audio data between professional audio equipment and software; WAV: a lossless audio format widely used for storing uncompressed audio data; MP3 / AAC: a lossy compressed audio format suitable for music distribution and streaming services; and USB / Bluetooth: interfaces used to transmit audio data to external devices. This embodiment, by using a long time domain related to the interface parameter duration, can be applied to various downstream audio export interface types for convenient downstream interface calls.
[0066] It should be further explained that, in order to improve the efficiency of audio generation, the above method may also include: determining whether the duration of each pitch information corresponding to the target audio is greater than the duration of the interface parameter; when the duration of the current pitch information is greater than the duration of the interface parameter, storing the cross-domain information corresponding to the long time domain into an array according to the pitch information, so as to control the end time of the target audio according to the cross-domain information in the array; wherein, the cross-domain information is the pitch information when the durations of two interface parameters are concatenated in the long time domain. In this embodiment, the duration of the pitch corresponding to the pitch information (duration information) is the length of time during which the pitch information remains stable within a certain period of time, or the time span during which a single note or sound maintains its pitch unchanged. For ease of understanding, if the length of the data frame is 512 sampling points and the duration of the receiving interface parameter is 441000 sampling points, when a sound signal with a pitch of 43 is received at a certain interface and lasts for more than one interface duration, the data will be recorded in a map array updated once every interface duration in the format of (pitch, number of frames to be sustained). The map array is a data structure that maps unique keys to values. In this embodiment, recording multiple audio segments in an array based on pitch information means that the data is recorded in pitch format in an array that is updated once every interface duration. It's understood that pitch information is a crucial parameter in audio processing, determining the fundamental frequency (FFM) of a note. By extracting and recording the pitch information of audio segments in an array, these data can be easily sorted, analyzed, or processed. This embodiment can store audio segments in the array based on pitch information when the duration of the pitch information exceeds the interface parameter duration, and not store them in the array when it is not greater than the interface parameter duration, instead reading the generated audio content normally. This reduces the storage frequency and controls the end time of the target audio based on cross-domain information in the array, improving the efficiency and accuracy of audio generation.
[0067] The audio generation method provided in this invention may include: S101, acquiring multiple pre-set timbre parameters; wherein each timbre parameter includes oscillator control information, filter control information, and effects control information; S102, generating multiple audio segments by combining the timbre parameters with the oscillator, filter, and effects; S103, superimposing the multiple audio segments to obtain the target audio. Compared with the current method that can only control timbre parameters based on the MIDI protocol, making the timbre parameters unchangeable, this invention allows for the design of timbre parameters, thereby generating audio based on the timbre parameters. Since the timbre parameters can be set according to requirements, expanding new timbres and modifying existing timbres only requires updating the text parameter data, and the timbre has great plasticity potential, possessing the ability to simulate a large number of timbres.
[0068] For a clearer understanding of this invention, please refer to the following details. Figure 2 , Figure 2 This is a flowchart illustrating an audio generation method provided in an embodiment of the present invention, which may specifically include:
[0069] S201 constructs multiple timbre parameters and uses these parameters, oscillators, filters, and effects to create multiple monophonic sound processing units; each monophonic sound processing unit produces a different timbre.
[0070] In this embodiment, a set of parameters is constructed for each tone and saved in XML (eXtensible Markup Language) format. Specifically, a complete tone parameter should contain information in the following formats: oscillator control information, filter control information, effects control information, and waveform modulation information. Some tones may require waveform information, such as custom wavetable information, low-frequency oscillator (LFO) oscillation information, sequencer information, etc. The storage method uses "id" as the parameter name identifier and "value" as the numerical information. The numerical information is stored in text format but will be converted to the corresponding data format after being read. A basic example of a tone parameter representation: id="osc1-freq" value="0.99999341114"; where id="osc1-freq" represents the wavelet length of the first envelope, and value="0.99999341114" represents the numerical form corresponding to the tone parameter text.
[0071] To ensure the capability and diversity of timbre generation, this embodiment uses three oscillator modules for waveform generation. Some or all of the oscillators can be selected based on different needs. Each module supports multiple oscillation methods, including analog oscillators (generating basic sine waves, square waves, etc.), frequency modulation oscillators (controlling waveform information through modulation signals), and wavetable oscillators (accepting customized waveform wavetables), as well as digital signal processing algorithms. The algorithm includes built-in wavetable files of basic waveforms for easy basic waveform overlay. During oscillation, the waveform can be deformed and modulated based on the information. After determining the wavetable, audio sampling point data can be extracted based on index values. All oscillator parameters (mainly fundamental frequency control, modulation ratio, etc.) and oscillator types can be controllably changed. Inputting the timbre parameters into the oscillator yields audio sampling point data; this waveform data is the original waveform data. For easier understanding, please refer to [link to documentation]. Figure 3 , Figure 3 This is a schematic diagram of an oscillator processing method provided in an embodiment of the present invention, wherein the parameter data is the parameter data corresponding to the input timbre parameters.
[0072] In this embodiment, the oscillator output is transmitted to a filter bank. Three filters are also reserved for waveform correction during this process, each supporting different filtering algorithms. The filters can modify the target audio by altering different frequency bands. This solution uses a biquad filter to implement basic high-pass, low-pass, and band-pass functions. Simultaneously, this method integrates unique filtering techniques such as comb filtering and formant filtering to quickly create a distinctive timbre. Like the oscillator, the filter parameters are all controllable. After processing by the filter bank, the oscillator's audio data output yields waveform data with harmonic patterns relatively similar to the target timbre.
[0073] This embodiment also provides a series of effects processing schemes to adjust the generated sound. These include effects chains such as distortion, delay, reverb, and flanger. The essence of distortion is to selectively destroy the waveform at a certain threshold, and this destruction method includes methods such as clamping, folding, and zeroing. Figure 4 This diagram illustrates a Clamp method provided in an embodiment of the present invention. In the diagram, Threshold represents the threshold value. The horizontal line represents the volume threshold for Clamping, and the entire curve represents a waveform graph, with time on the horizontal axis and volume on the vertical axis. Delay, flanging, and reverb algorithms essentially fuse the modified signal (wet signal) with the original signal (dry signal) to obtain the final signal. Delay buffers and delays audio in the time domain, providing a basic sound effect. Specifically, users can control the delay time, delay energy, and delay duration to adjust the value of the feedback signal after energy transformation and filtering. Reverb is very similar to delay, both processing the dry signal to obtain a wet signal and then superimposing it. Its function is also to simulate the formation of reflected sound to construct a sound field environment. Flanging adds the signal with the time difference to the original signal, but unlike the original signal, the time difference varies, resulting in a different phase cancellation effect. The waveform data generated by the filter will be processed by the effects unit in sequence and finally sent to an audio buffer for streaming. Users can also control the on / off state and relative position of each effect unit through parameters.
[0074] S202, activate multiple single-tone sound processing units in the single-tone sound processing module according to a preset time period, and sequentially cycle through multiple activated single-tone sound processing units to obtain multiple audio segments.
[0075] This application only requires timbre parameters, pitch information, and duration information for streaming generation, making it highly adaptable to different upstream tasks. This solution enables real-time rendering output and can be deployed on the client-side. Unlike MIDI, this solution does not require maintaining a continuous audio stream; resource allocation only occurs when a call signal is generated, and the process can be terminated freely. In this solution, the smallest waveform generation unit is a single-tone processing model (unit sound processing module). A single-tone processing model can emit a single tone, and each tone encapsulates the aforementioned timbre parameter reading, oscillator, filter, and effects methods, and independently receives external information parameters. Within the overall sound processing module, multiple activated single-tone processing models (up to 24) are sequentially looped and their audio is superimposed. The final result is input into an audio buffer for downstream interface calls. For easier understanding, please refer to [link to relevant documentation]. Figure 5 , Figure 5 This diagram illustrates a single-tone sound processing module according to an embodiment of the present invention. As shown, multiple single-tone sound processing modules constitute a single sound processing module. Each single-tone sound processing module includes timbre parameters, an oscillator, a filter, and effects. Users can customize the input timbre parameter information (including oscillator control information, filter control information, effects control information, waveform modulation information, custom wavetable information, low-frequency oscillator oscillation information, and sequencer information), as well as duration and pitch information. The input parameters are processed by the oscillator, filter, and effects to generate a frame-level audio stream file, which is then fed into the audio buffer. Streaming output is adaptively performed based on the temporal domain size of the interface.
[0076] In this embodiment, the oscillator wavetable index can be updated first by setting external BPM information, thus ensuring the continuity of the basic waveform. Secondly, ADSR envelope design is used to ensure that the attack and end of each note change regularly. When the sound processing module starts receiving parameters, it automatically enters a three-stage envelope state process, which mainly simulates the onset of the sound. When the sound processing module ends, it automatically enters the decay process. In the effects chain, because signals with time differences need to be mixed, there is also a buffer for the dry audio within the effects chain, with continuously updated parameter smoothing to handle the time-domain variations of the effects.
[0077] S203: Superimpose multiple audio segments to obtain the target audio, and input the target audio into the audio buffer.
[0078] S204, determine the duration of interface parameters based on the downstream interface type.
[0079] S205: Based on the duration of the interface parameters, the target audio in the audio buffer is spliced into data frames to obtain the long time domain, which is then played as audio.
[0080] Understandably, since the input of interface information may not be a frame-level streaming input, this solution also designs a two-layer timing logic in the temporal generation logic of data frames to facilitate the transmission of interface parameter information. The overall temporal logic is as follows: the interface releases generation signals to the generation module according to controllable (integer multiples of frame) interface parameter duration intervals; the sound processing module (multiple monotone sound processing modules) decomposes the generation task into data frame outputs and responds to the upper-level fetch commands through the audio buffer. It should be noted that this embodiment can combine the macro definition of basic parameters (some specific parameters) (binding multiple basic parameters as a group for collective modification) as the interaction interface (the interaction interface refers to the interface with the user), making it more controllable. This solution proposes a streaming timbre rendering method from parameters to audio, which can become a downstream interface for symbolic audio generation, making the audio generation process more complete.
[0081] In this embodiment, during time-domain processing at the unit's interface parameter level, in addition to processing the relevant information received by the interface, it also retrieves unfinished audio termination requests inherited from the past and controls their termination time. For easier understanding, please refer to... Figure 6 , Figure 6 This is a schematic diagram illustrating the processing of long time domains according to an embodiment of the present invention. Figure 6 The interface cross-domain information refers to information that does not exist independently in the temporal domain of a single interface. The duration of the interface parameter is definable. The interface cross-domain information continuously affects the audio of each interface segment. In this embodiment, the duration of the interface parameter is determined according to requirements. For example, 5 seconds corresponds to 210 frames. If we take a common interface, the short temporal domain of an interface is 5 seconds. 5 seconds x 441000 sampling points divided by 1024 (the number of sampling points per frame) corresponds to an interface parameter duration of 214.
[0082] This solution mainly comprises four aspects: data reading, waveform generation, effects processing, and streaming logic. This embodiment of the invention only requires pitch and duration information for streaming generation, making it highly adaptable to different upstream tasks. Furthermore, the method achieves lightweight design through parameter-based synthesis, eliminating the need for waveform sampling. For expanding to new timbres and modifying existing timbres, this method only requires updating text parameter data, and it possesses significant potential for timbre simulation, capable of simulating a large number of timbres. This method can serve as a downstream interface for symbol generation tasks, enabling frame-level real-time timbre rendering and audio generation.
[0083] The audio generation apparatus provided in the embodiments of the present invention will be described below. The audio generation apparatus described below and the audio generation method described above can be referred to in correspondence.
[0084] Please refer to the details. Figure 7 , Figure 7 A schematic diagram of an audio generation device provided in an embodiment of the present invention may include:
[0085] The timbre parameter acquisition module 100 is used to acquire multiple preset timbre parameters; wherein each timbre parameter includes oscillator control information, filter control information, and effects control information;
[0086] The audio segment generation module 200 is used to generate multiple audio segments by combining the timbre parameters with an oscillator, a filter and an effects processor.
[0087] The audio determination module 300 is used to superimpose multiple audio segments to obtain the target audio.
[0088] Furthermore, based on the above embodiments, the timbre parameters further include waveform information; wherein, the waveform information includes at least one of self-made wavetable information, low-frequency oscillator oscillation information, and sequencer information; the waveform information is the timbre parameter required by the oscillator.
[0089] Furthermore, based on the above embodiments, the audio segment generation module 200 may include:
[0090] An oscillator processing unit is used to process the oscillator control information using the oscillator to obtain audio sampling point data;
[0091] A filter processing unit is used to process the audio sampling point data using the filter based on the filter control information to obtain waveform data;
[0092] The effects processing unit is used to process the waveform data using the effects unit based on the effects unit control information to obtain the audio segment.
[0093] Furthermore, based on any of the above embodiments, the audio generation apparatus may further include:
[0094] The input to the audio buffer module is used to input the target audio into the audio buffer;
[0095] The interface parameter duration determination module is used to determine the interface parameter duration based on the downstream interface type; wherein, the downstream interface type is the type to which the interface used for exporting audio belongs;
[0096] The long-term domain determination module is used to concatenate the target audio in the audio buffer according to the duration of the interface parameters, and obtain the long-term domain; wherein, the long-term domain is unsegmented, smooth audio.
[0097] Furthermore, based on the above embodiments, the audio generation apparatus may further include:
[0098] The judgment module is used to determine whether the duration of each pitch information corresponding to the target audio is greater than the duration of the interface parameter.
[0099] The storage to array module is used to store the cross-domain information corresponding to the long time domain into the array according to the pitch information when the duration of the current pitch information is greater than the duration of the interface parameter, so as to control the end time of the target audio according to the cross-domain information in the array; wherein, the cross-domain information is the pitch information when the durations of two interface parameters are concatenated in the long time domain.
[0100] Furthermore, based on the above embodiments, the audio segment generation module 200 may include:
[0101] An encapsulation module is used to encapsulate multiple timbre parameters with the oscillator, the filter, and the effect unit respectively to obtain multiple single-tone sound processing models; wherein, each single-tone sound processing model is a model for generating a single timbre;
[0102] A three-stage envelope information encapsulation module is used to encapsulate three-stage envelope information into the single-tone sound processing model; wherein, the three-stage envelope information is a parameter describing the process of signal change over time;
[0103] Accordingly, the audio segment generation module 200 includes:
[0104] The audio segment generation unit based on three-stage envelope information is used to generate the multiple audio segments by using a single-tone vocalization processing model that encapsulates the three-stage envelope information according to the timbre parameters.
[0105] Furthermore, based on any of the above embodiments, the audio generation apparatus may further include:
[0106] The wavetable index determination module is used to set the wavetable index of the oscillator by the number of beats per minute, and update the generated waveform data based on the wavetable index; wherein, the wavetable index is a file that determines the order of waveform samples.
[0107] Furthermore, based on any of the above embodiments, the audio generation apparatus may further include:
[0108] The oscillator and filter determination module is used to select some or all of the oscillators from the oscillators, and to select some or all of the filters as oscillators for audio streaming generation according to different requirements; wherein, when there are multiple oscillators, each oscillator supports a different oscillation method, and when there are multiple filters, each filter supports a different filtering algorithm.
[0109] It should be noted that the order of the modules and units in the aforementioned audio generation device can be changed without affecting the logic.
[0110] The audio generation device provided in this embodiment of the invention may include: a timbre parameter acquisition module 100, used to acquire a plurality of pre-set timbre parameters; wherein each timbre parameter includes oscillator control information, filter control information, and effects control information; an audio segment generation module 200, used to perform audio streaming generation based on the timbre parameters in combination with the oscillator, filter, and effects to obtain a plurality of audio segments; and an audio determination module 300, used to superimpose the plurality of audio segments to obtain a target audio. Compared with the current method that can only control timbre parameters based on the MIDI protocol, making the timbre parameters unchangeable, this invention allows for the design of timbre parameters, thereby generating audio based on the timbre parameters. Since the timbre parameters can be set according to requirements, the expansion of new timbres and the modification of existing timbres only require updating the text parameter data, and the plasticity potential of the timbres is huge, possessing the ability to simulate a large number of timbres.
[0111] The following describes an audio generation device provided by an embodiment of the present invention. The audio generation device described below can be referred to in correspondence with the audio generation method described above.
[0112] Please refer to Figure 8 , Figure 8 A schematic diagram of an audio generation device provided in an embodiment of the present invention may include:
[0113] Memory 10 is used to store computer programs;
[0114] Processor 20 is used to execute computer programs to implement the audio generation method described above.
[0115] The memory 10, processor 20, and communication interface 30 all communicate with each other through the communication bus 40.
[0116] In this embodiment of the invention, the memory 10 is used to store one or more programs. The programs may include program code, which includes computer operation instructions. In this embodiment of the invention, the memory 10 may store programs for implementing the following functions:
[0117] Acquire multiple pre-set tone parameters; each tone parameter includes oscillator control information, filter control information, and effects control information;
[0118] Based on the timbre parameters, combined with oscillators, filters, and effects, multiple audio segments are generated using audio streaming.
[0119] Multiple audio segments are superimposed to obtain the target audio.
[0120] In one possible implementation, the memory 10 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; and the data storage area may store data created during use.
[0121] Furthermore, memory 10 may include read-only memory and random access memory, providing instructions and data to the processor. A portion of the memory may also include NVRAM. The memory stores operating systems and operating instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, wherein the operating instructions may include various operating instructions for implementing various operations. The operating system may include various system programs for implementing various basic tasks and handling hardware-based tasks.
[0122] Processor 20 can be a central processing unit (CPU), an application-specific integrated circuit, a digital signal processor, a field-programmable gate array, or other programmable logic device. Processor 20 can be a microprocessor or any conventional processor. Processor 20 can call programs stored in memory 10.
[0123] The communication interface 30 can be an interface for the communication module, used to connect with other devices or systems.
[0124] Of course, it should be noted that, Figure 8 The structure shown does not constitute a limitation on the audio generation device in the embodiments of the present invention. In practical applications, the audio generation device may include devices such as audio generators. Figure 8 More or fewer components as shown, or combinations of certain components.
[0125] The following describes the computer-readable storage medium provided in the embodiments of the present invention. The computer-readable storage medium described below can be referred to in correspondence with the audio generation method described above.
[0126] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described audio generation method.
[0127] The computer-readable storage medium may include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0128] The present invention also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the audio generation method described above.
[0129] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0130] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0131] Finally, it should be noted that in this document, relationships such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0132] The above provides a detailed description of an audio generation method, apparatus, device, and computer-readable storage medium provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. An audio generation method, characterized in that, include: Acquire multiple pre-set tone parameters; wherein each tone parameter includes oscillator control information, filter control information, and effects control information; Based on the timbre parameters, combined with oscillators, filters, and effects, audio streaming is generated to obtain multiple audio segments. The target audio is obtained by superimposing multiple audio segments. The target audio is input into the audio buffer; The duration of the interface parameters is determined based on the downstream interface type; wherein, the downstream interface type is the type of the interface used to export audio. The target audio in the audio buffer is concatenated into data frames according to the duration of the interface parameters to obtain a long-term domain; wherein, the long-term domain is unsegmented, smooth audio. Determine whether the duration of each pitch information corresponding to the target audio is greater than the duration of the interface parameter; When the duration of the current pitch information is longer than the duration of the interface parameter, the cross-domain information corresponding to the long time domain is stored in an array according to the pitch information, so as to control the end time of the target audio according to the cross-domain information in the array; wherein, the cross-domain information is the pitch information when the durations of two interface parameters in the long time domain are concatenated.
2. The audio generation method according to claim 1, characterized in that, The timbre parameters also include waveform information; wherein, the waveform information includes at least one of self-made wavetable information, low-frequency oscillator oscillation information, and sequencer information; the waveform information is the timbre parameter required by the oscillator.
3. The audio generation method according to claim 1, characterized in that, The audio streaming generation is performed based on the timbre parameters, combined with oscillators, filters, and effects to obtain multiple audio segments, including: The oscillator control information is processed using the oscillator to obtain audio sampling point data; The audio sampling point data is processed using the filter based on the filter control information to obtain waveform data; The waveform data is processed using the effector based on the effector control information to obtain the audio segment.
4. The audio generation method according to claim 1, characterized in that, The audio streaming generation is performed based on the timbre parameters, combined with oscillators, filters, and effects to obtain multiple audio segments, including: The multiple timbre parameters are respectively encapsulated with the oscillator, the filter, and the effect to obtain multiple single-tone vocal processing models; wherein, each single-tone vocal processing model is a model for generating a single timbre; The three-stage envelope information is encapsulated into the single-tone vocalization processing model; wherein, the three-stage envelope information is a parameter describing the change of the signal over time; Accordingly, the audio streaming generation is performed based on the timbre parameters combined with oscillators, filters, and effects to obtain multiple audio segments, including: Based on the timbre parameters, an audio streaming generation is performed using a single-tone vocalization processing model that encapsulates the three-stage envelope information to obtain the multiple audio segments.
5. The audio generation method according to claim 1, characterized in that, In the process of processing using the oscillator, the method further includes: The wavetable index of the oscillator is set by the number of beats per minute, and the generated waveform data is updated based on the wavetable index; wherein, the wavetable index is a file that determines the order of waveform samples.
6. The audio generation method according to claim 1, characterized in that, Before generating multiple audio segments by combining the timbre parameters with oscillators, filters, and effects, the process also includes: Depending on the specific requirements, some or all of the oscillators can be selected, along with some or all of the filters, as oscillators for audio streaming generation. When there are multiple oscillators, each oscillator supports a different oscillation method, and when there are multiple filters, each filter supports a different filtering algorithm.
7. An audio generation apparatus, characterized in that, include: The timbre parameter acquisition module is used to acquire multiple pre-set timbre parameters; wherein each timbre parameter includes oscillator control information, filter control information, and effects control information; The audio segment generation module is used to generate multiple audio segments by combining the timbre parameters with an oscillator, filter and effects. An audio determination module is used to superimpose multiple audio segments to obtain a target audio. The input to the audio buffer module is used to input the target audio into the audio buffer; The interface parameter duration determination module is used to determine the interface parameter duration based on the downstream interface type; wherein, the downstream interface type is the type to which the interface used for exporting audio belongs; The long-term domain determination module is used to concatenate the target audio in the audio buffer according to the duration of the interface parameters, and obtain the long-term domain; wherein, the long-term domain is unsegmented, smooth audio; The judgment module is used to determine whether the duration of each pitch information corresponding to the target audio is greater than the duration of the interface parameter. The storage to array module is used to store the cross-domain information corresponding to the long time domain into the array according to the pitch information when the duration of the current pitch information is greater than the duration of the interface parameter, so as to control the end time of the target audio according to the cross-domain information in the array; wherein, the cross-domain information is the pitch information when the durations of two interface parameters are concatenated in the long time domain.
8. An audio generation device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the audio generation method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the audio generation method as described in any one of claims 1 to 6.