Drum audio processing method and apparatus and electronic drum
Through the audio processing method of generating and modulating oscillating waves and noise waves in real time, the problem of unrealistic electronic drum tone simulation and limited custom tone is solved, high-quality and low-latency tone output is achieved, and the drummer's expression ability is improved.
Patent Information
- Application Number
- PCT/CN2024/139817
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-25
- Filing Date
- 2024-12-17
- Publication Date
- 2025-07-03
AI Technical Summary
The tone simulation of existing electronic drums is not realistic, the quantity and quality of tones are limited by memory capacity and processing power, and the inability to customize the tones, limiting the drummer's creativity.
By obtaining digital audio signals, oscillating waves and noise waves are generated in real time, and modulated to control pitch, volume and tone, neural network training and filters are used to perform sound effects processing to achieve mixing output.
It realizes dynamic and realistic tone output, reduces dependence on storage, reduces latency, and improves drummer's expressiveness and creativity.
Smart Images

Figure CN2024139817_03072025_PF_FP_ABST
Abstract
Description
Drum audio processing method, device and electric drum Technical Field
[0001] The present application relates to the technical field of music equipment, and in particular to a method and device for processing drum audio, and an electric drum. Background Art
[0002] With the development of musical equipment, musical instruments are changing with each passing day. Among them, electronic drums have already occupied a place in the musical instrument market. They are loved by many people for their portability, small size, low noise and variable tone.
[0003] In the prior art, the working principle of electronic drums is: when the drummer hits the drumhead, the trigger converts the hitting action into electronic signals; then, the sound module receives these electronic signals and selects and plays corresponding drum samples or synthesized drum sounds based on the strength of these electronic signals.
[0004] Although this working principle of electronic drums can already provide quite good tone and expressiveness, there are still some limitations.
[0005] 1) Electronic drum timbres typically consist of pre-recorded drum samples (or synthesized timbres) stored in the drum's internal memory. When a drummer strikes the electronic drum, sensors (such as pressure sensors or resistive touch sensors) detect the intensity of the strike and, based on the detected intensity, select and play the corresponding drum sample from the internal memory. However, pre-recorded drum samples or timbres cannot fully simulate the timbres of real drums, particularly in terms of dynamic expression and timbre transitions.
[0006] 2) Because each drum sample must be recorded and stored individually, the number and quality of electronic drum tones are limited by the internal memory capacity and sound processing capabilities of the electronic drum. If the internal memory capacity or sound processing capabilities of the electronic drum are insufficient, the drummer will not be able to achieve satisfactory sound effects.
[0007] 3) In existing electronic drum technology, the timbre processing is fixed. Drummers can only choose preset timbres and cannot customize the timbre or adjust the details of the timbre, which limits the drummer's creativity and expression.
[0008] In summary, current electronic drum technology has some problems in terms of sound simulation, sound quantity and quality, and sound customization. Summary of the Invention
[0009] Based on this, it is necessary to provide a drum audio processing method, device and electric drum to address the above technical problems, which can perform sound effect processing and mixing output in real time, and dynamically output a tone that is very close to the real one, with sound quality that does not depend on storage and has low latency.
[0010] A drum audio processing method includes: obtaining a drum digital audio signal, performing real-time sound effect processing on the digital audio signal, and performing mixed sound output.
[0011] In one embodiment, obtaining a digital audio signal of a drum, performing real-time sound effect processing on the digital audio signal, and mixing and outputting the mixed sound include:
[0012] acquiring a digital audio signal of a drum, and generating an oscillation wave and a noise wave according to the digital audio signal;
[0013] Modulating the oscillation wave to obtain oscillation wave parameters for controlling pitch and volume;
[0014] Modulating the noise wave to obtain noise wave parameters for controlling timbre;
[0015] Real-time sound effect processing is performed according to the oscillation wave parameters and the noise wave parameters, and mixed sound is output.
[0016] In one embodiment, the oscillation wave is modulated to obtain oscillation wave parameters for controlling pitch and volume, including:
[0017] There are multiple oscillation waves, and different envelopes are used to modulate each oscillation wave to obtain multiple envelope-modulated oscillation waves, and the envelope-modulated oscillation waves include: oscillation wave parameters for controlling pitch and volume; i1 (t) = A i sin(2πf i t) V i (t)=max(E(t),V i (t-dt)*a i ) x i2 (t)=(A i +d i1 V i (t))sin(2π(f i +d i2 V i (t))t)
[0018] Where x i1 (t) is the i-th oscillation wave, A i is the steady-state amplitude of the i-th oscillation wave, f i is the steady-state frequency of the i-th oscillation wave, V i(t) is the envelope of the i-th oscillation wave at time point t, E(t) is the drum signal energy received at time point t, V i (t-dt) is the envelope of the i-th oscillation wave at the time point t-dt, dt is a fixed short time, a i is the envelope attenuation coefficient of the i-th oscillation wave, x i2 (t) is the i-th envelope modulated oscillation wave, d i1 is the amplitude modulation depth of the i-th envelope modulated oscillation wave, d i2 is the frequency modulation depth of the i-th envelope modulated oscillation wave;
[0019] The oscillation wave parameters include steady-state amplitude, steady-state frequency, amplitude modulation depth, frequency modulation depth, and envelope attenuation coefficient.
[0020] In one embodiment, the noise wave is modulated to obtain noise wave parameters for controlling the timbre, including:
[0021] The noise wave is divided into a plurality of noise frequency bands, and each noise frequency band is modulated using a different envelope to obtain a plurality of envelope modulated noise waves, wherein the envelope modulated noise wave includes: noise wave parameters for controlling the timbre; V noise, j(t)=max(E(t),V noise,j (t-dt)*a noisej )
[0022] Where V noise,j (t) is the jth noise frequency band, V noise,j (t-dt) is the envelope of the jth noise band at the time point t-dt, dt is a fixed short time, a noisej is the jth noise band envelope attenuation coefficient, x noise,j (t) is the envelope modulated noise wave of the jth noise frequency band, A noise,j is the base volume of the j-th noise band, d noise,j is the jth noise band modulation depth, V noise,j (t) is the envelope of the j-th noise band at time t, H j is j-1 frequency division filter, rand(t) is a random number generator;
[0023] Noise wave parameters include: noise band base volume, noise band modulation depth, and noise band envelope attenuation coefficient.
[0024] In one embodiment, performing real-time sound effect processing and mixing output according to the oscillation wave parameters and the noise wave parameters includes:
[0025] Designing a forward propagation function according to the oscillation wave parameters and the noise wave parameters; performing time-spectrum calculation on the forward propagation function to obtain a target time-spectrum matrix;
[0026] Get the real drum sound and perform spectrogram calculation to get the real spectrogram matrix;
[0027] The oscillation wave parameters and the noise wave parameters are input into a neural network for training, and the similarity between the target time spectrum matrix and the real time spectrum matrix is used as a cost function. When it is judged that the cost function meets a preset condition, the neural network training stops, and the optimal oscillation wave parameters and the optimal noise wave parameters are output to obtain the optimal oscillation wave and the optimal noise wave;
[0028] The optimal oscillation wave and the optimal noise wave are input into the filter to control the timbre, then input into the amplifier to control the volume, and finally the mixed signal is output.
[0029] In one embodiment, designing a forward propagation function according to the oscillation wave parameters and the noise wave parameters includes:
[0030] Where S(t) is the forward propagation function, x i (t) is the i-th envelope modulated oscillation wave, x noise (t) is the envelope modulated noise wave.
[0031] In one embodiment, performing a time-spectrum graph calculation on the forward propagation function to obtain a target time-spectrum matrix includes:
[0032] The forward propagation function is uniformly sampled, a Hanning window is added, and a fast Fourier transform is performed to obtain a target time-spectrum matrix.
[0033] In one embodiment, the similarity between the target time spectrum matrix and the real time spectrum matrix is used as the cost function, including: J = ∑(F(t,f)-F target (t,f)) 2
[0034] Where J is the cost function, F(t,f) is the true time spectrum matrix, and F target (t,f) is the target time-spectrum matrix.
[0035] In one embodiment, the preset conditions include:
[0036] The number of training times reaches 300,000 or the cost function value is less than 0.01%.
[0037] A drum audio processing device, comprising:
[0038] an acquisition module, configured to acquire an audio signal of a drum and generate an oscillation wave and a noise wave according to the audio signal;
[0039] an oscillation wave modulation module, for modulating the oscillation wave to obtain oscillation wave parameters for controlling pitch and volume;
[0040] A noise wave modulation module, used to modulate the noise wave to obtain noise wave parameters for controlling the timbre;
[0041] The output module is used to perform real-time sound effect processing and mixed sound output according to the oscillation wave parameters and the noise wave parameters.
[0042] Electroacoustic drums, including: drum pad, receiver, collector and processor;
[0043] The drum pad generates an audio signal according to the striking sound;
[0044] The microphone collects the audio signal and sends it to the collector;
[0045] The collector converts the audio signal into an audio digital signal and sends the signal to the processor;
[0046] The processor uses the drum audio processing method to perform real-time sound effect processing and output a mixed signal.
[0047] The aforementioned drum audio processing method, device, and electronic drum set offer the advantages of realistic dynamics, sound quality independent of storage, and low latency. First, traditional triggering requires quantifying the force of the strike and converting it into a specific wavetable for playback. In this application, however, physical modeling is performed based on the collected vibration signal. Different recorded signals result in different output signals, resulting in subtle differences in the timbre of each strike. This is especially true for unique techniques such as brushing and rolls, which are faithfully reflected in the final output timbre. Even differences in drumsticks and the tightness of the mesh can affect the sound. This approach closely resembles real drums and is unattainable with traditional electronic drums. Second, the sound quality of traditional wavetables directly relies on extensive storage. For example, different strike positions and forces correspond to different wavetables. This application, however, directly utilizes sound effects, ensuring high precision in all processing and achieving lossless drum sound quality without relying on storage. Finally, because drums are rhythm instruments, latency requirements are particularly stringent. This application does not have any trigger judgment and does not need to play the wavetable. Instead, it converts the beating audio signal into drum sound in real time, thereby achieving ultra-low latency. This ultra-low latency will give the drummer a very good improved experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] FIG1 is a schematic flow chart of a drum audio processing method according to an embodiment;
[0049] FIG2 is a schematic diagram of the architecture of a drum audio processing method according to one embodiment;
[0050] FIG3 is a block diagram of a drum audio processing apparatus according to an embodiment;
[0051] FIG. 4 is a schematic diagram of an electric acoustic drum according to an embodiment. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely for explaining this application and are not intended to limit this application. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments in this application without creative work are within the scope of protection of this application.
[0053] It should be noted that all directional indications in the embodiments of the present application (such as up, down, left, right, front, back, etc.) are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.
[0054] In addition, the terms "first," "second," and so on, used in this application are for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, features specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this application, "multiple groups" means at least two groups, such as two groups, three groups, and so on, unless otherwise specifically defined.
[0055] In this application, unless otherwise specified or limited, the terms "connect," "fix," etc. should be understood in a broad sense. For example, "fix" can mean a fixed connection, a detachable connection, or an integral connection; it can mean a mechanical connection, an electrical connection, a physical connection, or a wireless communication connection; it can mean a direct connection or an indirect connection through an intermediate medium; it can mean internal communication between two elements or an interaction between two elements, unless otherwise specified. Those skilled in the art will understand the specific meanings of the above terms in this application based on the specific circumstances.
[0056] In addition, the technical solutions between the various embodiments of the present application can be combined with each other, but it must be based on the fact that ordinary technicians in this field can implement it. When the combination of technical solutions is mutually contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.
[0057] The present application provides a method for drum audio processing, comprising: obtaining a digital audio signal of the drum, performing real-time sound effect processing on the digital audio signal, and performing mixing and outputting the mixed signal.
[0058] As shown in FIG1 , in one embodiment, specifically, it includes:
[0059] Step 102: Acquire a digital audio signal of a drum, and generate an oscillation wave and a noise wave according to the digital audio signal.
[0060] In this step, multiple oscillation waves are generated with different frequencies, and there is one noise wave. The oscillation wave is a wave with periodic characteristics, which can be emitted by an oscillator. Specifically, it can be a sine wave, a triangle wave, a pulse wave, etc. It can also form resonances of different pitches through digital wave guide technology.
[0061] For example, each drum pad always has 10 sine wave oscillators that emit 10 sine waves of different frequencies, as well as a noise wave with timbre characteristics.
[0062] As for how to obtain the digital audio signal and how to generate the oscillation wave and the noise wave according to the digital audio signal, these are all existing technologies and will not be described in detail here.
[0063] Step 104: modulate the oscillation wave to obtain oscillation wave parameters for controlling pitch and volume.
[0064] Specifically:
[0065] There are multiple oscillation waves, and different envelopes are used to modulate each oscillation wave to obtain multiple envelope-modulated oscillation waves. The envelope-modulated oscillation waves include: oscillation wave parameters that control pitch and volume; x i1 (t) = A i sin(2πf i t) V i (t)=max(E(t),V i (t-dt)*a i ) x i2 (t)=(A i +d i1 V i (t))sin(2π(f i +d i2 V i (t))t)
[0066] Where x i1 (t) is the i-th oscillation wave (since the system has no specific initial time and the frequencies of the oscillation waves are often not multiples of each other, there is no need to consider the initial phase problem, so the phase coefficient is ignored. Here, a sine wave is used as an example), Ai is the steady-state amplitude of the i-th oscillation wave, f i is the steady-state frequency of the i-th oscillation wave, V i (t) is the envelope of the i-th oscillation wave at time point t, E(t) is the drum signal energy received at time point t, V i (t-dt) is the envelope of the i-th oscillation wave at the time point t-dt, which is also the envelope of the previous time point. dt is a fixed short time, a i is the envelope attenuation coefficient of the i-th oscillation wave, x i2 (t) is the i-th envelope modulated oscillation wave, d i1 is the amplitude modulation depth of the i-th envelope modulated oscillation wave, d i2 is the frequency modulation depth of the i-th envelope modulated oscillation wave;
[0067] Oscillation wave parameters include: steady-state amplitude A i , steady-state frequency f i , amplitude modulation depth d i1 , frequency modulation depth d i2 and the envelope attenuation coefficient a i .
[0068] In this step, multiple envelopes with different envelope attenuation coefficients are used to modulate the oscillation wave. Specifically, the amplitude and frequency of each oscillation wave are modulated separately, and an independent control depth is given to obtain a modulated oscillation wave, namely an envelope-modulated oscillation wave. The dynamic timbre of the modulated oscillation wave is determined by five dynamic parameters, namely the oscillation wave parameters.
[0069] For different oscillation waves, the modulation process is the same, but the envelopes are different, the envelope attenuation coefficients are different, and the corresponding oscillation wave parameters are also different. Of course, for different oscillation waves, the modulation process is the same, the envelopes are the same, and the corresponding oscillation wave parameters are also different.
[0070] Step 106: modulate the noise wave to obtain noise wave parameters for controlling the timbre.
[0071] Specifically:
[0072] The noise wave is divided into multiple noise frequency bands, and each noise frequency band is modulated using different envelopes to obtain multiple envelope modulated noise waves. The envelope modulated noise waves include: noise wave parameters for controlling the timbre; V noise,j (t)=max(E(t),V noise,j (t-dt)*a noisej )
[0073] Where V noise,j (t) is the jth noise frequency band, V noise,j(t-dt) is the envelope of the jth noise band at the time point t-dt, dt is a fixed short time, a noisej is the jth noise band envelope attenuation coefficient, x noise,j (t) is the envelope modulated noise wave of the jth noise frequency band, A noise,j is the base volume of the j-th noise band, d noise,j is the jth noise band modulation depth, V noise,j (t) is the envelope of the j-th noise band at time t, H j is a j-1 frequency division filter, rand(t) is a random number generator used to generate white noise;
[0074] Noise wave parameters include: noise band base volume A noise,j , noise band modulation depth d noise,j And the noise band envelope attenuation coefficient a noisej .
[0075] In this step, multiple envelopes with different envelope attenuation coefficients are used to modulate the noise wave. Specifically, the base volume of each noise frequency band in the noise wave is modulated separately to obtain a modulated noise wave, namely, an envelope modulated noise wave. The modulated noise wave is determined by three dynamic parameters, namely, the noise wave parameters, thereby completing the dynamic processing of the noise timbre and volume.
[0076] For different noise frequency bands, the modulation process is the same, but the envelopes are different, the envelope attenuation coefficients are different, and the corresponding noise wave parameters are also different.
[0077] A crossover filter with j-1 different frequencies can generate j noise bands. For example, a crossover filter with four frequencies of 100, 400, 2000, and 4000 can divide a noise wave into five noise bands. j can be any value between 2 and 256. A larger value for j increases the computational complexity but improves accuracy.
[0078] Step 108 : Perform real-time sound effect processing according to the oscillation wave parameters and the noise wave parameters, and perform mixed sound output.
[0079] Specifically:
[0080] According to the oscillation wave parameters and the noise wave parameters, a forward propagation function is designed; the time spectrum of the forward propagation function is calculated to obtain the target time spectrum matrix;
[0081] Get the real drum sound and perform spectrogram calculation to get the real spectrogram matrix;
[0082] The oscillation wave parameters and noise wave parameters are input into the neural network for training, and the similarity between the target time spectrum matrix and the real time spectrum matrix is used as the cost function. When the cost function is judged to meet the preset conditions, the neural network training stops and the optimal oscillation wave parameters and optimal noise wave parameters are output to obtain the optimal oscillation wave and optimal noise wave.
[0083] The optimal oscillation wave and the optimal noise wave are input into the filter to control the timbre, then input into the amplifier to control the volume, and finally the mixed signal is output.
[0084] More specifically:
[0085] According to the oscillation wave parameters and noise wave parameters, the forward propagation function is designed:
[0086] Where S(t) is the forward propagation function, x i (t) is the i-th envelope modulated oscillation wave (i.e. x i2 (t)), x noise (t) is the envelope modulated noise wave;
[0087] Perform time spectrum calculation on the forward propagation function, that is, uniformly sample the forward propagation function, add a Hanning window, and perform fast Fourier transform to obtain the target time spectrum matrix; for example, take the total length of S(t) as 44100 (44100 sampling rate is discussed, 1 second), take a data of length 512 every 128 samples, add a Hanning window and perform FFT calculation to obtain the target time spectrum matrix;
[0088] Get the real drum sound and perform spectrogram calculation to get the real spectrogram matrix;
[0089] The oscillation wave parameters and noise wave parameters are input into the neural network for training, and the similarity between the target time spectrum matrix and the real time spectrum matrix is used as the cost function: J = ∑(F(t,f)-F target (t,f)) 2
[0090] Where J is the cost function, F(t,f) is the real time spectrum matrix (t is discrete and is sampled every 2.9ms (128 / 44100 seconds), which means that after point t, a 512-length signal is taken, the Hanning window is added, the FFT is taken, and the absolute value is taken). target (t,f) is the target time spectrum matrix;
[0091] When the cost function is determined to meet preset conditions, the neural network training stops and the optimal oscillation wave parameters and optimal noise wave parameters are output to obtain the optimal oscillation wave and optimal noise wave. The preset conditions include: the number of training times reaches 300,000 or the cost function value is less than 0.01%;
[0092] The optimal oscillation wave and the optimal noise wave are input into the filter to control the timbre, then input into the amplifier to control the volume, and finally the mixed signal is output.
[0093] In this step, based on the envelope's modulation of the oscillation wave and noise wave, a parameter combination is obtained: 5m + 3n dynamic parameters, where 5 represents five dynamic parameters for each oscillation wave, m represents the number of oscillation waves, 3 represents three dynamic parameters for each noise band within the noise wave, and n represents the number of noise bands within the noise wave. The dynamic parameters of the oscillation wave and the dynamic parameters of the noise band within the noise wave are independent of each other. Using these parameters as a vector, a forward propagation function is designed and time-spectrogram calculations are performed to obtain two time-spectrogram matrices. Using existing neural networks (such as TensorFlow and PyTorch), the forward propagation process is reproduced through custom nodes. Backward propagation is then automatically performed, and training iterations and continuous updates are performed to obtain parameter combinations with high similarity.
[0094] In this embodiment, as shown in FIG2 , first, a drum plate vibration signal is collected, and the volume of the sound received by the drum plate is obtained according to the vibration signal, and an oscillation wave and a noise wave are emitted.
[0095] Then, based on the volume of the drum pad, an envelope with controlled attenuation is generated. The envelope generation formula is: V(t) = max(E(t), V(t-dt)*a)
[0096] Where V(t) is the envelope of the oscillation wave at time t; E(t) is the drum signal energy received at time t; dt is a fixed short time, such as 1ms; V(t-dt) is the envelope of the oscillation wave at time t-dt; and a is the envelope attenuation coefficient of the oscillation wave, which controls the speed of envelope attenuation and is specifically a number very close to 1 but less than 1.
[0097] Finally, multiple envelopes with different attenuation coefficients are used to simultaneously control all oscillation waves, the noise wave, the filter after the mixture of all oscillation waves and noise, and the final amplifier, and then mix and output the sound. Specifically, the envelope modulates the amplitude modulation depth of the oscillation wave to control the volume and, therefore, the timbre; the frequency modulation depth of the oscillation wave to control the pitch; the noise band floor volume of the noise wave to control the timbre; the envelope modulates the filter to further control the timbre; and the envelope modulates the amplifier to further control both the volume and timbre. In this way, as the drum is struck, the pitch, timbre, and volume of the output drum sound can be controlled to change in real time with the strike signal (vibration signal), and can be used to simulate the realistic timbre of different drum pieces.
[0098] It should be noted that modulation refers to the use of envelopes to control the various parameters of the tone in real time, and modulation depth refers to the degree of influence on the parameters when using envelopes to control the parameters. The greater the modulation depth, the greater the influence.
[0099] The above-described drum audio processing method offers the advantages of realistic dynamics, sound quality independent of storage, and low latency. First, traditional triggering requires quantifying the force of the strike and converting it into a specific wavetable for playback. In this application, however, physical modeling is performed based on the collected vibration signal. Different recorded signals result in different output signals, resulting in subtle differences in the timbre of each strike. This is especially true when using unique techniques such as brushing and rolls, which are faithfully reflected in the final output timbre. Even differences in drumsticks and the tightness of the mesh can affect the sound quality. This approach closely resembles real drums and is unattainable with traditional electronic drums. Second, the sound quality of traditional wavetables directly relies on extensive storage. For example, different strike positions and forces correspond to different wavetables. This application, however, directly utilizes sound effects, ensuring high precision in all processing and achieving lossless drum sound quality without relying on storage. Finally, since drums are rhythm instruments, latency requirements are particularly stringent. This application does not have any trigger judgment and does not need to play the wavetable. Instead, it converts the beating audio signal into drum sound in real time, thereby achieving ultra-low latency. This ultra-low latency will give the drummer a very good improved experience.
[0100] It should be understood that, although the various steps in the flowchart of FIG1 are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in FIG1 may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these sub-steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0101] The present application also provides a drum audio processing device, as shown in FIG3 . In one embodiment, the device includes: an acquisition module 302 , an oscillation wave modulation module 304 , a noise wave modulation module 306 , and an output module 308 , wherein:
[0102] an acquisition module 302 for acquiring an audio signal of a drum and generating an oscillation wave and a noise wave according to the audio signal;
[0103] an oscillation wave modulation module 304 for modulating the oscillation wave to obtain oscillation wave parameters for controlling pitch and volume;
[0104] The noise wave modulation module 306 is used to modulate the noise wave to obtain noise wave parameters for controlling the timbre;
[0105] The output module 308 is used to perform real-time sound effect processing and mixed sound output according to the oscillation wave parameters and the noise wave parameters.
[0106] The specific definition of a drum audio processing device can be found in the definition of a drum audio processing method described above and will not be further elaborated here. Each module in the above-described device may be implemented in whole or in part via software, hardware, or a combination thereof. Each of these modules may be embedded in or independent of a processor in a computer device in hardware form, or may be stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.
[0107] The present application also provides an electric drum, the structural design of which refers to a real drum rather than an electric drum. As shown in Figure 4, in one embodiment, it includes: a drum plate, a sound receiver, a collector and a processor. The drum plate, the sound receiver, the collector and the processor are connected in sequence, that is, the drum plate is connected to one end of the sound receiver, the other end of the sound receiver is connected to one end of the collector, and the other end of the collector is connected to the processor.
[0108] Drum pad: Generates an audio signal based on the striking sound (strike signal) and sends it to the receiver via multiple audio cables (such as a stereo audio cable). Considering the crosstalk between tones, the drum skin on the drum pad uses a silent mesh design to achieve silence.
[0109] Receiver: Collects audio signals and sends them to the collector via multiple audio cables. Receivers (e.g., buzzers or silicon microphones) are located at iconic locations on the drum pad (e.g., the center and edge of the snare drum pad). Multiple receivers can be used to obtain multiple audio signals and can be connected to the collector via the audio input port on the main control device.
[0110] Collector: converts the audio signal into an audio digital signal (IIS signal) and sends it to the processor; there can be multiple collectors (for example: Codec, AD array, multi-channel AD converter, FPGA) to convert multiple audio signals into multiple audio digital signals.
[0111] Processor: Use the drum audio processing method to perform real-time sound effect processing on the audio digital signal and output a mixed signal (i.e., convert the audio signal of the drum plate into a timbre similar to that of the original acoustic drum, that is, perform timbre modeling, for example: snare drum timbre modeling, bass drum timbre modeling, to become the timbre of the snare drum, bass drum, and cymbals); the processor can be a DSP unit, which performs real-time sound effect processing on each audio digital signal, converts the original striking sound of the drum plate mesh into the corresponding timbre of each drum plate in the real drum, and then mixes these sounds into audio and outputs them through the stereo audio output port (i.e., outputs the mixed signal).
[0112] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0113] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A method for audio processing of a drum, characterized in that, Including: Obtain the digital audio signal of the drum, perform real-time sound effect processing on the digital audio signal, and perform mixing output.
2. The method for audio processing of a drum according to claim 1, wherein Obtain the digital audio signal of the drum, perform real-time sound effect processing on the digital audio signal, and perform mixing output, including: Obtain the digital audio signal of the drum, and generate an oscillation wave and a noise wave according to the digital audio signal; Modulate the oscillation wave to obtain oscillation wave parameters for controlling pitch and volume; Modulate the noise wave to obtain noise wave parameters for controlling timbre; Perform real-time sound effect processing and mixing output according to the oscillation wave parameters and the noise wave parameters.
3. The method for audio processing of a drum according to claim 2, characterized in that, Modulate the oscillation wave to obtain oscillation wave parameters for controlling pitch and volume, including: There are multiple oscillation waves, and different envelopes are used to modulate each oscillation wave respectively to obtain multiple envelope-modulated oscillation waves. The envelope-modulated oscillation waves include: oscillation wave parameters for controlling pitch and volume; x i1 x(t) = A i sin(2πft i t) V i V(t) = max(E(t), V i (t - dt) * a i ) x i2 (t) = (A i + d i1 V i (t)) sin(2π(f i + d i2 V i (t)) t) where x i1 (t) is the i-th oscillatory wave, A i is the steady-state amplitude of the i-th oscillatory wave, f i is the steady-state frequency of the i-th oscillatory wave, V i (t) is the envelope of the i-th oscillatory wave at time point t, E(t) is the energy of the drum signal received at time point t, V i (t - dt) is the envelope of the i-th oscillatory wave at time point t - dt, dt is a fixed short time, a i is the envelope attenuation coefficient of the i-th oscillatory wave, x i2 (t) is the i-th envelope-modulated oscillatory wave, d i1 is the amplitude modulation depth of the i-th envelope-modulated oscillatory wave, d i2 is the frequency modulation depth of the i-th envelope-modulated oscillatory wave; The oscillation wave parameters include: steady-state amplitude, steady-state frequency, amplitude modulation depth, frequency modulation depth, and envelope attenuation coefficient.
4. The method for audio processing of a drum according to claim 3, characterized in that, Modulate the noise wave to obtain noise wave parameters for controlling timbre, including: Divide the noise wave into multiple noise frequency bands, and use different envelopes to modulate each noise frequency band respectively to obtain multiple envelope-modulated noise waves. The envelope-modulated noise waves include: noise wave parameters for controlling timbre; V noise,j V(t) = max(E(t), V noise,j (t - dt) * a noisej ) Where, V noise,j (t) is the j-th noise frequency band, V noise,j (t - dt) is the envelope of the j-th noise frequency band at time point t - dt, dt is a fixed short time, a noisej is the envelope attenuation coefficient of the j-th noise frequency band, x noise,j (t) is the envelope modulation noise wave of the j-th noise frequency band, A noise,j is the base volume of the j-th noise frequency band, d noise,j is the modulation depth of the j-th noise frequency band, V noise,j (t) is the envelope of the j-th noise frequency band at time point t, H j is the (j - 1)-th frequency division filter, rand(t) is a random number generator; The noise wave parameters include: noise frequency band base volume, noise frequency band modulation depth, and noise frequency band envelope attenuation coefficient.
5. The method for audio processing of a drum according to any one of claims 2 to 4, characterized in that Perform real-time sound effect processing and mixing output according to the oscillation wave parameters and the noise wave parameters, including: Design a forward propagation function according to the oscillation wave parameters and the noise wave parameters; perform spectrogram calculation on the forward propagation function to obtain a target spectrogram matrix; Obtain the real drum sound, and perform spectrogram calculation to obtain a real spectrogram matrix; Input the oscillation wave parameters and the noise wave parameters into a neural network for training, and use the similarity between the target spectrogram matrix and the real spectrogram matrix as a cost function. When it is judged that the cost function meets the preset conditions, the neural network training stops, and the optimal oscillation wave parameters and the optimal noise wave parameters are output to obtain the optimal oscillation wave and the optimal noise wave; Input the optimal oscillation wave and the optimal noise wave into a filter to control timbre, then input them into an amplifier to control volume, and finally output a mixed signal.
6. The method for audio processing of a drum according to claim 5, characterized in that, Design a forward propagation function according to the oscillation wave parameters and the noise wave parameters, including: where \(S(t)\) is the forward propagation function, \(x i (t)\) is the \(i\)-th envelope-modulated oscillation wave, \(x noise (t)\) is the envelope-modulated noise wave.
7. The method for audio processing of a drum according to claim 5, wherein Perform spectrogram calculation on the forward propagation function to obtain a target spectrogram matrix, including: Perform uniform sampling on the forward propagation function, add a Hanning window, and perform a fast Fourier transform to obtain a target spectrogram matrix.
8. The method for audio processing of a drum according to claim 5, characterized in that Use the similarity between the target spectrogram matrix and the real spectrogram matrix as a cost function, including: J = ∑(F(t, f) - F target (t, f)) 2 Where J is the cost function, F(t, f) is the true time-frequency spectrum matrix, and F target (t, f) is the target time-frequency spectrum matrix.
9. The method for audio processing of a drum according to claim 5, characterized in that, The preset conditions include: The number of training times reaches 300,000 times or the cost function value is less than 0.01%.
10. An apparatus for audio processing of a drum, characterized in that, Including: An acquisition module, configured to acquire the audio signal of the drum, and generate an oscillation wave and a noise wave according to the audio signal; An oscillation wave modulation module, configured to modulate the oscillation wave to obtain oscillation wave parameters for controlling pitch and volume; A noise wave modulation module, configured to modulate the noise wave to obtain noise wave parameters for controlling timbre; An output module, configured to perform real-time audio effect processing according to the oscillation wave parameters and the noise wave parameters, and perform mixed audio output.
11. Electronic drum, characterized in that, Comprising: A drum disc, a microphone, a collector and a processor; The drum disc generates an audio signal according to the percussion sound; The microphone collects the audio signal and sends it to the collector; The collector converts the audio signal into an audio digital signal and sends it to the processor; The processor performs real-time audio effect processing by using the method for audio processing of the drum according to any one of claims 1 to 9, and outputs a mixed audio signal.
Citation Information
Patent Citations
Editable general tone synthesis analysis system and method
CN112289289A
Musical instrument timbre modeling method and device, sound module and storage medium
CN115331649A
Tone generation method and system
CN116758882A
Drum audio processing method and device and electro-acoustic drum
CN117765904A
Machine-Learned Differentiable Digital Signal Processing
US20220013132A1