Drum tone simulation model training method, simulation method, device, equipment and medium
Patent Information
- Application Number
- CN202511807429.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2045-12-03
AI Technical Summary
尽管延迟时间较短,但对于追求极致实时响应速度的专业演奏者而言,仍可能影响演奏的流畅性与即时反馈感
[0018]相比现有技术,本发明的技术效果,如下:
Smart Images

Figure CN121281476B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of audio signal processing technology, and in particular to a training method, simulation method, device, equipment and medium for a drum timbre simulation model. Background Technology
[0002] Electronic drums belong to the field of electronic musical instrument technology, specifically involving a drum kit simulation device based on sensor detection and digital audio technology, which aims to provide performers with a playing experience similar to that of traditional acoustic drum kits.
[0003] A typical electronic drum system usually includes several trigger pads, a hi-hat controller, and a core sound module. Its working principle is as follows: the player's striking motion on the trigger pads is captured by the built-in sensors and converted into electrical signals. These signals are analyzed by signal processing circuitry to identify characteristics such as the force and location of the strikes, and then a trigger signal is generated and sent to the sound module. The sound module ultimately uses pre-stored digital audio data or synthesizes it through algorithms to output the corresponding drum sound audio.
[0004] In terms of signal detection, most mainstream electronic drum pads currently use piezoelectric sensors as the core detection element. These sensors utilize the piezoelectric effect, generating a voltage signal proportional to the stress experienced upon mechanical impact. Depending on their functional complexity, drum pads can be categorized as single-zone, dual-zone, or multi-zone. Single-zone drum pads typically integrate only one piezoelectric element to sense the impact force; while dual-zone or triple-zone drum pads differentiate between center and edge strikes by placing additional piezoelectric elements or switching elements at specific locations (such as the edges). The raw signal from the sensor undergoes buffering, amplification, filtering, and analog-to-digital conversion. The built-in microprocessor then executes a series of algorithms, such as peak detection, threshold judgment, jitter reduction, and time window analysis, ultimately determining a valid strike event and quantifying it into a standard force value.
[0005] In terms of sound generation, existing electronic drum sound modules primarily employ a tone engine based on multi-layer sampling. This engine pre-stores multiple sample layers of the same drum tone recorded at different dynamic levels. The system calls upon the corresponding sample layer for playback based on the real-time detected dynamic values, and simulates dynamic changes in tone through crossfading or direct switching. To pursue greater realism, some high-end products have begun to introduce physical modeling technology, using mathematical algorithms to simulate the vibration patterns of the drumhead, cavity resonance, and the energy attenuation of the cymbals, in order to generate a more continuously dynamic tone.
[0006] Despite significant advancements in existing technology, there are still several notable shortcomings: The accuracy of effective strike identification is unsatisfactory: Current methods relying on peak detection and fixed threshold judgment have inherent limitations. In fast, dense performances, signal peaks may overlap or be too close, leading to omissions or misjudgments by the detection algorithm. Furthermore, the fixed threshold mechanism struggles to effectively identify strikes with weak force but clear musical intent, lacking a nuanced perception of the performer's dynamic intentions.
[0007] The system trigger latency needs further optimization: To ensure accurate peak capture, traditional peak detection algorithms need to wait a preset time window to confirm that the signal has reached its peak. This process inevitably introduces additional system latency. Although the latency is short, it may still affect the smoothness of the performance and the immediate feedback for professional performers who pursue the ultimate real-time response speed.
[0008] The flexibility and realism limitations of sound generation methods: Multi-sampling-based schemes are essentially playbacks of discrete samples. Switching between different sampling layers is difficult to achieve completely smooth and natural transitions, and they sound abrupt when simulating subtle dynamics such as continuous changes in drumstick contact position and force. While existing physical modeling algorithms provide continuity, their model parameters are often fixed or have limited adjustable ranges, making it difficult to faithfully reproduce the complex acoustic characteristics of the drum body under different striking conditions (such as force and position). This limits the richness and controllability of the sound.
[0009] Therefore, there is an urgent need in this field for a new electronic drum technology solution that can more accurately identify playing intentions, respond more quickly, and generate more dynamic details and high-fidelity timbres. Summary of the Invention
[0010] In view of the shortcomings of the existing technology, the purpose of this invention is to provide a training method, simulation method, device, equipment and medium for a drum timbre simulation model.
[0011] To achieve the above objectives, the technical solution adopted by the present invention is as follows: On one hand, the present invention provides a method for training a drum timbre simulation model, comprising the following steps: Multiple sets of training samples were obtained, each set of training samples including drumbeat signals and corresponding real drum sound recordings; A deep learning synthesis model is constructed. The input of the deep learning synthesis model is a drumbeat signal, and the output is a synthesized drum sound. The deep learning synthesis model represents the output synthesized drum sound as a superposition of an impulse response part, a harmonic sine part that varies with the envelope, and a noise part that varies with the envelope. A loss function is constructed with the goal of minimizing the difference between the synthesized drum sound and the real drum sound recording. A deep learning synthesis model is trained using training samples to learn the mapping relationship from the drum strike signal to the parameter set of the deep learning synthesis model, so as to obtain a well-trained drum timbre simulation model.
[0012] Furthermore, the drumbeat signal is the pickup output signal when the drum is struck.
[0013] On the other hand, the present invention provides a method for simulating drum timbre, comprising the following steps: Acquire the target drumbeat signal; The target drumbeat signal is input into the drum timbre simulation model obtained by the above-mentioned drum timbre simulation model training method to generate and output a synthesized drum sound.
[0014] On the other hand, the present invention provides a drum tone simulation device, comprising: The training sample acquisition module is used to acquire multiple sets of training samples, each set of training samples including drum beat signals and corresponding real drum sound recordings; The model building module is used to build a deep learning synthesis model. The input of the deep learning synthesis model is a drumbeat signal, and the output is a synthesized drum sound. The deep learning synthesis model represents the output synthesized drum sound as a superposition of an impulse response part, a harmonic sine part that varies with the envelope, and a noise part that varies with the envelope. The model training module is used to construct a loss function with the goal of minimizing the difference between the synthesized drum sound and the real drum sound recording. It uses training samples to train a deep learning synthesis model and learns the mapping relationship from the drum strike signal to the parameter set of the deep learning synthesis model in order to obtain a trained drum timbre simulation model. The drum timbre simulation module inputs the target drum impact signal into the trained drum timbre simulation model to generate and output a synthesized drum sound.
[0015] On the other hand, the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of a drum timbre simulation model training method, or to implement the steps of a drum timbre simulation method.
[0016] On the other hand, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of a drum timbre simulation model training method, or the steps of a drum timbre simulation method.
[0017] On the other hand, the present invention provides a computer program product stored on a computer-readable storage medium and including computer instructions that, when executed by a processor, cause a computer device to implement a step of a drum timbre simulation model training method, or a step of implementing a drum timbre simulation method.
[0018] Compared with the prior art, the technical effects of the present invention are as follows: This invention constructs a deep learning synthesis model. The input to this model is a drumbeat signal, and the output is a synthesized drum sound. The deep learning synthesis model represents the output synthesized drum sound as a superposition of an impulse response component, a harmonic sine wave component varying with the envelope, and a noise component varying with the envelope. The deep learning synthesis model optimizes the parameters of each component to achieve high-fidelity reproduction and controllable synthesis of the drum sound characteristics. The synthesized drum sound is highly consistent with the real drum sound in both subjective listening experience and objective spectrum. Compared with traditional sampling-based schemes, this invention provides accurate identification of effective impact events, lower envelope detection latency compared to peak detection, and better dynamic response and physical interpretability. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.
[0020] Figure 1 This is a flowchart of a drum timbre simulation model training method in one embodiment. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0022] Reference Figure 1 One embodiment proposes a method for training a drum timbre simulation model, comprising the following steps: Multiple sets of training samples were obtained, each set of training samples including drumbeat signals and corresponding real drum sound recordings; A deep learning synthesis model is constructed. The input of the deep learning synthesis model is a drumbeat signal, and the output is a synthesized drum sound. The deep learning synthesis model represents the output synthesized drum sound as a superposition of an impulse response part, a harmonic sine part that varies with the envelope, and a noise part that varies with the envelope. A loss function is constructed with the goal of minimizing the difference between the synthesized drum sound and the real drum sound recording. A deep learning synthesis model is trained using training samples to learn the mapping relationship from the drum strike signal to the parameter set of the deep learning synthesis model, so as to obtain a well-trained drum timbre simulation model.
[0023] The drum strike signal is the pickup output signal when the drum is struck. The corresponding real drum sound recording is a real drum sound recording made using the same playing method and striking position.
[0024] This invention uses a deep learning synthesis model to characterize the output synthesized drum sound as a superposition of an impulse response component, an envelope-varying harmonic sine wave component, and an envelope-varying noise component. The impulse response component represents the resonance transfer function of the drum body and drumhead after being subjected to an instantaneous impact, corresponding to the spatial response characteristics formed by the drum cavity structure and the drum shell material. The envelope-varying harmonic sine wave component simulates the vibration mode of the drumhead and cavity, with different harmonics corresponding to different orders of drumhead resonance frequencies, and the amplitude and frequency can vary with the striking force and striking position. The noise component simulates the high-frequency transients and air disturbances at the moment of impact, as well as the snare drum's string structure.
[0025] Specifically, the deep learning synthesis model represents the output synthesized drum sound as follows: ; in express The impulse response function of the drum body at any given moment reflects the resonance characteristics after the drum cavity and drumhead are coupled; express Constantly impacting the energy envelope signal; express The envelope signal of the harmonic components is controlled at all times; express The envelope signal of the noise component is controlled at all times; This represents the amplitude of the k-th harmonic component, where k = 1, 2, ..., N, and N is the number of harmonic components. This represents the frequency of the k-th harmonic component; This represents the phase of the k-th harmonic component; Indicates the noise amplitude. express A white noise sequence with unit variance at time step.
[0026] For three envelope signals , , Each envelope signal All are corresponding Parameter control, among which Defined as: ; in for The moment of the drumbeat signal, that is, the pickup output signal when the drum is struck.
[0027] To achieve adaptive modeling of drum sound features, this invention introduces a deep learning synthesis model. The deep learning synthetic model is a neural network model. Its input is the drumbeat signal, and its output is the set of parameters from the aforementioned model. ; in, Noise amplitude, These are the parameters of the neural network model.
[0028] The loss function is constructed with the objective of minimizing the difference between the synthesized drum sound and the real drum sound recording, as follows: ; in It is the output of the deep learning synthesis model. Synthetic drum sounds at all times yes Real drum sound recordings at all times It is a Mel-spectrum transform. These are weighting coefficients, where the first term on the right-hand side of the equation constrains the time-domain characteristics to be consistent, and the second term constrains the frequency-domain characteristics to be consistent.
[0029] The above embodiments combine interpretable physical acoustic models with data-driven deep learning. The model can automatically learn the complex, nonlinear mapping from the drumbeat signal to the underlying physical model parameters (harmonic frequencies, amplitudes, phases, noise envelopes, noise amplitudes, etc.). This makes the synthesized drum sound possess the realism of the physical model, surpassing traditional solutions in both timbre fidelity and performance controllability.
[0030] In one embodiment, a method for simulating drum timbre is provided, comprising the following steps: Acquire the target drumbeat signal; The target drumbeat signal is input into the trained drum timbre simulation model to generate and output a synthesized drum sound. The trained drum timbre simulation model is obtained based on the drum timbre simulation model training method of any of the foregoing embodiments.
[0031] In another embodiment, a drum tone simulation device is provided, comprising: The training sample acquisition module is used to acquire multiple sets of training samples, each set of training samples including drum beat signals and corresponding real drum sound recordings; The model building module is used to build a deep learning synthesis model. The input of the deep learning synthesis model is a drumbeat signal, and the output is a synthesized drum sound. The deep learning synthesis model represents the output synthesized drum sound as a superposition of an impulse response part, a harmonic sine part that varies with the envelope, and a noise part that varies with the envelope. The model training module is used to construct a loss function with the goal of minimizing the difference between the synthesized drum sound and the real drum sound recording. It uses training samples to train a deep learning synthesis model and learns the mapping relationship from the drum strike signal to the parameter set of the deep learning synthesis model in order to obtain a trained drum timbre simulation model. The drum timbre simulation module inputs the target drum impact signal into the trained drum timbre simulation model to generate and output a synthesized drum sound.
[0032] On the other hand, the present invention provides a computer device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the drum timbre simulation model training method provided in any of the above embodiments, or to implement the steps of the drum timbre simulation method provided in any of the above embodiments. The computer device may be a server. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device stores sample data. The network interface of the computer device is used for communication with external terminals via a network connection.
[0033] On the other hand, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it implements the steps of the drum timbre simulation model training method provided in any of the above embodiments, or implements the steps of the drum timbre simulation method provided in any of the above embodiments.
[0034] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0035] Matters not covered in this invention are common knowledge.
[0036] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0037] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these modifications and improvements all fall within the protection scope of this application.
[0038] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A training method for a drum timbre simulation model, characterized in that, Includes the following steps: Multiple sets of training samples were obtained, each set of training samples including drum beat signals and corresponding real drum sound recordings; A deep learning synthesis model is constructed. The input to the deep learning synthesis model is a drum impact signal, and the output is a synthesized drum sound. The deep learning synthesis model represents the output synthesized drum sound as a superposition of an impulse response component, an envelope-varying harmonic sine wave component, and an envelope-varying noise component. The impulse response component represents the resonance transfer function of the drum body and drumhead after an instantaneous impact, corresponding to the spatial response characteristics formed by the drum cavity structure and drum shell material. The envelope-varying harmonic sine wave component simulates the vibration mode of the drumhead and cavity; different harmonics correspond to different orders of drumhead resonance frequencies, and the amplitude and frequency can vary with the impact force and impact position. The noise component simulates the high-frequency transients and air disturbances at the moment of impact, as well as the snare drum's string structure. The deep learning synthesis model is a neural network model, and the deep learning synthesis model represents the output synthesized drum sound as follows: in express The impulse response function of the drum body at any given moment reflects the resonance characteristics after the drum cavity and drumhead are coupled; express Constantly impacting the energy envelope signal; express The envelope signal of the harmonic components is controlled at all times; express The envelope signal of the noise component is controlled at all times; This represents the amplitude of the k-th harmonic component, where k = 1, 2, ..., N, and N is the number of harmonic components. This represents the frequency of the k-th harmonic component; This represents the phase of the k-th harmonic component; Indicates the noise amplitude. express A white noise sequence with unit variance at time step; A loss function is constructed with the goal of minimizing the difference between the synthesized drum sound and the real drum sound recording. A deep learning synthesis model is trained using training samples to learn the mapping relationship from the drum strike signal to the parameter set of the deep learning synthesis model, so as to obtain a well-trained drum timbre simulation model.
2. The drum timbre simulation model training method according to claim 1, characterized in that, The drumbeat signal is the output signal of the pickup when the drum is struck.
3. The drum timbre simulation model training method according to claim 1 or 2, characterized in that, For three envelope signals , , Each envelope signal All are corresponding Parameter control, among which Defined as: in for The drumbeat signal at the right moment.
4. The drum timbre simulation model training method according to claim 1 or 2, characterized in that, The loss function is constructed with the objective of minimizing the difference between the synthesized drum sound and the real drum sound recording, as follows: in It is the output of the deep learning synthesis model. Synthetic drum sounds at all times yes Real drum sound recordings at all times It is a Mel-spectrum transform. It is the weighting coefficient.
5. A method for simulating drum timbre, characterized in that, Includes the following steps: Acquire the target drumbeat signal; The target drumbeat signal is input into the drum timbre simulation model obtained by the drum timbre simulation model training method of claim 1, and a synthesized drum sound is generated and output.
6. A drum tone simulation device, used to implement the drum tone simulation method of claim 5, characterized in that, include: The training sample acquisition module is used to acquire multiple sets of training samples, each set of training samples including drum beat signals and corresponding real drum sound recordings; The model building module is used to build a deep learning synthesis model. The input of the deep learning synthesis model is a drumbeat signal, and the output is a synthesized drum sound. The deep learning synthesis model represents the output synthesized drum sound as a superposition of an impulse response part, a harmonic sine part that varies with the envelope, and a noise part that varies with the envelope. The model training module is used to construct a loss function with the goal of minimizing the difference between the synthesized drum sound and the real drum sound recording. It uses training samples to train a deep learning synthesis model and learns the mapping relationship from the drum strike signal to the parameter set of the deep learning synthesis model in order to obtain a trained drum timbre simulation model. The drum timbre simulation module inputs the target drum impact signal into the trained drum timbre simulation model to generate and output a synthesized drum sound.
7. An electronic device comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the drum timbre simulation model training method as described in claim 1, or the steps of the drum timbre simulation method as described in claim 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the drum timbre simulation model training method as described in claim 1, or the steps of the drum timbre simulation method as described in claim 5.
Citation Information
Patent Citations
Miniature wave table phonics method, system and electronic musical instrument
CN106356047A
Drum audio processing method and device and electro-acoustic drum
CN117765904A