Midi-based human voice timbre driving method, device, medium and system

CN122676784APending Publication Date: 2026-09-01GUANGZHOU ENYA INNOVATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610881319.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-17
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

但是,在电子乐器领域,通过电子乐器来实现人声演奏的技术还无法实现

Benefits of technology

本发明能够将人声音色演奏从专业录音棚扩展到实时现场演奏和大众音乐创作的场景中,提高用户演奏体验;以及以标准的MIDI协议为基础,兼容现有的所有MIDI控制器设备,无需借助专用硬件即可实现,可广泛应用于智能乐器、DAW编曲、Live演出及音乐教育等领域。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122676784A_ABST
    Figure CN122676784A_ABST
Patent Text Reader

Abstract

The application discloses a human voice timbre driving method, device, medium and system based on MIDI. The method comprises the following steps: recording human voice audio according to a pronunciation element and pitch information to obtain a plurality of audio sampling samples; each audio sampling sample comprises an audio segment, the pronunciation element and the pitch information; configuring a trigger parameter for each audio sampling sample according to a preset rule to construct a mapping relationship between the trigger parameter and the audio sampling sample; obtaining a MIDI parameter through a MIDI controller and analyzing the MIDI parameter to obtain the pitch information and the trigger parameter; and playing the audio segment of the corresponding audio sampling sample according to the trigger parameter and the pitch information of the corresponding audio sampling sample. The application can realize real-time performance of human voice.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic music devices, and more particularly to a method, apparatus, medium, and system for driving human voice timbre based on MIDI. Background Technology

[0002] In traditional music production and live performances, vocals are typically sung live by the performer or pre-recorded in a professional recording studio and then played back. However, in the field of electronic musical instruments, the technology to perform vocals using electronic instruments is not yet feasible. Summary of the Invention

[0003] In order to overcome the shortcomings of the prior art, one of the objectives of this invention is to provide a MIDI-based human voice timbre driving method that enables real-time performance of human voices.

[0004] The second objective of this invention is to provide a MIDI-based human voice timbre driving device that enables real-time performance of human voices.

[0005] A third objective of this invention is to provide a computer-readable medium that enables real-time performance of human voices.

[0006] The fourth objective of this invention is to provide a MIDI-based human voice timbre driving system that enables real-time performance of human voices.

[0007] One of the objectives of this invention is achieved through the following technical solution: A MIDI-based human voice timbre driving method includes: Audio source package acquisition steps: Record human voice audio based on pronunciation elements and pitch information to obtain multiple audio sample samples; each audio sample sample includes an audio segment, pronunciation elements, and pitch information; Configuration steps: Configure trigger parameters for each audio sample according to preset rules to establish a mapping relationship between trigger parameters and audio sample samples; Triggering steps: Obtain MIDI parameters through the MIDI controller and parse the MIDI parameters to obtain pitch information and trigger parameters; Matching steps: Based on the trigger parameters and mapping relationship, several candidate audio sample samples are obtained. Based on the pitch information and several candidate audio sample samples, the corresponding audio sample sample is obtained. Then, the audio segment of the corresponding audio sample sample is played.

[0008] Furthermore, the configuration steps specifically include: determining the maximum and minimum values ​​of each velocity layer based on the velocity range, and simultaneously constructing a mapping relationship between each velocity layer and each audio sample component; one velocity layer is matched with several audio sample samples.

[0009] Furthermore, the triggering step also includes: obtaining CC control parameters and obtaining corresponding modification parameters according to the CC control parameters and the CC control parameter configuration table, processing the corresponding audio segment according to the corresponding modification parameters to generate a human voice audio signal, and transmitting the human voice audio signal to the subsequent audio playback device for playback through the output interface.

[0010] Furthermore, the specific steps for creating an audio source package include: Steps for determining pronunciation elements: Determine pronunciation elements based on requirements; the set of pronunciation elements should include at least standard vowels and commonly used filler sounds; standard vowels include A, E, I, O, and U; commonly used filler sounds include murmurs, auspicious sounds, velar sounds, and auspicious-velar sounds; Pitch configuration steps: Determine the target pitch range of the human voice source, divide the target pitch range into multiple pitches according to the preset segmentation rules, and combine each pitch with each articulation element in the articulation element set to form multiple pitch and articulation element combinations; Recording steps: Sample the audio segment corresponding to each pitch and vocal element combination according to the preset sampling rules; the preset sampling rules include the duration of the sampling segment, the acquisition target, the recording environment, the signal-to-noise ratio, and the sampling rate.

[0011] Furthermore, the recording steps also include: recording human voice audio based on articulation elements and pitch information to obtain multiple audio sample samples, specifically including: setting multiple velocity layers and configuration sampling rules for each velocity layer, and then recording human voice audio based on each pitch according to the configuration sampling rules of each velocity layer to obtain multiple audio sample samples; each audio sample sample includes pitch, velocity layer, articulation elements, and audio segments.

[0012] Furthermore, multiple intensity levels include a light intensity level, a medium intensity level, and a strong intensity level. Each intensity level includes a maximum intensity value and a minimum intensity value. When the intensity level is a light intensity level, the sampling rule is configured to use breathy and soft singing methods for sampling. When the intensity level is a medium intensity level, the sampling rule is configured to use standard natural singing methods for sampling. When the intensity level is a strong intensity level, the sampling rule is configured to use loud, shouting, or head voice methods for sampling.

[0013] The second objective of this invention is achieved by the following technical solution: A MIDI-based human voice timbre driving device includes a memory and a processor. The memory stores a human voice timbre driver program that runs on the processor. The human voice timbre driver program is a computer program. When the processor executes the human voice timbre driver program, it implements the steps of a MIDI-based human voice timbre driving method as one of the objectives of this invention.

[0014] The third objective of this invention is achieved by the following technical solution: A computer-readable storage medium storing a human voice timbre driver program thereon, the human voice timbre driver program being a computer program, the human voice timbre driver program being executed by a processor to implement the steps of a MIDI-based human voice timbre driver method as one of the objectives of this invention.

[0015] The fourth objective of this invention is achieved by the following technical solution: MIDI-based human voice timbre driving system, including: MIDI controller; The storage module is used to store the human voice source package and mapping configuration file; A processing module, electrically connected to a MIDI controller, is used to acquire MIDI parameters from the MIDI controller and derive audio signals using a MIDI-based human voice timbre driving method employed according to one of the purposes of this invention. The audio output module is used to input audio signals into subsequent audio playback devices for playback.

[0016] Furthermore, the MIDI controller is at least one of the following: MIDI guitar, MIDI keyboard, MIDI drum pad, MIDI wind instrument, MIDI glove, and motion controller.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention extends vocal performance from professional recording studios to real-time live performances and popular music creation scenarios, improving the user's performance experience; and based on the standard MIDI protocol, it is compatible with all existing MIDI controller devices, can be implemented without the need for dedicated hardware, and can be widely used in fields such as smart musical instruments, DAW arrangement, live performances and music education. Attached Figure Description

[0018] Figure 1 The flowchart of the MIDI-based vocal production control method provided by the present invention. Detailed Implementation

[0019] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments. It should be noted that, without conflict, the various embodiments or technical features described below can be arbitrarily combined to form new embodiments. Example 1

[0020] This invention constructs a human voice timbre driving system based on the MIDI protocol, allowing users to perform real-time human voices using any standard controller that supports the MIDI protocol, such as a MIDI keyboard, MIDI guitar, MIDI drum pad, or MIDI wind instrument. This achieves a musical expressiveness that closely resembles that of a real human voice performance, thereby enhancing the user experience.

[0021] like Figure 1 As shown, the present invention provides a preferred embodiment of a MIDI-based method for controlling the vocal production of a performer, comprising: Step S1: Record human voice audio based on pronunciation elements and pitch information to obtain multiple audio sample samples; each audio sample sample includes an audio segment, pronunciation elements, and pitch information.

[0022] This invention pre-records human voices to form human voice audio segments, constructs a human voice source package, and stores it in the system in advance. During performance, the corresponding audio segment is retrieved by triggering parameters to achieve human voice playback.

[0023] Specifically, the articulation elements generally include standard vowels and commonly used filler sounds; among them, standard vowels include A, E, I, O, and U, and commonly used filler sounds include murmurs, auspicious sounds, whimpers, and auspicious-whimper mixtures. By collecting a corresponding audio sample for each articulation element, when a corresponding sound needs to be played, the corresponding audio sample can be found directly through the corresponding sound to achieve audio playback.

[0024] Furthermore, this invention also considers pitch information when sampling human voice audio. Multiple combinations of pitch and articulation elements are constructed by combining each pitch information with each articulation element. Then, human voice recording is performed separately for each combination of pitch and articulation element to obtain the corresponding audio segment. Moreover, during human voice recording, the duration of the sampling segment, the acquisition target, the recording environment, the signal-to-noise ratio, and the sampling rate are also set.

[0025] In addition, pitch can be determined based on the target pitch range of the human voice source, and multiple pitch information can be obtained by segmenting the target pitch range.

[0026] Preferably, during audio sampling, the present invention further introduces velocity information to set multiple velocity layers and configure the sampling rules for each velocity layer. Then, based on each pitch, human voice audio is recorded separately according to the configuration sampling rules of each velocity layer to obtain multiple audio sample samples. Each audio sample sample includes pitch, velocity layer, articulatory element, and audio segment.

[0027] Specifically, the multiple force layers in this embodiment include a light force layer, a medium force layer, and a strong force layer.

[0028] Specifically, when the intensity level is light, the sampling rule is configured to use breathy and soft singing techniques; when the intensity level is medium, the sampling rule is configured to use standard natural singing techniques; and when the intensity level is strong, the sampling rule is configured to use loud, shouting, or head voice techniques. Furthermore, the intensity level settings are not limited to those given in this embodiment and can be configured according to actual circumstances.

[0029] Step S2: Configure trigger parameters for each audio sample according to preset rules to establish a mapping relationship between trigger parameters and audio sample.

[0030] Specifically, the triggering parameter in this embodiment is velocity information. The velocity information is mapped to the velocity layer in each audio sample, and the pitch information is also mapped to the pitch information in the audio sample. Furthermore, the velocity layer can be set via a velocity range, including a maximum velocity value and a minimum velocity value. Thus, when the user plays, if the current velocity value is between the maximum and minimum velocity values ​​of the corresponding velocity layer, it is considered that the current velocity value matches the corresponding velocity layer.

[0031] Step S3: Obtain MIDI parameters through the MIDI controller and parse the MIDI parameters to obtain pitch information and trigger parameters.

[0032] Step S4: Match several candidate audio samples according to the trigger parameters and mapping relationship, and obtain the corresponding audio sample based on the pitch information and several candidate audio samples, and then play the audio segment according to the corresponding audio sample.

[0033] When a note is triggered, the MIDI controller receives note information, which includes pitch and velocity information. Pitch information indicates which note the user is currently playing or pressing; velocity information indicates the strength of the pressure applied to the note. For example, when a user presses a key, the MIDI controller receives MIDI parameters. Parsing these parameters yields note information including middle C and a velocity of 80. Middle C represents the pitch, and velocity 80 indicates the strength of the pressure applied to the key.

[0034] In other words, when the MIDI controller receives MIDI parameters, it parses the MIDI parameters to obtain pitch and velocity information, that is, pitch and velocity magnitude. First, it selects the corresponding audio sample based on the pitch, and then matches the corresponding velocity layer based on the velocity magnitude. Then, it determines the final audio sample based on the corresponding velocity layer, and then obtains the audio segment of the final audio sample and plays it based on the audio segment.

[0035] More preferably, in order to achieve more expressions in human voice audio playback, the present invention also obtains CC control parameters through a MIDI controller, obtains corresponding modification parameters according to the CC control parameters and the CC control parameter configuration table, processes the corresponding audio segment according to the modification parameters to generate human voice audio signals, and transmits the human voice audio signals to subsequent audio playback devices, such as DWM, headphones, speakers, etc., for playback through the output interface.

[0036] The CC control parameters, typically ranging from CC1 to CC127, each represent a specific control function used to modify audio segments and enrich vocal playback. For example, CC1—the modulation wheel—is usually used to control the intensity of vibrato or expression; for instance, pushing it up on a flute will make the sound more trembling, or it can make the sound of strings fuller by adding vibrato to the audio segment. Other examples include CC7—channel volume, equivalent to the master volume fader for the instrument channel; CC10—panning, determining whether the sound is on the left or right; CC11—expression, a tool for crescendo and diminuendo; and CC64—the sustain pedal, referring to the right pedal commonly used on a piano, its value indicating whether the pedal is pressed or released. In other words, when a note is pressed, the MIDI controller receives the note information and simultaneously receives the CC control parameters. These parameters are then used to adjust the vibrato depth, breathiness ratio, resonance, and other expression parameters of the audio segment to enrich vocal playback.

[0037] The MIDI controller in this invention can take many forms, as long as it supports the MIDI protocol, such as a MIDI guitar, MIDI keyboard, MDI drum pad, MIDI wind instrument, MIDI glove, and motion controller. Furthermore, this invention also provides operation settings for the pitch, velocity, and CC control parameters of the MID controller. Specifically: when the MIDI controller is a MIDI guitar, pitch information is obtained by pressing the corresponding fret, velocity information is triggered by plucking the corresponding string, and CC control parameters are obtained by plucking the guitar wand or modulation wheel; when the MIDI controller is a MIDI keyboard, pitch and velocity information are triggered by pressing different strings, and CC control parameters are obtained by pressing the touch sensor or pitch bend wheel; when the MIDI controller is a MIDI drum pad, pitch and velocity information are triggered by tapping the drum pad; when the MIDI controller is a wind instrument, pitch and velocity information are triggered by blowing air and using the keyboard; when the MIDI controller is a MID glove or motion controller, pitch and velocity information are triggered by hand movements; when the MIDI controller is a virtual MIDI controller within DAW software, pitch and velocity information are triggered by mouse clicks or MIDI Learn.

[0038] This invention extends vocal performance from professional recording studios to real-time live performances and popular music creation scenarios, enhancing the user's performance experience. Based on the standard MIDI protocol, it is compatible with all existing MIDI controller devices and can be implemented without the need for dedicated hardware. This invention can be widely applied in fields such as smart musical instruments, DAW arrangement, live performances, and music education. Example 2

[0039] Based on Embodiment 1, the present invention also provides an embodiment of a MIDI-based human voice timbre driver device, including a memory and a processor. The memory stores a human voice timbre driver program that runs on the processor. The human voice timbre driver program is a computer program. When the processor executes the human voice timbre driver program, it performs the following steps: Audio source package acquisition steps: Record human voice audio based on pronunciation elements and pitch information to obtain multiple audio sample samples; each audio sample sample includes an audio segment, pronunciation elements, and pitch information; Configuration steps: Configure trigger parameters for each audio sample according to preset rules to establish a mapping relationship between trigger parameters and audio sample samples; Triggering steps: Obtain MIDI parameters through the MIDI controller and parse the MIDI parameters to obtain pitch information and trigger parameters; Matching steps: Based on the trigger parameters and mapping relationship, several candidate audio sample samples are obtained. Based on the pitch information and several candidate audio sample samples, the corresponding audio sample sample is obtained. Then, the audio segment of the corresponding audio sample sample is played.

[0040] Furthermore, the triggering step also includes: obtaining CC control parameters and obtaining corresponding modification parameters according to the CC control parameters and the CC control parameter configuration table, processing the corresponding audio segment according to the corresponding modification parameters to generate a human voice audio signal, and transmitting the human voice audio signal to the subsequent audio playback device for playback through the output interface.

[0041] Furthermore, the specific steps for creating an audio source package include: Steps for determining pronunciation elements: Determine pronunciation elements based on requirements; the set of pronunciation elements should include at least standard vowels and commonly used filler sounds; standard vowels include A, E, I, O, and U; commonly used filler sounds include murmurs, auspicious sounds, velar sounds, and auspicious-velar sounds; Pitch configuration steps: Determine the target pitch range of the human voice source, divide the target pitch range into multiple pitches according to the preset segmentation rules, and combine each pitch with each articulation element in the articulation element set to form multiple pitch and articulation element combinations; Recording steps: Sample the audio segment corresponding to each pitch and vocal element combination according to the preset sampling rules; the preset sampling rules include the duration of the sampling segment, the acquisition target, the recording environment, the signal-to-noise ratio, and the sampling rate.

[0042] Furthermore, the process of recording human voice audio based on articulation elements and pitch information to obtain multiple audio sample samples specifically includes: setting multiple velocity layers and configuration sampling rules for each velocity layer, and then recording human voice audio for each pitch and articulation element combination according to the configuration sampling rules of each velocity layer to obtain multiple audio sample samples; one velocity layer corresponds to several audio sample samples.

[0043] Furthermore, multiple intensity levels include a light intensity level, a medium intensity level, and a strong intensity level. When the intensity level is a light intensity level, the sampling rule is configured to use breathy and soft singing methods for sampling. When the intensity level is a medium intensity level, the sampling rule is configured to use standard natural singing methods for sampling. When the intensity level is a strong intensity level, the sampling rule is configured to use loud, shouting, or head voice methods for sampling.

[0044] Furthermore, the configuration steps specifically include: determining the maximum and minimum velocity values ​​of each velocity layer based on the velocity range of the MIDI controller, and then constructing a mapping relationship between each velocity layer and the audio sampling sample based on the maximum and minimum velocity values ​​of the velocity layers. Example 3

[0045] Based on Embodiment 1, the present invention also provides an embodiment: a computer-readable storage medium storing a human voice timbre driver program, wherein the human voice timbre driver program is a computer program, and when executed by a processor, the human voice timbre driver program performs the following steps: Audio source package acquisition steps: Record human voice audio based on pronunciation elements and pitch information to obtain multiple audio sample samples; each audio sample sample includes an audio segment, pronunciation elements, and pitch information; Configuration steps: Configure trigger parameters for each audio sample according to preset rules to establish a mapping relationship between trigger parameters and audio sample samples; Triggering steps: Obtain MIDI parameters through the MIDI controller and parse the MIDI parameters to obtain pitch information and trigger parameters; Matching steps: Based on the trigger parameters and mapping relationship, several candidate audio sample samples are obtained. Based on the pitch information and several candidate audio sample samples, the corresponding audio sample sample is obtained. Then, the audio segment of the corresponding audio sample sample is played.

[0046] Furthermore, the triggering step also includes: obtaining CC control parameters and obtaining corresponding modification parameters according to the CC control parameters and the CC control parameter configuration table, processing the corresponding audio segment according to the corresponding modification parameters to generate a human voice audio signal, and transmitting the human voice audio signal to the subsequent audio playback device for playback through the output interface.

[0047] Furthermore, the specific steps for creating an audio source package include: Steps for determining pronunciation elements: Determine pronunciation elements based on requirements; the set of pronunciation elements should include at least standard vowels and commonly used filler sounds; standard vowels include A, E, I, O, and U; commonly used filler sounds include murmurs, auspicious sounds, velar sounds, and auspicious-velar sounds; Pitch configuration steps: Determine the target pitch range of the human voice source, divide the target pitch range into multiple pitches according to the preset segmentation rules, and combine each pitch with each articulation element in the articulation element set to form multiple pitch and articulation element combinations; Recording steps: Sample the audio segment corresponding to each pitch and vocal element combination according to the preset sampling rules; the preset sampling rules include the duration of the sampling segment, the acquisition target, the recording environment, the signal-to-noise ratio, and the sampling rate.

[0048] Furthermore, the process of recording human voice audio based on articulation elements and pitch information to obtain multiple audio sample samples specifically includes: setting multiple velocity layers and configuration sampling rules for each velocity layer, and then recording human voice audio based on each pitch according to the configuration sampling rules of each velocity layer to obtain multiple audio sample samples; each audio sample sample includes pitch, velocity layer, articulation elements, and audio segments.

[0049] Furthermore, multiple intensity levels include a light intensity level, a medium intensity level, and a strong intensity level. When the intensity level is a light intensity level, the sampling rule is configured to use breathy and soft singing methods for sampling. When the intensity level is a medium intensity level, the sampling rule is configured to use standard natural singing methods for sampling. When the intensity level is a strong intensity level, the sampling rule is configured to use loud, shouting, or head voice methods for sampling.

[0050] Furthermore, the configuration steps specifically include: determining the maximum and minimum velocity values ​​of each velocity layer based on the velocity range of the MIDI controller, and then constructing a mapping relationship between each velocity layer and the audio sampling sample based on the maximum and minimum velocity values ​​of the velocity layers. Example 4

[0051] Based on Embodiment 1, the present invention also provides an embodiment of a MIDI-based human voice timbre driving system, comprising: a MIDI controller; The storage module is used to store the human voice source package and mapping configuration file; A processing module, electrically connected to a MIDI controller, is used to acquire MIDI parameters of the MIDI controller and derive an audio signal using the MIDI-based human voice timbre driving method according to any one of claims 1-5; The audio output module is used to input audio signals into subsequent audio playback devices for playback.

[0052] Furthermore, the MIDI controller is at least one of the following: MIDI guitar, MIDI keyboard, MIDI drum pad, MIDI wind instrument, MIDI glove, and motion controller.

[0053] The above embodiments are merely preferred embodiments of the present invention and should not be construed as limiting the scope of protection of the present invention. Any non-substantial changes and substitutions made by those skilled in the art based on the present invention shall fall within the scope of protection claimed by the present invention.

Claims

1. A MIDI-based method for driving human voice timbre, characterized in that, include: Audio source package acquisition steps: Record human voice audio based on pronunciation elements and pitch information to obtain multiple audio sample samples; each audio sample sample includes an audio segment, pronunciation elements, and pitch information; Configuration steps: Configure trigger parameters for each audio sample according to preset rules to establish a mapping relationship between trigger parameters and audio sample samples; Triggering steps: Obtain MIDI parameters through the MIDI controller and parse the MIDI parameters to obtain pitch information and trigger parameters; Matching steps: Based on the trigger parameters and mapping relationship, several candidate audio sample samples are obtained. Based on the pitch information and several candidate audio sample samples, the corresponding audio sample sample is obtained. Then, the audio segment of the corresponding audio sample sample is played.

2. The MIDI-based human voice timbre driving method according to claim 1, characterized in that, The triggering steps also include: obtaining CC control parameters and obtaining corresponding modification parameters according to the CC control parameters and CC control parameter configuration table, processing the corresponding audio segment according to the corresponding modification parameters to generate a human voice audio signal, and transmitting the human voice audio signal to the subsequent audio playback device for playback through the output interface.

3. The MIDI-based human voice timbre driving method according to claim 1, characterized in that, The specific steps for creating a sound source package include: Steps for determining pronunciation elements: Determine pronunciation elements based on requirements; the set of pronunciation elements should include at least standard vowels and commonly used filler sounds; standard vowels include A, E, I, O, and U; commonly used filler sounds include murmurs, auspicious sounds, velar sounds, and auspicious-velar sounds; Pitch configuration steps: Determine the target pitch range of the human voice source, divide the target pitch range into multiple pitches according to the preset segmentation rules, and combine each pitch with each articulation element in the articulation element set to form multiple pitch and articulation element combinations; Recording steps: Sample the audio segment corresponding to each pitch and vocal element combination according to the preset sampling rules; the preset sampling rules include the duration of the sampling segment, the acquisition target, the recording environment, the signal-to-noise ratio, and the sampling rate.

4. The MIDI-based human voice timbre driving method according to claim 1, characterized in that, Recording human voice audio based on articulation elements and pitch information to obtain multiple audio sample samples specifically includes: setting multiple velocity layers and configuration sampling rules for each velocity layer, and then recording human voice audio based on each pitch according to the configuration sampling rules of each velocity layer to obtain multiple audio sample samples; each audio sample sample includes pitch, velocity layer, articulation elements, and audio segments.

5. The MIDI-based human voice timbre driving method according to claim 4, characterized in that, Multiple intensity levels include light intensity, medium intensity, and strong intensity; when the intensity level is light intensity, the sampling rule is configured to use breathy and soft singing techniques for sampling. When the intensity level is medium, the sampling rule is configured to use the standard natural singing method for sampling; when the intensity level is high, the sampling rule is configured to use the loud voice, shouting, or head voice for sampling.

6. The MIDI-based human voice timbre driving method according to claim 4, characterized in that, The configuration steps specifically include: determining the maximum and minimum velocity values ​​for each velocity layer based on the velocity range of the MIDI controller, and then constructing a mapping relationship between each velocity layer and the audio sample based on the maximum and minimum velocity values ​​of the velocity layers.

7. A MIDI-based human voice timbre driver device, comprising a memory and a processor, wherein the memory stores a human voice timbre driver program that runs on the processor, characterized in that, The human voice timbre driver is a computer program that, when executed by the processor, implements the steps of the MIDI-based human voice timbre driving method as described in any one of claims 1-6.

8. A computer-readable storage medium storing a human voice timbre driver thereon, characterized in that, The human voice timbre driver is a computer program that, when executed by a processor, implements the steps of the MIDI-based human voice timbre driving method as described in any one of claims 1-6.

9. A MIDI-based human voice timbre driving system, characterized in that, include: MIDI controller; The storage module is used to store the human voice source package and mapping configuration file; A processing module, electrically connected to a MIDI controller, is used to acquire MIDI parameters of the MIDI controller and derive an audio signal using the MIDI-based human voice timbre driving method according to any one of claims 1-5; The audio output module is used to input audio signals into subsequent audio playback devices for playback.

10. The MIDI-based human voice timbre driving system according to claim 9, characterized in that, The MIDI controller must be at least one of the following: MIDI guitar, MIDI keyboard, MIDI drum pad, MIDI wind instrument, MIDI glove, and motion controller.