Music making system based on deep learning

Through a music production system based on deep learning, using time-frequency domain hybrid dual U-Net network and full convolutional network for sound source separation and MIDI transcription, the problems of insufficient track separation accuracy and poor music transcription effect in traditional methods are solved, and efficient track separation and visual arrangement are achieved, which is suitable for the field of digital music production.

CN120564673APending Publication Date: 2025-08-29SHANGHAI SECOND POLYTECHNIC UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510868948.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

Traditional audio track separation methods are difficult to effectively decouple between human voices and multi-instrumental parts in complex mixing scenarios, and the music transcription algorithm has poor results in processing music materials with rich harmony and complex rhythms. Professionals are cumbersome to debug parameters, which restricts the efficiency of music creation.

Method used

The music production system based on deep learning is adopted, and the audio source separation is performed using a time-frequency domain hybrid dual U-Net network, and MIDI transcription is realized through a full convolutional network, integrating the effect calculation module, display module and audio output module, providing equalization, modulation, reverb and distortion effects, supporting track visualization and arrangement operations.

Benefits of technology

It realizes high-precision separation and visualization of common sound tracks, lowers the professional threshold, enables non-professional users to quickly complete the preprocessing and arrangement of music materials, and improves the efficiency of music creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564673A_ABST
    Figure CN120564673A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of digital music production, in particular to a music production system based on deep learning, and the system comprises a core calculation module which is used for carrying out the effect calculation according to the demands of a user, and carrying out the sound source separation and MIDI transcription processing based on a deep learning network; the display module is used for displaying the output of the core calculation module; the key module is used for carrying out lifting adjustment on the output of the core calculation module; the audio output module is used for outputting final audio; wherein the display module, the key module and the audio output module are integrated on the core calculation module. According to the method, high-precision separation of common audio tracks can be realized, and meanwhile, the separated audio tracks are visualized through MIDI transcription.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital music production, and in particular to a music production system based on deep learning. Background Art

[0002] In the field of digital music production, track separation and music transcription, as core pre-processing steps, directly affect the quality and efficiency of subsequent creative processes such as arrangement and mixing. Traditional track separation mainly relies on signal processing methods such as Fourier transform and non-negative matrix decomposition, which have problems such as limited separation accuracy and insufficient generalization ability. Especially when dealing with complex mixing scenarios, it is difficult to effectively decouple the human voice and multiple instrument parts. At the same time, in terms of music transcription, traditional algorithms based on acoustic feature extraction are not effective for processing music materials with rich harmonies and complex rhythms. Professionals are often required to repeatedly adjust parameters, which seriously restricts the efficiency of music creation.

[0003] With breakthroughs in deep learning technology in areas like image semantic segmentation and speech recognition, neural network-based audio processing has brought new possibilities to music production. Therefore, to address the common challenges of current music production tools, such as high professional barriers to entry and insufficient intelligent processing capabilities, this paper proposes a music production system based on deep learning. Summary of the Invention

[0004] The purpose of this invention is to provide a music production system based on deep learning, which can achieve high-precision separation of common audio tracks and visualize the separated audio tracks through MIDI transcription.

[0005] To achieve the above object, the present invention provides the following solutions:

[0006] A music production system based on deep learning, comprising:

[0007] The core computing module is used to perform effect calculations based on user needs, as well as perform sound source separation and MIDI transcription processing based on deep learning networks;

[0008] A display module, configured to display the output of the core computing module;

[0009] A key module, used to adjust the output of the core computing module;

[0010] Audio output module, used to output the final audio;

[0011] Among them, the display module, key module, and audio output module are integrated on the core computing module.

[0012] Optionally, the core computing module includes:

[0013] The effect calculation module is used to calculate the effect according to the user's needs and generate mixed audio with the corresponding effect;

[0014] A sound source separation module is used to process the mixed audio based on a preset sound source separation model and output a plurality of single-track sound sources, wherein the sound source separation model is constructed by a hybrid U-Net network based on time domain and frequency domain;

[0015] A MIDI transcription module is used to process the mixed audio based on a preset MIDI transcription model and output a note sequence under the standard MIDI protocol, wherein the MIDI transcription model is constructed through a fully convolutional network.

[0016] Optionally, the effect calculation module includes:

[0017] The equalization effect unit is used to adjust the audio frequency band by cascading several filters to achieve the equalization effect;

[0018] Modulation effect unit, used to superimpose audio frequencies through several low-frequency oscillators to achieve modulation effects;

[0019] The reverb effect unit is used to add a sense of three-dimensionality and space to the audio by simulating the physical space and designing the mixing ratio of dry and wet sounds to achieve a reverberation effect;

[0020] The distortion effect unit is used to change the audio harmonic structure through dynamic compression and waveform deformation of nonlinear functions to achieve a distortion effect.

[0021] Optionally, the sound source separation model includes a time domain path and a frequency domain path, the time domain path and the frequency domain path are connected through a feature fusion layer, the time domain path includes a time domain coding layer and a time domain decoding layer that are mirrored and jump-connected, and the frequency domain path includes a frequency domain coding layer and a frequency domain decoding layer that are mirrored and jump-connected.

[0022] Optionally, processing the mixed audio based on the sound source separation model includes:

[0023] Input the mixed audio to the time domain path, perform feature extraction through several time domain coding layers, and output time domain features;

[0024] The mixed audio is subjected to short-time Fourier transform and then input into the frequency domain path, subjected to feature extraction through several frequency domain coding layers, and outputting frequency domain features;

[0025] After the time domain features and the frequency domain features are fused by the feature fusion layer, they are respectively passed through a plurality of time domain decoding layers and a plurality of frequency domain decoding layers to obtain a time domain output and a frequency domain output;

[0026] After performing inverse short-time Fourier transform on the frequency domain output, the output is added to the time domain output to obtain a final output, namely the single-track sound source.

[0027] Optionally, the MIDI transcription model includes an audio preprocessing unit, a feature extraction unit and a post-processing unit, wherein the audio preprocessing unit is used to resample, peak normalize, and time-frequency convert the input signal to obtain a spectrogram; the feature extraction unit is used to extract features from the spectrogram based on a fully convolutional network to obtain pitch features and note features; and the post-processing unit is used to convert the pitch features and note features into discrete MIDI note events and multi-pitch tracks.

[0028] Optionally, the fully convolutional network includes a feature extraction branch, a jump connection branch and a splicing layer, the feature extraction branch is used to extract deep pitch features based on the spectrogram, the jump connection branch is used to extract high-frequency transient features based on the spectrogram, and the splicing layer is used to splice the deep pitch features and high-frequency transient features to output the final features.

[0029] Optionally, the feature extraction branch includes a first feature extraction layer and a second feature extraction layer, wherein the first feature extraction layer is used to extract the probability of note existence based on the spectrum graph; and the second feature extraction layer is used to extract deep pitch features based on the note existence probability.

[0030] The beneficial effects of the present invention are:

[0031] This paper designs a time-frequency hybrid dual U-Net network for sound source separation, develops four effects units: equalization, modulation, reverberation, and distortion, and implements MIDI transcription based on CQT and a fully convolutional network. This achieves high-precision separation of common audio tracks and visualizes the separated tracks through MIDI transcription, enabling non-professional users to quickly pre-process music materials and perform simple composing operations, paving the way for music adaptation and creation for a wider range of music enthusiasts. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0033] Figure 1 This is a diagram of the overall architecture of the equalization effect unit according to an embodiment of the present invention;

[0034] Figure 2 A flow chart of a real-time processing mechanism according to an embodiment of the present invention;

[0035] Figure 3 is an overall structural diagram of a modulation effect unit according to an embodiment of the present invention;

[0036] Figure 4 1. This is a flowchart of a chorus effect processing according to an embodiment of the present invention;

[0037] Figure 5 This is a flow chart of the rimming effect processing according to an embodiment of the present invention;

[0038] Figure 6 is a flow chart of phase effect processing according to an embodiment of the present invention;

[0039] Figure 7 This is a diagram showing the overall architecture of a reverberation effect unit according to an embodiment of the present invention;

[0040] Figure 8 is an overall structural diagram of a distortion effect unit according to an embodiment of the present invention;

[0041] Figure 9 This is a diagram showing the overall architecture of a sound source separation model according to an embodiment of the present invention;

[0042] Figure 10 This is a diagram of the overall architecture of a fully convolutional network according to an embodiment of the present invention;

[0043] Figure 11 This is a workflow diagram of the post-processing unit according to an embodiment of the present invention. DETAILED DESCRIPTION

[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0045] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0046] This embodiment provides a music production system based on deep learning, including: a core computing module, used to perform effect calculations according to user needs, and to perform sound source separation and MIDI transcription processing based on a deep learning network; a display module, used to display the output of the core computing module; a key module, used to adjust the pitch of the output of the core computing module; and an audio output module, used to output the final audio; wherein the display module, key module, and audio output module are integrated on the core computing module.

[0047] Specifically, this embodiment designs a time-frequency domain hybrid dual U-Net network for sound source separation, develops four types of effects: equalization, modulation, reverberation, and distortion, and implements MIDI transcription based on CQT and a fully convolutional network. This achieves high-precision separation of common audio tracks and visualizes the separated tracks through MIDI transcription, allowing non-professional users to quickly pre-process music materials and perform simple composing operations, providing technical conditions for music adaptation and creation for more music enthusiasts.

[0048] This example uses a Raspberry Pi 5 as the core computing module, the official Raspberry Pi 7-inch IPS touchscreen as the display module, and a three-level independent key module to simulate a single octave of a piano, corresponding to eight natural notes (C, D, E, F, G, A, B, C) and five semitones (up / down). The audio output module uses a USB sound card with a CM108B chip as its core.

[0049] Furthermore, the core computing module includes: an effect calculation module, which is used to perform effect calculations according to user needs and generate mixed audio with corresponding effects; a sound source separation module, which is used to process the mixed audio based on a preset sound source separation model and output several single-track sound sources, wherein the sound source separation model is constructed through a hybrid U-Net network based on time domain and frequency domain; a MIDI transcription module, which is used to process the mixed audio based on a preset MIDI transcription model and output a note sequence under the standard MIDI protocol, wherein the MIDI transcription model is constructed through a fully convolutional network.

[0050] The effect calculation module includes: an equalization effect unit, which adjusts the audio frequency band by cascading several filters to achieve an equalization effect; a modulation effect unit, which superimposes the audio frequency by several low-frequency oscillators to achieve a modulation effect; a reverberation effect unit, which adds a sense of three-dimensionality and space to the audio by simulating the physical space and designing the dry and wet sound mixing ratio to achieve a reverberation effect; and a distortion effect unit, which changes the audio harmonic structure by dynamic compression and waveform deformation through nonlinear functions to achieve a distortion effect.

[0051] (1) Equalization effect unit:

[0052] The function of applying different gains to different frequency bands of real-time audio is achieved by using a multi-level parametric equalizer. Its core modules include audio input / output, filter parameter calculation, multi-threaded real-time processing and GUI interaction. After the user loads the audio file, the system converts the format, forces mono and matches the sampling rate to ensure consistency with the hardware output. The equalization effect unit designs a GUI, which contains parameters for adjusting the frequency band, such as center frequency, gain and quality factor, and can calculate the coefficients of the IIR filter in real time. The background thread processes the original audio through all filters in advance to avoid delays caused by users operating the equalization effect module in real time. The overall architecture of the equalization effect unit is as follows: Figure 1 shown.

[0053] In this embodiment, the equalization effect unit uses a second-order peak filter, and each frequency band corresponds to an independently controllable peak or notch filter. Its transfer function is:

[0054]

[0055] Describes the input and output relationship of the filter, where b0 represents the weight coefficient of the current input sample, b1 represents the weight coefficient of the previous input sample, and b2 represents the weight coefficient of the previous two input samples. Similarly, a1 and a2 represent the weight coefficients of the previous and previous two output samples, respectively. -1 Represents the value of the input signal at the previous sampling moment, z -2 is the value of the input signal at the first two sampling moments. Convert it into a time domain difference equation:

[0056] y[n]=b0x[n]+b1x[n-1]+b2x[n-2]-a1y[n-1]-a2y[n-2] (2);

[0057] The current output y[n] is the weighted sum of the current and previous two input samples minus the weighted sum of the previous two output samples. Different coefficient combinations will change the amplitude-frequency characteristics of the filter. Formulas (3)-(11) describe the calculation method of b0b1b2a1a2 in the transfer function, and its value is calculated from the center frequency f0, gain G and quality factor Q.

[0058] A=10 G / 40 (3);

[0059]

[0060] b0=1+αA (6);

[0061] b1=-2cos(ω0) (7);

[0062] b2=1-αA (8);

[0063] a0=1+α / A (9);

[0064] a1=-2cos(ω0) (10);

[0065] a2=1-αA (11);

[0066] The final normalization coefficient is:

[0067]

[0068] This embodiment uses three filters, corresponding to low frequency, medium frequency and high frequency respectively, which are connected in series. The output of the previous filter corresponds to the input of the next filter, which is called "cascade". This method realizes the multi-band adjustment function of the equalizer.

[0069] In this embodiment, N filters are cascaded, and their transfer functions are H1(z), H2(z), ..., H N (z), u is:

[0070] H 总 (z)=H1(z)×H2(z)×…×H N (z) (13);

[0071] In the actual signal processing process, the amplitude response of each frequency band is expressed as the algebraic sum on the decibel scale, then:

[0072] |H 总 (f)|=|H1(f)|+|H2(f)|+…+|H N (f)| (14).

[0073] The real-time processing mechanism required by this embodiment can be achieved by setting up a pre-calculation buffer and multi-threaded collaborative operation, separating the filtering calculation from the real-time playback thread, pre-processing all possible audio data, generating a playable buffer to store data, and directly reading the memory during playback to reduce the delay problem caused by "real calculation". The real-time processing mechanism flow is shown in Figure 2 .

[0074] Buffer design: When the parameter changes significantly, the buffer thread is awakened, the event flag is set, and the cascade IIR filtering calculation in the thread is started at the same time. After processing, the data is updated to the buffer and the event flag is cleared, indicating that the parameter update is completed. The buffer data is then spliced ​​together to achieve loop playback. When the parameter changes slightly, the original parameters are kept unchanged and the audio playback position is updated in real time, that is, the audio is played.

[0075] (2) Modulation effect unit:

[0076] The modulation effect unit uses three low-frequency oscillators (LFOs) to process the audio to realize the functions of chorus, flanger, and phaser in the modulation effect. Low-frequency oscillator (LFO) is a commonly used tool in music production. It generates a periodic waveform below the frequency range that the human ear can hear (0-20Hz) and superimposes it with the original audio to realize the function of modulating the audio. In this embodiment, after the user loads the audio file, he can choose to use mono or stereo. The modulation effect unit designs a GUI, which contains 4 parameters for adjusting the independent control of the three modulation effects: rate, depth, feedback, mixing and other combinations. The modulation effect unit can calculate the parameters of the LFO in real time. The data passing through the LFO is mixed and applied to the output level control, and the LFO waveform is fed back to the graphical interface in real time. The overall architecture of the modulation effect unit is as follows Figure 3 shown.

[0077] The core goal of the Chorus effect is to simulate the stereo spatial sense of multiple people singing together. Therefore, LFO is used to create multiple slightly delayed copies of the original signal with a small amount of phase shift to simulate the effect of multiple sound sources sounding at the same time. Figure 4 shown.

[0078] The LFO of the chorus effect modulates the delay time of the signal copy, making the delay amount periodic and mixed with the original signal. The formula for the LFO to generate the delay time in this embodiment is:

[0079]

[0080] rate=0.1+4.9·chorus rate ,depth=40·chours depth (16);

[0081] The LFO's rate is used to control the delay time, creating a chorus effect. The rate and depth represent the actual calculated rate and depth values ​​for the chorus effect, respectively. The rate controls the LFO's oscillation frequency, which is expressed as the speed of the delay time change, and ranges from 0.1 to 5 Hz. The depth controls the delay time range of this oscillation frequency, which determines the amplitude of the pitch fluctuation, and ranges from 0 to 40 ms. The delayed signal is followed by the delayed sample:

[0082] signal delayed =input(n-delay(n)) (17);

[0083] Wherein, φ represents the phase of the oscillation frequency. In this embodiment, three φs are applied to the original signal respectively: Generates a 3-way delayed signal that changes the phase of the original processed signal And take the weighted average to get:

[0084]

[0085] Afterwards, chorus out After feedback coefficient feedback chorus , get the feedback signal signal feedback , as shown in formula (19), which represents the intensity of repeated superposition of signals. The range of feedback coefficient feedback chorus ∈(0,0.9).

[0086] signal feedback =chorus out Feedback chorus (19);

[0087] input(n)=input original (n)+signal feedback (n) (20);

[0088] As shown in formula (20), the role of the feedback signal is to inject part of the original processed signal into the original signal of the next frame, which will produce an effect similar to echo superposition, which meets the requirements of the chorus effect.

[0089] output=input original (n)·(1-mix)+chorus out ·mix (21);

[0090] Formula (21) is the final signal output of the current frame, which mixes the dry and wet sounds through the mixing parameter mix, and the final output is the mixed signal output after the chorus effect chorus .

[0091] The flanger effect uses LFO to generate a delayed modulation signal of less than 20ms, and reverberates with the original signal to produce a comb filter effect, making the sound have a "sweep frequency" or "metal" effect. Figure 5 shown.

[0092] The delay time generated by the flanger effect LFO is:

[0093]

[0094] The design of the flanger effect is similar to that of the chorus, but the difference is that there is no phase φ that generates and changes the original processed signal. Therefore, the flanger delay signal is also similar to the chorus effect, but the flanger does not have a three-way delay signal. The rate and depth of the flanger effect represent the rate and depth values ​​actually calculated in the flanger effect (Flanger), respectively. Compared with the chorus, the range of rate and depth in the flanger effect is quite different from that in the chorus effect: the depth∈(0.1~10) of the flanger is shorter, and the depth∈(0.05~10) is also faster, which satisfies the "sweep" effect that the flanger effect needs to produce. The mapping relationship between the two parameters in the flanger effect is also different:

[0095] rate=0.05+9.95·flanger rate , depth=0.1+9.9·flanger depth (twenty three);

[0096] The feedback formula (24) and the mixing ratio formula (25) in the flanger effect have the same design concept and structure as the chorus effect:

[0097] feedback signal =input[n]+feedback·feedback signal [n-delay(n)] (24);

[0098] output=input original (n)·(1-mix)+chorus out ·mix (25).

[0099] The phase effect is different from the above two modulation effects in that, in addition to applying a low-frequency oscillator (LFO), it also requires an all-pass filter (APF) to produce periodic dips in the spectrum, creating a sound effect similar to a "sweep frequency". The all-pass filter only changes the phase of the signal without changing its amplitude or other characteristics. In this modulation effect, the LFO is designed to modulate the center frequency of the filter. See the phase effect processing flow for details. Figure 6 , specifically:

[0100] The all-pass filter is used in the phase effect, and the transfer function is:

[0101]

[0102] Where α represents the filter coefficient, and the center frequency f modulated by the LFO c The formula is:

[0103]

[0104] Among them, f min ,f max Indicates the frequency range. In this embodiment, the frequency range is preset between 200Hz and 2000Hz. rate indicates the rate, and depth indicates the depth. The larger the rate, the more frequent the α changes, the faster the "sweep" speed, and the more "rapid" the effect; the larger the depth, the larger the α change amplitude, the wider the range of the spectrum notch, and the more obvious the effect. The mapping relationship between rate and depth is:

[0105] rate=0.05+9.95·phaser rate ,depth=phaser depth ×100% (28);

[0106] f c The conversion relationship with α is:

[0107]

[0108] The phase effect is enhanced by cascading multiple all-pass filters to enhance the phase shift. In this embodiment, the number of stages is designed to represent the number of cascaded all-pass filters, which directly affects the intensity and complexity of the phase effect. For an N-stage cascaded all-pass filter, its transfer function is:

[0109]

[0110] From this we can get the transfer function of the single-stage filter:

[0111]

[0112] Convert it into a time domain difference equation:

[0113] y k (n) = -α·x k (n)+x k (n-1)+α·y k (n-1) (33);

[0114] Among them, x k (n) is the input of the k-th filter and the output y of the k-1th filter. k (n-1), y k (n) is the output of the kth filter. The total output after cascading N all-pass filters is:

[0115] y(n)=y N (n) (34);

[0116] Subtracting the input x(n) from the output y(n), when the phase difference approaches 0°, the frequencies in that region are canceled out, creating a "spectral notch." When the phase difference approaches 180°, the frequencies in that region are doubled. The LFO controls the center frequency of the all-pass filter, sweeping it across the spectrum and creating a dynamic phase modulation effect.

[0117] The mix parameter (mix) is the same as the chorus and flanger effects, which are parameters for adjusting the intensity of the effect. The formula is:

[0118] output(n)=x(n)·(1-mix)+(x(n)-y(n))·mix (35);

[0119] To prevent clipping, the system also performs soft limiting on output(n):

[0120] output(n)=tanh(output(n)·0.95) (36).

[0121] (3) Reverberation effect unit:

[0122] This embodiment designs a simple reverberation effector, which adds a sense of three-dimensionality and space to the signal by simulating the number and size of the physical space and designing the mixing ratio of dry and wet sounds. Figure 7 The design GUI interface is integrated into the system. After the audio is imported, it will first be filtered through the high / low cut filter to remove the high / low frequency noise that may need to be filtered. In this embodiment, a 4th-order Butterworth low-pass filter is designed to achieve the high-cut function. The normalized cutoff frequency calculation formula is:

[0123]

[0124] Among them, f highcut Indicates the cutoff frequency of high cut, f s Indicates the sampling rate. The system default sampling rate is 44.1kHz. n It represents the normalized cutoff frequency.

[0125]

[0126] Equation (38) is the analog transfer function of the Butterworth filter, where s k is the position of the pole, and N is the filter order. For the 4th order Butterworth, k=0,1,2,3, after substituting the poles, the actual transfer function is:

[0127]

[0128] The continuous time domain is mapped to the discrete time domain through bilinear transformation:

[0129]

[0130] T=1 / f s (42);

[0131] Substitute into H(s), expand and organize into the polynomial form of H(z):

[0132]

[0133] Bidirectional filtering involves applying a filter to the input signal x[n], yielding y1[n]. This filter is then inverted to yield y1[-n], and the same filter is applied again to yield y2[n]. This filter is then inverted to yield y2[n], yielding the final output y[n] = y2[-n]. Bidirectional filtering doubles the order of the original filter, resulting in a steeper transition band.

[0134] The design ideas of low-cut and high-cut are basically the same. The low-cut design is a 4th-order Butterworth high-pass filter, and its transfer function is:

[0135]

[0136] Unlike the low-pass filter, the numerator of its transfer function is s 4 When s→∞, H(s)→1, so that the filter allows high-frequency signals to pass through, and when s→0, H(s)→0, achieving the effect of suppressing low-frequency signals.

[0137] In reverberation processing, wet sound processing usually requires cutting off the low-frequency portion to prevent a large amount of energy accumulation at low frequencies, resulting in muddy sound. To optimize the reverberation effect in this regard, this embodiment designs a high-pass filter after performing high / low-cut filtering on the audio. The transfer function is:

[0138]

[0139] The calculation method of a3, a2, a1, and a0 is the same as that of the high-cut filter.

[0140] The core of the reverb effect is the impulse response (IR), which is expressed as a time function consisting of early reflections ER(t), reverb tail Tail(t) and random noise Noise(t):

[0141] h(t)=ER(t)+Tail(t)+Noise(t) (46);

[0142] The impulse response characteristic describes the characteristics of sound reflection and attenuation in a specific space. In this embodiment, the size (SIZE), diffusion (DIFF) and pre-delay (DELAY) of the design space are used to control the impulse response characteristic.

[0143] The expression of early reflection is:

[0144]

[0145] Represents discrete and easily distinguishable initial reflections, such as the first and second reflections. i represents the amplitude attenuation coefficient of the i-th reflection. In this embodiment, the expression for the influence of SIZE on it is:

[0146] A i =0.85×(0.9-0.05×(SIZE-1.0)) i (48);

[0147] DIFF determines the number and density of reflections, which affects the total number of early reflections N. The greater the number of reflections, the denser the reflection intervals. The actual calculation expression of DIFF is:

[0148] N = 5 + 15 × DIFF (49);

[0149] And τ i There is no mapping relationship with DELAY. The value of pre-delay is τ i size.

[0150] The reverberation tail is generated by random noise after attenuation and filtering, and its expression is:

[0151] Tail(t)=e -αt ·n(t) (50);

[0152] It manifests as dense and gradually decaying multiple reflections. α is the decay coefficient, which controls the decay rate of the reverberation tail. The effect of SIZE on the decay coefficient is:

[0153] α = 5 - 0.6 × SIZE (51);

[0154] The larger the SIZE, the larger the space, the smaller the attenuation coefficient, and the lower the tail decay rate. n(t) is expressed as the Gaussian white noise required for the reverberation tail:

[0155] n(t)~N(0,σ 2 )(52);

[0156] The effect of DIFF on the reverberation tail is reflected in the control of noise variance:

[0157] σ = 0.2 + 0.6 × DIFF (53);

[0158] After the impulse response characteristic h(t) is known, independent convolution calculations are performed for the early reflections and wet sound control of the signal, respectively, as follows:

[0159]

[0160] The final output signal y(t) requires the dry sound, early reflections and wet sound to be mixed and output in proportion:

[0161] y(t)=x(t)·dry_level+y ER (t)·er_level+y Tail (t)·wet_scale (56);

[0162] Among them, wet_scale is mapped to the actual gain value with a nonlinear function to optimize the listening experience:

[0163] wet_scale=0.15·(1-e -2.5·wet_level ) (57);

[0164] Through all the above parameter calculations, the specific parameter ranges and default values ​​of the reverb effect unit are shown in Table 1. Under the default parameter settings, the reverb effect processing is performed using separated audio containing only drum.

[0165] Table 1

[0166] Parameter name scope default value High-cut frequency HCut (20,20000)Hz 20000Hz Low-cut frequency LCut (20,20000)Hz 20Hz Space size SIZE (1,5) 1 Spatial surface number DIFF (3,12) plane 7 sides Delay time (0,500)ms 0ms Dry sound ratio DRY (0,1) 0.7 Early Reflection Ratio ER (0,1) 0.3 Wet sound ratio WET (0,1) 0.3

[0167] (4) Distortion effect unit:

[0168] Distortion is a nonlinear distortion of an audio signal during transmission or processing, including distortion, compression, and clipping of the original waveform. Distortion is an unnecessary effect at the original sound quality level, but in music and sound production, distortion effects are often used to highlight emotions, unique timbre, etc. The distortion effect unit framework designed in this embodiment is as follows: Figure 8 shown.

[0169] The core of the distortion effect unit designed in this embodiment is the design of a nonlinear function. In audio processing, this function changes the harmonic structure of the signal through dynamic compression and waveform deformation, thereby producing a distortion effect. This embodiment designs an A / B mode distortion. Mode A represents hard clipping of the signal. Its function form is:

[0170]

[0171] The distortion threshold (threshold) indicates the critical value at which the signal is clipped. Any signal part outside the ±threshold will be truncated to ±threshold. Before the signal is actually hard clipped, it often needs to be amplified by the pre-gain (pre_gain):

[0172] x pre = x·pre_gain (59);

[0173] The design and algorithm of the mixing ratio (mix) in the distortion effect are the same as the previous method:

[0174] x mix =dry×(1-mix)+wet×mix (60);

[0175] Post-gain is similar to pre-gain in that it amplifies the mixed signal to compensate for the volume loss caused by clipping:

[0176] output=output·post_gain (61).

[0177] B mode means the distortion effect performs soft clipping on the signal. Unlike hard clipping, it uses a hyperbolic tangent function to compress the original signal:

[0178]

[0179] The design principles and algorithms for pre-gain, blend ratio, and post-gain in soft clipping are essentially the same as those for hard clipping. However, the effects of pre-gain and distortion threshold on soft clipping differ from those on hard clipping. In this mode, pre-gain determines how long the signal enters the compression region; higher values ​​result in earlier compression; distortion threshold controls the inflection point of the compression curve; higher values ​​result in a flatter compression curve.

[0180] Furthermore, the sound source separation model includes a time domain path and a frequency domain path, which are connected by a feature fusion layer. The time domain path includes a time domain coding layer and a time domain decoding layer that are mirrored and jump-connected, and the frequency domain path includes a frequency domain coding layer and a frequency domain decoding layer that are mirrored and jump-connected.

[0181] Processing the mixed audio based on the sound source separation model includes: inputting the mixed audio to the time domain path, performing feature extraction through several time domain coding layers, and outputting time domain features; performing short-time Fourier transform on the mixed audio and inputting it to the frequency domain path, performing feature extraction through several frequency domain coding layers, and outputting frequency domain features; after the time domain features and the frequency domain features are fused through the feature fusion layer, they are respectively passed through several time domain decoding layers and several frequency domain decoding layers to obtain time domain output and frequency domain output; after the frequency domain output is inversely short-time Fourier transformed, it is added to the time domain output to obtain the final output, that is, the single-track sound source.

[0182] Specifically, this embodiment needs to convert the input signal from the time domain to the frequency domain, and uses Short-Time Fourier Transform (STFT) to map the global time domain signal to the frequency domain.

[0183] Let the original signal be x(t), and split it into multiple frames of length N, with a defined step size of M. Then the starting position of the mth frame is t = m·M, so the signal of the mth frame is:

[0184] x m (τ)=x(τ+m·M),τ=0,1,...N-1 (63);

[0185] A window function is applied to each frame of signal. This embodiment uses a Hanning window, so:

[0186]

[0187] After the above processing is completed, the windowed signal of each frame can be subjected to discrete Fourier transform (DFT) to obtain:

[0188]

[0189] Where k = 0, 1, ..., N-1 corresponds to the frequency index, and the complete STFT definition is:

[0190]

[0191] This embodiment designs a sound source separation model based on a hybrid U-Net in the time domain and frequency domain. Its overall architecture is as follows: Figure 9 As shown:

[0192] The model primarily consists of an encoding layer, a decoding layer, and a feature fusion layer. The input audio is processed through two parallel paths. The first is the time domain path, where the original waveform is directly input into the time domain encoding layer for processing. The second is the frequency domain path, where the waveform undergoes a short-time Fourier transform (STFT) to generate a complex frequency domain, which is then input into the frequency domain encoding layer for feature extraction. When the feature dimensions (time resolution, frequency resolution, and number of channels) of the two paths are aligned, the time domain and frequency domain feature representations are fused through element-by-element addition. The decoding layer is constructed with a symmetric encoding layer structure. The output spectrum of the frequency domain decoding layer is converted to a time domain waveform via the iSTFT, which is then added to the output of the time domain decoder to form the final separation result.

[0193] 1) Encoder:

[0194] The encoding layer includes encoders based on time domain and frequency domain. The frequency domain encoder is composed of multiple encoding layers stacked together, and the operations of each layer are convolution downsampling, group normalization, and residual DConv branch. Let the input be The convolution downsampling formula of a single-layer encoder is:

[0195] Xconv=Wconv*X+bconv (67);

[0196] Among them, W conv is the convolution kernel weight, the size is determined by the learnable kernel_size, input In the example, B represents the calculation batch, C in Indicates the number of input channels, H×W indicates the height and width of the input. If in the time domain, the input is Where L represents the length of the input. It represents a bias term, which ensures that the output dimension and the number of output channels C out consistent.

[0197] Group normalization divides the input feature channels into several groups, and performs independent normalization in each group to reduce the dependence on batch size. Group G, then the number of channels in each group is C g =C / G. For each group of g features have:

[0198]

[0199]

[0200] Among them, μ g Indicates group mean calculation, which calculates the mean of all elements in the g-th group feature map. gRepresents the g-th group of data of the input feature map, B represents the batch size, and the summation range is: traversal of the index b (batch), c (channel), h (height), and w (width) of all elements in the group. Represents the group variance, which calculates the square difference between all elements in the g-th group feature map and the mean, and measures the degree of dispersion of the data within the group. Normalized output In, γ g and β g are the learnable scale and shift parameters, and ∈ is a small constant to prevent possible division by 0. After group normalization, it also needs to go through the activation function:

[0201]

[0202] The GELU activation function is non-zero in the negative region, and its gradient is relatively continuous. It can dynamically adjust the activation threshold according to the input distribution to enhance the flexibility of the model.

[0203] In this embodiment, the residual branch is designed as a dilated convolution, which inserts gaps between convolution kernel elements to expand the receptive field and is used to model long-term dependencies in audio signals. The mathematical description of dilated convolution is:

[0204]

[0205] Where W[i] represents the convolution kernel weight, and d represents the dilation rate, which controls the sampling interval. Dilated convolution significantly expands the model's receptive field through dilated sampling without increasing computational cost, enabling it to effectively model long-term dependencies in music signals.

[0206] 2) Decoder:

[0207] The decoder structure is a mirror image of the encoder structure, and the high-dimensional features are gradually restored to time domain waveforms or spectrograms through deconvolution upsampling operations. The deconvolution upsampling formula for a single decoding layer is:

[0208] X deconv =W deconv *X in +b deconv (73);

[0209] The upsampling design idea of ​​the decoding layer is the same as the downsampling design idea of ​​the encoding layer. The weight matrix W deconv Different from the encoding layer, the decoding layer also receives the features X from the corresponding encoding layer. skip , called skip connection fusion:

[0210] X fused =X deconv +X skip(74);

[0211] The design ideas of the decoder's group normalization and activation function are exactly the same as those of the encoder, which can be seen in Equations (68), (69), (70) and (71).

[0212] The spectrum graph restored to the original number of channels needs to be restored to the time domain graph through the inverse short-time Fourier transform (iSTFT), and then mixed with the waveform of the time domain branch as the final output result. The calculation formula of iSTFT is:

[0213]

[0214] Where ω[n] is the Hanning window function, x[n] is the inverse Fourier transform result after framing, and H is the jump length, which represents the temporal sampling interval between adjacent frames.

[0215] Furthermore, the MIDI transcription model includes: an audio preprocessing unit for resampling, peak normalization, and time-frequency conversion of the input signal to obtain a spectrogram; a feature extraction unit for extracting features from the spectrogram based on a fully convolutional network to obtain pitch features and note features; a post-processing unit for converting pitch features and note features into discrete MIDI note events and multi-pitch tracks. Among them, the fully convolutional network includes a feature extraction branch, a jump connection branch, and a splicing layer. The feature extraction branch is used to extract deep pitch features based on the spectrogram, the jump connection branch is used to extract high-frequency transient features based on the spectrogram, and the splicing layer is used to splice the deep pitch features and high-frequency transient features to output the final features. Specifically, the feature extraction branch includes a first feature extraction layer and a second feature extraction layer. Among them, the first feature extraction layer is used to extract the probability of note existence based on the spectrogram; the second feature extraction layer is used to extract deep pitch features based on the probability of note existence. Specifically as follows:

[0216] 1) Audio preprocessing unit:

[0217] The audio preprocessing unit is the entry point for the MIDI transcription model. To ensure the uniformity of subsequent inputs, the input signal must first be resampled, its sampling rate converted to 22050Hz, and its amplitude scaled to [-1, 1] using peak normalization. The resampling process includes anti-aliasing filtering:

[0218]

[0219] Indicates that the signal is above the Nyquist frequency f N The transfer function of the part to be filtered out, where, due to the sampling theorem, the resampled signal frequency must be greater than or equal to twice the Nyquist frequency, otherwise spectrum aliasing will occur. In this embodiment, the resampled signal frequency is designed to be 22050 Hz, so f N The calculation method is:

[0220]

[0221] For the filtered signal x f Perform integer multiple decimation. Since the common input audio sampling rate is 44100Hz, the preset decimation factor D is 2, which means that one point is retained every other sample, resulting in:

[0222] x d [n] = x f [n·D] (80);

[0223] After resampling, the sampling rate of the input signal is fixed to 22050Hz, and its Nyquist frequency is 11025Hz, covering the fundamental frequencies of all notes from B1 to C9.

[0224] After resampling, the signal undergoes peak normalization to scale the amplitude of the waveform to [-1,1]. Its mathematical description is:

[0225]

[0226] Wherein, max(|x|) represents the maximum absolute value of the amplitude in the signal, and A represents the threshold of the peak amplitude, which is set to 1 in this embodiment.

[0227] Since the fully convolutional network (FCN) requires the input signal to be represented in the frequency domain, the constant-Q transform (CQT) is used to perform time-frequency conversion on the input signal in the time domain. The calculation formula of the constant-Q transform is:

[0228]

[0229] Where, k represents the kth frequency band, N k Represents the window length of the frequency band, ω k [n] represents the window function. This module uses the Hanning window, and the window function is:

[0230]

[0231] Q is the core parameter in the constant Q transform, which determines the different combinations of time resolution and frequency resolution for each frequency band. The calculation formula for Q is:

[0232]

[0233] Among them, f k represents the center frequency of the kth frequency band, Δf k Indicates the bandwidth of the frequency band. Since the value of Q is fixed, the bandwidth can be obtained as:

[0234]

[0235] When the center frequency increases, the bandwidth also becomes wider, and the window length N k The formula is:

[0236]

[0237] When the center frequency f k As increases, the window length decreases, and the time span of the corresponding frequency band k decreases, thereby improving the time resolution and making the function more precise in capturing the time characteristics of the high-frequency band. Combining equations (78) and (80), we can obtain that the window length is inversely proportional to the bandwidth:

[0238]

[0239] 2) Feature extraction unit:

[0240] After audio preprocessing, the MIDI transcription model outputs a spectrogram of shape (T × F × 1) to the feature extraction unit, or fully convolutional network. T represents the number of time frames after the audio is framed according to a fixed time window, reflecting the signal's temporal changes. F represents the number of frequency bands per time frame, reflecting the frequency band corresponding to the pitch. The channel number 1 indicates that the CQT outputs a single-channel grayscale spectrogram.

[0241] The architecture of the fully convolutional network (FCN) designed in this embodiment is as follows Figure 10 As shown, it has two branches and outputs three features: note onset / offset, quantized pitch (Pitch) and continuous offset (Pitch Bend).

[0242] Harmonics are frequency components that are integer multiples of the fundamental frequency. The harmonic stacking layer replicates the original CQT spectrum multiple times across the number of channels and shifts it along the frequency axis to align the corresponding frequency points of the harmonics. Finally, all these shifted spectra are stacked into a three-dimensional tensor (time × frequency × harmonic channels). This stacked three-dimensional tensor serves as the input to the subsequent convolutional layer.

[0243] The full convolutional network has two branches, one of which is Y output through three convolutional layers p, is the posterior graph of Multipitch Estimation, which is used to predict whether a note exists at each time and frequency point, that is, the pitch activation probability. Among them, the first layer of convolution is a two-dimensional convolution with 32 output channels, a convolution kernel size of 5*5, and a stride of (1, 3). The input tensor is For the output tensor Y1∈R T×[F / 3]×32 , and its calculation formula is:

[0244]

[0245] Among them, i,j controls the convolution kernel weight W c (i, j, k), k traverses all input channels. In the input index X(t+i,3f+j,k), t+i means that the current time point t is the center, and 2 frames are extended before and after. 3f+j means that the current frequency point 3f is the center, and 2 original frequency points are extended before and after. In this way, there are 5 frequency points in time and frequency. c A bias term is introduced to ensure that even if the input is all 0, it can still be activated by the activation function.

[0246] The input tensor of the second convolutional layer is the output tensor Y1∈R of the first layer. T×[F / 3]×32 , calculated as:

[0247]

[0248] The second convolution layer is similar to the first convolution layer. The stride of the second layer is the default (1, 1), which means no downsampling is done to avoid losing some information about the start and end of the notes. The number of convolution kernels c in the second layer is set to 16, so its output tensor Y2∈R T×[F / 3]×16 .

[0249] The third convolutional layer is used to generate Y p The bottom layer of the convolution layer is to map the multi-channel features extracted by the first two layers to a single-channel probability map to represent the probability of the output note. Its input tensor is the output tensor Y2∈R of the second convolution layer. T ×[F / 3]×16 , and its calculation formula is:

[0250]

[0251] The third convolution layer is similar to the second convolution layer, and there is no downsampling. The output tensor Y p (t,f) does not have an output channel parameter c; its purpose is to weight the features of the 16 channels in the input tensor into a single channel to align with the output requirements of the posterior graph. Furthermore, the third convolutional layer uses the sigmoid activation function σ, which maps the output to the range [0, 1], thereby mapping the features to note presence probabilities.

[0252] Preliminary note existence probability posterior map Y p After two layers of two-dimensional convolution, the deep pitch feature Y is output n , for the input feature Y p ∈R B×T×F×1 Perform a two-dimensional convolution with an output channel number of 32, a convolution kernel size of 7*7, and a stride of 1*3:

[0253]

[0254] i, j represent the index on the time axis and frequency axis respectively, (-3, 3) a total of seven points, k represents the number of input channels, here by Y p k is 1, c represents the number of output channels, and this layer is set to 32. Each output channel is calculated independently, and the cumulative sum of c is not included in the formula. The convolution kernel size designed for this layer ensures its ability to extract contextual features over a large time span, corresponding to the duration and beat of the notes.

[0255] The second two-dimensional convolutional layer takes the output of the first layer Y3∈R B×T×F×32 As input, design a two-dimensional convolution with an output channel of 1 and a convolution kernel size of 7*3:

[0256]

[0257] From Y p to Y n The two-layer two-dimensional convolution uses a wider convolution kernel on the time axis to capture notes with continuous characteristics in time, while the second layer of convolution uses a narrower convolution kernel on the frequency axis to refine the pitch features with more significant frequency boundaries and reduce the possibility of harmonic misjudgment.

[0258] Deep Pitch Feature Y n It will enter a splicing layer (Contact) and be spliced ​​with the output of the jump connection branch to generate Y o The input of the skip connection comes from the output of the harmonic stack layer By designing a two-dimensional convolution layer with 32 output channels, a convolution kernel size of 5*5, and a stride of 1*3:

[0259]

[0260] For this network, the 5*5 convolution kernel size can cover the local time-frequency region to capture high-frequency transient features, such as the impulse response of the note onset. The concatenation layer receives the high-frequency transient features Y from the jump connection branch. c ∈R B×T×F×32 With deep pitch characteristics Y n ∈R B×T×F×1, which superimposes the two outputs in the channel dimension:

[0261]

[0262] The obtained Y con ∈R B×T×F×33 Set the Y c With 32 channels of Y n Merged into 33 channels. After merging, Y con After a two-dimensional convolution layer with an output channel of 1 and a convolution kernel size of 3*3:

[0263]

[0264] The response of the high-frequency transient channel is weighted and combined with the probability value of the deep pitch channel through the weight matrix W, and the final output is the note detection event probability Y of the multi-task output o .

[0265] This embodiment applies batch normalization and activation function ReLu between some two-dimensional convolutional layers in the network. In deep neural networks, the input of a certain middle layer is usually the output of the previous layer, and the parameters of the previous layer will cause its distribution to change significantly when updated, which is called internal covariate shift. It will cause subsequent layers to constantly adapt to the change, reducing training efficiency. The reason for introducing batch normalization is that it can effectively solve this problem by standardizing the input distribution of each layer. It divides the input into B = {x1, x2, ..., x m} Mini-batch data and calculate its mean and variance:

[0266]

[0267] It is then normalized by adding a small constant ∈ to ensure numerical stability:

[0268]

[0269] Introducing linear learnable parameter scaling γ and offset β, we get the final output:

[0270]

[0271] After batch normalization, an activation function ReLu is required. The calculation formula of ReLu is:

[0272]

[0273] A Sigmoid activation function (σ) is also applied before each output head, and its mathematical expression is:

[0274]

[0275] 3) Post-processing unit:

[0276] The post-processing unit needs to convert the FCN output Y o Converts to discrete MIDI note events (Note Events) and multi-pitch tracks (Multipitch). It accepts the multitasking output header (Y o ,Y n ,Y p ).like Figure 11 As shown, it can be divided into the following core processing steps: starting point detection, note generation and correction, and multi-pitch trajectory generation.

[0277] Starting point detection will start the probability map Y o Converted into a set of reliable note starting point candidates, providing the starting time point and pitch positioning point for subsequent note information generation. o It is a two-dimensional matrix with time and frequency as the axes. Each point represents the probability of the note starting at the corresponding time and pitch corresponding frequency position. Peak Picking compares the probability of the current time frame with the probability of the adjacent frames before and after on the time axis. If the probability of the current frame is higher than the probability of the adjacent frames before and after and exceeds the confidence threshold of 0.5, it is marked as a peak. The candidate list of starting points is generated in descending order of time (t i ,f i ), t i Indicates the start time, f i Indicates the pitch, such as a MIDI number or the corresponding frequency value.

[0278] The starting point candidate list generated by starting point detection and Y n Combined and activated by comparing the threshold τ n and a minimum duration of 11 frames to generate note events with accurate start and end times. Note generation is performed for each starting point (t i ,f i ) Scan Y along the positive time axis n The corresponding pitch is f i The activation probability of τ is determined when the activation probability of 11 consecutive frames is less than τ n , the current event is recorded as a note-off event. Note merging requires checking whether two adjacent notes have the same pitch and are separated by less than 50ms. If so, the two notes are merged, with the start and end times of the merged note being the earliest start time and the latest stop time. After testing, the region is set to 0 to prevent repeated testing of the same region.

[0279] Multi-pitch trajectory generation will be pitch probability map Yp Converts to a continuous and smooth pitch trajectory to capture subtle changes in notes such as glissando and vibrato. p , where each time frame contains 3 frequency bins corresponding to the subdivision within a semitone. Peak detection scans the Y axis according to the frequency axis p , to filter out the frequency segment with the highest probability. To generate a smooth pitch trajectory, a quadratic curve fitting (Parabolic Interpolation) is used. First, for each time frame, the bin with the highest probability is selected, along with its adjacent bins. Assuming the probabilities of these bins are p1, p2, and p3, respectively, the corresponding center frequencies are f1, f2, and f3, respectively. The actual pitch f is:

[0280]

[0281] This embodiment uses a U-Net network to fuse time domain and frequency domain features, and adopts a dual U-Net network for sound source extraction and separation, thereby improving the accuracy and speed of sound source separation; adopts a CQT+ fully convolutional network to extract features from the time domain and frequency domain, and realizes the start and end detection, pitch extraction and rhythm quantization functions of multi-pitch notes in a multi-instrument scenario; realizes four effects of equalization, reverberation, modulation and distortion, and the parameters of these four effects can be manually controlled, and a visual effect preview function is adopted; adopts bidirectional filtering and LFO to compensate for the time delay after the note generation, effectively performing real-time audio analysis on the Raspberry Pi 5, and making it meet the requirements of professional-level music production.

[0282] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.

Claims

1. A music production system based on deep learning, characterized in that: include: The core computing module is used to perform effect calculations based on user needs, as well as perform sound source separation and MIDI transcription processing based on deep learning networks; A display module, configured to display the output of the core computing module; A key module, used to adjust the output of the core computing module; Audio output module, used to output the final audio; Among them, the display module, key module, and audio output module are integrated on the core computing module.

2. The music production system based on deep learning according to claim 1, characterized in that The core computing module includes: The effect calculation module is used to calculate the effect according to the user's needs and generate mixed audio with the corresponding effect; A sound source separation module is used to process the mixed audio based on a preset sound source separation model and output a plurality of single-track sound sources, wherein the sound source separation model is constructed by a hybrid U-Net network based on time domain and frequency domain; A MIDI transcription module is used to process the mixed audio based on a preset MIDI transcription model and output a note sequence under the standard MIDI protocol, wherein the MIDI transcription model is constructed through a fully convolutional network.

3. The music production system based on deep learning according to claim 2, characterized in that The effect calculation module includes: The equalization effect unit is used to adjust the audio frequency band by cascading several filters to achieve the equalization effect; Modulation effect unit, used to superimpose audio frequencies through several low-frequency oscillators to achieve modulation effects; The reverb effect unit is used to add a sense of three-dimensionality and space to the audio by simulating the physical space and designing the mixing ratio of dry and wet sounds to achieve a reverberation effect; The distortion effect unit is used to change the audio harmonic structure through dynamic compression and waveform deformation of nonlinear functions to achieve a distortion effect.

4. The music production system based on deep learning according to claim 2, characterized in that The sound source separation model includes a time domain path and a frequency domain path, the time domain path and the frequency domain path are connected by a feature fusion layer, the time domain path includes a time domain coding layer and a time domain decoding layer in a mirror structure and jump connection, and the frequency domain path includes a frequency domain coding layer and a frequency domain decoding layer in a mirror structure and jump connection.

5. The deep learning-based music production system according to claim 4, characterized in that: Processing the mixed audio based on the sound source separation model includes: Input the mixed audio to the time domain path, perform feature extraction through several time domain coding layers, and output time domain features; The mixed audio is subjected to short-time Fourier transform and then input into the frequency domain path, subjected to feature extraction through several frequency domain coding layers, and outputting frequency domain features; After the time domain features and the frequency domain features are fused by the feature fusion layer, they are respectively passed through a plurality of time domain decoding layers and a plurality of frequency domain decoding layers to obtain a time domain output and a frequency domain output; After performing inverse short-time Fourier transform on the frequency domain output, the output is added to the time domain output to obtain a final output, namely the single-track sound source.

6. The music production system based on deep learning according to claim 2, characterized in that The MIDI transcription model includes an audio preprocessing unit, a feature extraction unit, and a post-processing unit. The audio preprocessing unit is used to resample, peak normalize, and time-frequency convert the input signal to obtain a spectrogram; the feature extraction unit is used to extract features from the spectrogram based on a fully convolutional network to obtain pitch features and note features; and the post-processing unit is used to convert the pitch features and note features into discrete MIDI note events and multi-pitch tracks.

7. The deep learning-based music production system according to claim 6, characterized in that: The fully convolutional network includes a feature extraction branch, a jump connection branch and a splicing layer. The feature extraction branch is used to extract deep pitch features based on the spectrogram, the jump connection branch is used to extract high-frequency transient features based on the spectrogram, and the splicing layer is used to splice the deep pitch features and high-frequency transient features to output the final features.

8. The deep learning-based music production system according to claim 7, characterized in that: The feature extraction branch includes a first feature extraction layer and a second feature extraction layer, wherein the first feature extraction layer is used to extract the probability of note existence based on the spectrum graph; the second feature extraction layer is used to extract deep pitch features based on the note existence probability.