Multi-instrument collaborative synthesis method based on matlab and related devices

By setting the standard sampling frequency and time vector on the MATLAB platform, establishing the scale frequency mapping, generating the target fundamental frequency, and performing multi-track time axis alignment and mixing on the audio signal of the target instrument, the problems of complex operation and insufficient multi-instrument collaborative synthesis of existing tools are solved, and high-precision multi-instrument collaborative synthesis is achieved.

CN122073111APending Publication Date: 2026-05-22JIANGHAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-04
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing digital music synthesis tools have high operational barriers, lack a unified integrated framework, cannot meet the precision requirements of scientific research and professional creation, and have insufficient multi-instrument collaborative synthesis capabilities, which limits the industry's innovation and upgrading.

Method used

A MATLAB-based multi-instrument collaborative synthesis method is proposed. This method constructs a time vector corresponding to the duration of a note by setting a standard sampling frequency, establishes a scale frequency mapping relationship, generates a target fundamental frequency, and performs multi-track time axis alignment and mixing on the audio signal of the target instrument to achieve multi-instrument collaborative synthesis.

Benefits of technology

It lowers the operational threshold, improves pitch consistency and auditory layering, enhances the controllability of timbre differences among multiple instruments, suppresses energy fluctuations, and meets the needs of high-precision professional creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122073111A_ABST
    Figure CN122073111A_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-instrument collaborative synthesis method and related equipment based on MATLAB.The method comprises the following steps: setting standard sampling frequency, and constructing time vector corresponding to note duration based on the standard sampling frequency, establishing the frequency mapping relationship of scale covering the preset sound range, and determining the target fundamental frequency based on the input pitch parameter;For at least two target instruments, respectively call the acoustic parameter configuration corresponding to the target instrument, generate the single-instrument synthesis signal of the target instrument based on the target fundamental frequency;The audio signals of each target instrument are aligned on the multi-track time axis, and the audio signals of each target instrument are mixed based on the alignment result to generate a multi-instrument collaborative synthesis audio signal.Various solutions on the market can solve the problem that the operation threshold is high, lacks a unified integrated framework, and restricts the industry innovation and upgrading.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and more specifically, to a multi-instrument collaborative synthesis method and related equipment based on MATLAB. Background Technology

[0002] Digital music synthesis and intelligent interaction technologies have become core supports for the cross-integration of multiple fields such as audio engineering, pan-entertainment, intelligent education, and immersive media. With the global digital music market exceeding $100 billion, the accelerated implementation of the metaverse and VR / AR industries, and the continuous increase in the penetration rate of intelligent education, the industry has created an unprecedented and urgent demand for high-precision, integrated, and cross-scenario audio synthesis tools. However, current market solutions suffer from numerous fatal bottlenecks and structural pain points, severely restricting industry innovation and upgrading, and creating a huge gap between technological supply and market demand.

[0003] First, professional-grade software such as Cubase and Logic Pro have high barriers to entry. They are essentially black-box products designed for professional musicians, with completely closed underlying acoustic algorithms. Users cannot call core modules for secondary development or academic verification, which forces researchers to spend months or even years repeatedly building the basic framework. At the same time, these software programs are extremely complex to operate, requiring professional music theory knowledge and practical experience. Ordinary users and researchers face the dilemma of being discouraged as soon as they start learning. Furthermore, the development of customized functions requires licensing fees of hundreds of thousands of yuan, which significantly increases the innovation costs for small and medium-sized teams and research institutions.

[0004] Secondly, as the most mainstream platform in the global scientific research and engineering fields, MATLAB has a powerful audio processing toolbox, but the existing related tools are severely fragmented. Functions such as audio synthesis, interactive interface, and resource management are scattered across hundreds of independent functions and toolkits, lacking a unified integrated framework. The interactive experience is extremely outdated and cannot meet the modern operational needs of WYSIWYG. Furthermore, the technical architecture is outdated, with most tools remaining at the basic synthesis logic of superimposed single sine waves, failing to incorporate core acoustic characteristics such as nonlinear overtones, natural vibrato, and dynamic envelopes, and thus unable to meet the precision requirements of scientific research and professional creation.

[0005] Third, most simple synthesis tools only support single-instrument simulation, and their acoustic models do not consider details such as nonlinear overtones, natural vibrato, and the textures of plucking or striking strings of traditional instruments, resulting in a lack of realism in the synthesized audio. Furthermore, existing tools lack a unified scheduling framework for multi-instrument collaborative synthesis and lack integrated user permission management and music resource navigation capabilities. The professionalism, ease of use, and integration required for music synthesis are difficult to satisfy in the current market. Summary of the Invention

[0006] The summary section introduces a series of simplified concepts, which will be further explained in detail in the detailed description section. The summary section of this invention is not intended to limit the key features and essential technical features of the claimed technical solution, nor is it intended to determine the scope of protection of the claimed technical solution.

[0007] To address the issues of high operational barriers and a lack of unified integration frameworks in current market solutions, which hinder industry innovation and upgrading, this invention proposes a multi-instrument collaborative synthesis method based on MATLAB, comprising: Set a standard sampling frequency, construct a time vector corresponding to the note duration based on the standard sampling frequency, establish a scale frequency mapping relationship covering a preset range, and determine the target fundamental frequency based on the input pitch parameters; For at least two target musical instruments, the acoustic parameter configurations corresponding to the target musical instruments are called respectively, and a single-instrument synthesized signal of the target musical instrument is generated based on the target fundamental frequency; Multitrack time axis alignment is performed on the audio signals of each of the target instruments, and the audio signals of each of the target instruments are mixed based on the alignment results to generate a multi-instrument collaborative synthesized audio signal.

[0008] Secondly, this invention also proposes a multi-instrument collaborative synthesis device based on MATLAB, comprising: The determining unit is used to set a standard sampling frequency, construct a time vector corresponding to the note duration based on the standard sampling frequency, establish a scale frequency mapping relationship covering a preset range, and determine the target fundamental frequency based on the input pitch parameters; The processing unit is configured to, for at least two target musical instruments, respectively call the acoustic parameter configuration corresponding to the target musical instrument, and generate a single-instrument synthesized signal of the target musical instrument based on the target fundamental frequency; A synthesis unit is used to perform multi-track time axis alignment on the audio signals of each of the target instruments, and to mix the audio signals of each of the target instruments based on the alignment results to generate a multi-instrument collaborative synthesized audio signal.

[0009] Thirdly, an electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program stored in the memory to implement the steps of the MATLAB-based multi-instrument collaborative synthesis method as described in any of the first aspects above.

[0010] Fourthly, the present invention also proposes a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the MATLAB-based multi-instrument collaborative synthesis method of any of the preceding claims of the first aspect.

[0011] In summary, the multi-instrument collaborative synthesis method based on MATLAB proposed in this application establishes a scale frequency mapping relationship covering a preset range by setting a standard sampling frequency and constructing a time vector corresponding to the note duration based on the standard sampling frequency, and determining the target fundamental frequency based on the input pitch parameters; for at least two target instruments, the acoustic parameter configuration corresponding to the target instrument is called respectively, and a single-instrument synthesized signal of the target instrument is generated based on the target fundamental frequency; multi-track time axis alignment is performed on the audio signals of each target instrument, and the audio signals of each target instrument are mixed based on the alignment result to generate a multi-instrument collaborative synthesized audio signal. A unified sampling frequency and a unified time base form the numerical foundation for multitrack synthesis. A mapping from scale to fundamental frequency ensures pitch consistency. Then, a parametric acoustic model of the instrument transforms the same fundamental frequency into single instrument tracks with different timbres. Finally, multitrack alignment and mixing strategies achieve part coordination and energy management. In MATLAB, it can be organized as follows: a global configuration module for sampling frequency, range, tempo, and timeline; a pitch-to-frequency module for mapping functions or looking up tables; an instrument synthesis module that provides a set of parameter structures and synthesis functions for each instrument; a multitrack arrangement module for placement, zero padding, and alignment; and finally, a mixing output module for a streamlined call from gain, normalization, dynamic control to export playback. This supports expansion from single notes to complete musical phrases and even multi-section arrangements. Therefore, firstly, because the pitch and rhythm are consistent across multiple instruments, the synthesized sound caused by inconsistent mapping or imprecise timing alignment is significantly reduced; secondly, because the timbre differences of multiple instruments are generated controllably by a parametric model, the functions of the main melody, harmony, and backing vocal parts can be clearly separated, improving the listening experience and recognizability; thirdly, through amplitude and dynamic management in the alignment and mixing stages, clipping can be suppressed and energy fluctuations caused by phase cancellation can be reduced, improving overall loudness and stability. Attached Figure Description

[0012] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit this specification. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 A schematic diagram of a multi-instrument collaborative synthesis method based on MATLAB is provided for an embodiment of this application; Figure 1a A schematic diagram of the software architecture for a MATLAB-based multi-instrument collaborative synthesis method provided in this application embodiment; Figure 1bA flowchart illustrating an acoustic synthesis algorithm based on a MATLAB-based multi-instrument collaborative synthesis method is provided in an embodiment of this application. Figure 1c A flowchart illustrating another MATLAB-based multi-instrument collaborative synthesis method provided for embodiments of this application; Figure 1d A schematic diagram of the software interface for a MATLAB-based multi-instrument collaborative synthesis method provided in this application embodiment; Figure 2 A schematic diagram of a multi-instrument collaborative synthesis device based on MATLAB is provided in the embodiments of this application; Figure 3 This is a schematic diagram of an electronic device structure for a multi-instrument collaborative synthesis method based on MATLAB, provided as an embodiment of this application. Detailed Implementation

[0013] The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus. The technical solutions of the embodiments of this application will now be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them.

[0014] To address the issues of high operational barriers and a lack of a unified integration framework in current market solutions, which hinder industry innovation and upgrading, please refer to [link / reference needed]. Figure 1 This is a schematic diagram of a multi-instrument collaborative synthesis method based on MATLAB provided in an embodiment of this application, which may specifically include steps S110 to S130.

[0015] S110, set a standard sampling frequency, construct a time vector corresponding to the note duration based on the standard sampling frequency, establish a scale frequency mapping relationship covering a preset pitch range, and determine the target fundamental frequency based on the input pitch parameters.

[0016] S120: For at least two target musical instruments, respectively call the acoustic parameter configuration corresponding to the target musical instrument, and generate a single-instrument synthesized signal of the target musical instrument based on the target fundamental frequency.

[0017] S130, perform multi-track time axis alignment on the audio signals of each of the target instruments, and mix the audio signals of each of the target instruments based on the alignment results to generate a multi-instrument collaborative synthesized audio signal.

[0018] For example, the sampling frequency can determine the density of discrete sampling points per unit time, thereby determining the highest frequency limit and time resolution that the synthesized signal can express; the time vector maps each sampling point to the physical time axis, so that subsequent time-domain operations such as envelope, vibrato, and phase accumulation can strictly fall on each sampling point; the scale frequency mapping relationship is used to stably convert the pitch input of musical semantics, such as note name, the scale number of the twelve equal temperament, or MIDI pitch number, into the target fundamental frequency that can be used for numerical synthesis. In MATLAB, you can first set a standard sampling frequency, commonly 44100 Hz or 48000 Hz, which are compatible with common audio playback and professional audio links, respectively. Then, for a note duration such as 0.5 seconds, 1 second, or 1 beat, construct a time vector from zero to that duration. The step of the time vector is equal to the sampling period, thus ensuring that the total number of sampling points equals the note duration multiplied by the sampling frequency and rounded. Next, establish a frequency table covering a preset pitch range, for example, using a reference note, usually middle A. Calculate the frequency of each pitch according to the exponential relationship of the twelve-tone equal temperament, and implement the mapping from pitch parameters to frequencies in the program using a lookup table or calculation function. The pitch parameters can be MIDI numbers, note name strings, or custom scale numbers. Finally, obtain the target fundamental frequency based on the input pitch parameters. In this way, on the one hand, the sampling frequency and time vector ensure that the synthesized signal is numerically stable and compatible in the playback chain, avoiding high-frequency distortion and timing jitter caused by insufficient sampling points; on the other hand, frequency mapping ensures that the same pitch uses the same fundamental frequency reference in different instrument modules, avoiding pitch drift or discrepancies in the sound of notes with the same name when synthesizing multiple instruments. For example, when the user inputs MIDI number 69, 440 Hz is mapped as the target fundamental frequency; when the user inputs a pitch parameter one octave higher than middle A, the target fundamental frequency automatically becomes 880 Hz; when the user inputs a pitch parameter in a lower range, such as around 110 Hz, the time vector contains more periodic information for the same note duration, which is beneficial for expressing a rich timbre through overtone superposition, while still maintaining accurate beat length.

[0019] For example, the same fundamental frequency can be used to generate single-instrument signals with different timbre characteristics through different acoustic models and parameter constraints. From a signal processing perspective, the differences in instrument timbre are mainly determined by the spectral structure (such as overtone distribution, the ratio of harmonic to non-harmonic components), the temporal envelope (such as attack speed, decay pattern, sustain energy, and release tail), and micro-perturbations (such as vibrato, small frequency drift, and noise texture). Therefore, in MATLAB, each instrument can be abstracted into a set of acoustic parameter configurations, including, for example, the overtone order range, the attenuation curves of each overtone amplitude, and the phase. Initialization methods, envelope parameters, vibrato frequency and depth, noise texture weights, filter parameters, etc.; when generating a single-instrument synthesized signal, a sine or cosine wave of the fundamental frequency component is first constructed on the time vector, and then multi-order overtone components are superimposed according to the parameter configuration, and different amplitudes and phases are assigned to form the target spectral profile. At the same time, an envelope function is applied to the entire signal to shape the energy change over time like a real instrument, and vibrato modulation, such as using a low-frequency oscillation signal to make small swings on the instantaneous frequency or phase, and texture noise, such as low-pass or band-passing white noise, can be superimposed according to the configuration to form a texture similar to plucking strings or airflow. Thus, even if multiple instruments use the same target fundamental frequency, the final signal can be significantly distinguished in terms of sound. For example, the piano emphasizes a clear attack and faster decay, strings emphasize smooth sustain and slow release, and the flute emphasizes fewer high-order overtones and a more stable fundamental frequency component. At the same time, through parameterization, the same instrument can be automatically adapted to different registers. For example, high-order overtones are appropriately enhanced in the high register to avoid weakness, and excessively strong high-order overtones are appropriately suppressed in the low register to avoid muddiness. For example, when synthesizing two instruments, a piano and a guitar, the piano parameter configuration can use a faster onset and decay, and add a slight percussion noise at the moment of onset to simulate string striking. The guitar parameter configuration can add more obvious string texture noise after onset and set a stronger fundamental frequency to the proportion of second and third overtones. When synthesizing strings and a flute, the string configuration can add more obvious vibrato and set a longer release tail, while the flute configuration can reduce non-harmonic noise and shape the formant characteristics of wind instruments through slight bandpass filtering. This allows two or more instruments to form a clear division of voices on the same melodic line, rather than simply superimposing the same frequencies and causing a blurred listening experience.

[0020] For example, multiple single instrument signals can be regarded as multi-track audio, synchronized at the same sampling frequency and the same global time base, and then energy and space are allocated according to a preset mixing strategy after synchronization, so as to form a synthesis result that sounds harmonious and does not conflict. If different tracks are inconsistent in terms of start time, note boundary or sampling point length, direct addition will result in beat misalignment, transient superposition abnormality or even energy fluctuations caused by phase cancellation. Therefore, time axis alignment is required first. In practical implementation, a unified global timeline can be established for each track in MATLAB. For example, a global sampling point index can be generated based on the beat and note sequence, and then each single instrument signal can be placed at the corresponding starting index position. For tracks of different lengths, zero padding or truncation can be used to make their lengths consistent. For scenarios requiring finer alignment, such as when a track has artificial delay or a start-note offset introduced by the model, the optimal alignment offset can be found through cross-correlation calculation or based on the start-note transient detection, and the track can be shifted accordingly. After alignment, mixing is performed. Mixing is not just a simple addition. A more common approach is to first normalize the amplitude of each track or set the gain according to the target loudness, and then assign weights according to the role of the voice part. For example, the main melody track has a higher weight, and the harmony base track has a lower weight. If necessary, dynamic range control can be applied to each track to avoid peak overload. Finally, the multiple tracks are superimposed to obtain a multi-instrument collaborative synthesized audio signal. In this way, the rhythmic structure can be kept consistent, and the instruments can rise and fall together on the same beat, avoiding dragging or rushing the beat. At the same time, through mixing gain and dynamic control, clipping distortion can be suppressed and the overall loudness utilization rate can be improved, making the final signal both layered and not harsh. For example, when the piano is responsible for harmonic decomposition, the strings are responsible for sustained notes, and the flute is responsible for the main melody, alignment can ensure that all three fall on the same beat in the first beat of each measure; by giving the flute a higher gain and appropriately suppressing the high-frequency energy of the piano, the flute's main melody can be prevented from being masked by the piano's overtones; when it is necessary to achieve an arrangement where the guitar starts first and then the strings enter, a clear entry moment can be achieved on the timeline by setting the track start index, while still maintaining the consistency of the overall beat grid, thus forming a clearer paragraph hierarchy.

[0021] In some examples, it also includes: Acquire and initialize a multi-instrument collaborative synthesis environment running on the MATLAB platform. The environment includes at least an interaction layer, a core algorithm layer, and a data layer. The interaction layer, the core algorithm layer, and the data layer communicate with each other through function handles and callback mechanisms. Set a standard sampling frequency and construct a time vector corresponding to the note duration based on the standard sampling frequency; establish a scale frequency mapping relationship covering a preset range and determine the target fundamental frequency based on the input pitch parameters; The interactive layer outputs the playback, visualization, and export results of the multi-instrument collaboratively synthesized audio signal.

[0022] In some examples, generating the single-instrument synthesized signal of the target instrument based on the target fundamental frequency includes: Dynamic ADSR envelope modulation is applied to the single-instrument synthesized signal; The fundamental frequency component and multiple overtone components are superimposed based on a nonlinear overtone superposition model. Randomized vibrato modulation is introduced into the single instrument synthesized signal and natural texture noise is superimposed.

[0023] For example, the system adopts a layered modular architecture, divided into an interaction layer, a core algorithm layer, and a data layer from top to bottom. The interaction layer employs identity verification and hierarchical permission algorithms, using user information hash matching such as SHA-256 encryption and role-based permission mapping logic to dynamically assign functional permissions to different users. Simultaneously, a tree-structured indexing algorithm and callback function mapping mechanism enable function menu scheduling, supporting one-click switching of function modules and event-triggered responses. The multi-instrument synthesis visualization part uses Short-Time Fourier Transform (STFT) for real-time audio rendering, while the music website navigation uses URL parsing and redirection algorithms, supporting fast link calls and accurate resource retrieval within the MATLAB environment. The core algorithm layer encapsulates a multi-instrument acoustic synthesis model. Based on nonlinear acoustic modeling algorithms, it integrates a scale mapping interpolation algorithm to ensure accurate pitch calibration, an ADSR dynamic envelope modulation algorithm to achieve dynamic volume changes, a nonlinear overtone superposition algorithm to achieve non-integer order overtone combinations, a natural vibrato parameter randomization algorithm to achieve normal distribution random generation of vibrato frequency and depth, and a string texture noise generation and filtering algorithm to achieve Gaussian white noise and first-order IIR filtering. The data layer stores user information based on a user identity encryption storage algorithm, as well as other system-related instrument acoustic parameters and navigation resource links. Data exchange and functional linkage between layers are achieved through MATLAB's App Designer and function handles.

[0024] In some examples, the scale frequency mapping relationship covers a preset range from the low register to the super high register, the pitch parameter adopts the MIDI pitch standard, and the target fundamental frequency is accurately calibrated through the scale frequency mapping relationship.

[0025] In some examples, the dynamic ADSR envelope modulation includes time-domain modulation of the onset period, decay period, duration period and release period, and controls the dynamic change of volume through configurable onset time, decay time, duration, release time and duration amplitude attenuation coefficient. The nonlinear overtone superposition model includes non-integer order overtone terms and executes an overtone supplementation strategy according to the pitch range. In the high-pitched range, preset high-order overtones are supplemented, and preset low-order overtones are supplemented in the low-pitched range to enhance the timbre consistency and realism of different pitch ranges. The natural texture noise is the string texture noise obtained by performing a low-pass filter on white noise. The low-pass filter uses a Butterworth filter and has a configurable cutoff frequency. The amplitude of the natural texture noise meets the preset noise amplitude constraint.

[0026] For example, using four instruments with higher perceived sensitivity—guitar, guzheng, piano, and flute—a nonlinear acoustic model is constructed to simulate timbre, and a standard sampling frequency is defined. The corresponding Nyquist frequency is: This frequency range covers the range of human hearing (20Hz-20kHz) and the highest harmonic frequency of musical instruments (≤15kHz), meeting the requirements of high-fidelity synthesis.

[0027] Based on a 16th note as the basic duration unit Define the time vector of any note duration as: Where N is the duration coefficient, such as N=4 for a quarter note, N=2 for an eighth note, and N takes a corresponding multiple for special durations, with the 16th note as the base duration unit. In the formula defining the time vector of any note duration, K is a simplified abbreviation notation in the mathematical expression. Its function is equivalent to declaring all intermediate terms or terms in the same way, used to simplify writing and avoid listing all the large number of sampling points. In the code... , The time vector length is: Establish a frequency mapping relationship covering the low, mid, high, and super-high registers. To achieve precise calibration of different pitches, among which For pitch parameters, This is a scale frequency mapping function; the scale frequency mapping table covers the C0 to C8 pitch range, with pitch parameters... Adopting the MIDI pitch standard, meeting .

[0028] Simultaneously, by dynamically adjusting the time-domain parameters of onset (0.01s), decay (0.3s), duration (0.3 times the volume), and release (0.6s), the dynamic changes in sound volume are simulated, achieving ADSR envelope modulation. The time-domain expression of the dynamic ADSR envelope modulation model is as follows: in, For the maximum amplitude, The start time of the sound. For decay time, For duration, For the time of pronunciation interpretation, This is the amplitude attenuation coefficient for the continuous segment; Combining audio characteristics and Fourier transform, a superposition model of the fundamental frequency and nonlinear overtones of the 2nd, 3.2nd, 5.7th, and 7.3rd orders is constructed. Simultaneously, 9.1st order (high register) or 1.5th order (low register) overtones are added to adapt the model based on pitch differences, thus reproducing the harmonic characteristics of the instrument. Its expression is as follows: in, For the fundamental frequency, It is a mapping function from the pitch parameter p to the fundamental frequency, and its explicit expression is: And the fundamental frequency of the musical tone corresponding to the pitch The frequency range covers 20Hz~20kHz (the range of human hearing) and is compatible with the high-fidelity synthesis requirements of musical instruments with the highest harmonic frequency ≤15kHz. Corresponding to the fundamental frequency component, Corresponding to overtone components, For overtone scales, supplementing the high register bass range supplement , For overtone amplitude weights, This is the initial phase of the overtone.

[0029] Overtone amplitude weighting satisfy and The weighting coefficients of each overtone were determined through experiments testing the acoustic characteristics of musical instruments.

[0030] Furthermore, randomized vibrato frequency and depth can be introduced, and string-filtering noise accumulated over a test period can be superimposed to enhance the realism of the timbre. The vibrato modulation expression is: in, For vibrato depth, Randomized vibrato frequency; superimposed string-filtered noise. The final audio signal is obtained: Among them, string filtering noise The white noise is filtered by a Butterworth low-pass filter, and the filter cutoff frequency is... The noise amplitude satisfies .

[0031] By normalizing the amplitude, the audio peak value is controlled at 0.7 times the full scale to avoid signal distortion, while supporting the collaborative splicing and mixing of multi-instrument audio tracks.

[0032] In some examples, the interaction layer includes identity verification and permission classification based on a user information database before performing the multi-instrument collaborative synthesis, and switches between multi-instrument synthesis, audio playback, audio export and music resource navigation functions through a tree-shaped function menu. The music resource navigation is implemented by jumping to the encapsulated music resource website links through MATLAB web calls.

[0033] For example, identity verification can be implemented based on a user information database, supporting hierarchical permission management such as for ordinary users and administrators, ensuring system data security. Simultaneously, to create a more operable system structure, a tree-shaped function menu is introduced, enabling one-click switching between multi-instrument synthesis, audio playback / export, and navigation modules. URLs of commonly used music resource websites are encapsulated, supporting quick navigation within the MATLAB environment.

[0034] A multi-dimensional permission matrix can be constructed using the SHA-256 hash encryption algorithm and the Role-Based Access Control (RBAC) model. User passwords are entered using the formula: Irreversible encryption is performed, where PW is the original password and S is a random salt value. An access control vector is generated by combining user roles, and the call interface is only open to authorized functional modules to ensure the security of system data and algorithms.

[0035] A user behavior-driven priority allocation algorithm can be introduced, using the formula... in, , Frequency of function operation Calculate the weight of each functional module for scene adaptability, dynamically adjust the menu sorting, realize the rapid access to high-frequency functions and scene-specific needs, and improve interaction efficiency.

[0036] Therefore, the system employs nonlinear overtone superposition, dynamic ADSR envelope, and a natural texture noise model to accurately reproduce the acoustic characteristics of traditional musical instruments. The timbre realism of its synthesized audio is significantly superior to traditional sine wave superposition algorithms, meeting the needs of professional music composition and acoustic research. For the first time, it integrates authentication, multi-instrument synthesis, and resource navigation functions on the MATLAB platform, breaking through the bottleneck of fragmented functions in traditional tools and constructing a complete workflow environment for algorithm verification, music creation, and resource retrieval. Based on App Designer... The system features a visual interactive interface, lowering the operational threshold for professional algorithms while maintaining accessibility to the underlying code. It caters to diverse needs across various fields, including public creation, academic research, and industrial applications. The modular algorithm architecture supports rapid integration of new instrument acoustic models, and standardized function interfaces provide a convenient experimental platform for iterative verification of audio algorithms. The scenario-adaptive parameter library can be continuously expanded based on emerging field needs. The system is adaptable to multiple fields such as immersive media, intelligent music education, academic research, accessibility services, and industrial testing. Through customized parameter adjustments and multi-channel audio output, it meets the personalized audio synthesis needs of different scenarios, expanding the system's practical value and application boundaries. It provides efficient tools for digital music creation and industrial audio testing, empowers public welfare sectors such as intelligent education and accessibility services, and offers customized audio solutions for special groups such as the hearing impaired and the elderly, combining technological and social value.

[0037] In some cases, considering that the spectra of different instruments may continuously overlap in several key frequency bands under the same harmonic progression, reducing the intelligibility of the melody or key voices, this phenomenon of spectral masking is related to the human ear's critical band. Therefore, short-time spectral analysis can be performed on each track on the time axis of a unified beat grid. The spectral energy is divided according to the critical band or equivalent bandwidth, and the energy overlap between the melody track and the accompaniment track in the target frequency band is calculated. When the overlap exceeds a threshold, a dynamic attenuation that smoothly changes over time is applied to the accompaniment track in that frequency band, while a limited transient enhancement or slight harmonic emphasis is applied to the melody track within the same beat window. This achieves a synergy of yielding and prominence without significantly altering the arrangement structure. Based on this, in some examples, the process of generating multi-instrument collaborative synthesized audio signals also includes: performing short-time spectrum analysis on the audio signals of each target instrument, which includes at least a melody track and an accompaniment track, under a unified sampling frequency and a unified time reference; estimating the energy overlap of the melody track and the accompaniment track in the target frequency band based on the short-time spectrum analysis results according to the critical band division rule; and when the energy overlap exceeds a preset threshold, applying band-selective dynamic attenuation aligned with the beat grid to the accompaniment track in the target frequency band and applying transient enhancement aligned with the beat grid to the melody track, so as to reduce spectrum masking and improve the clarity of the melody.

[0038] For example, in MATLAB, the analysis window length and step size can be determined first to balance time resolution and frequency resolution. For instance, a window length matching the beat resolution can be selected and the short-time Fourier transform can be calculated by moving the overlapping window. After obtaining the amplitude spectrum of each frame, the energy is calculated by grouping according to the preset critical bands and forming an energy vector. Then, the masking index of each critical band is calculated with the main melody track as a reference. For example, the masking intensity can be estimated by the ratio of accompaniment energy to main melody energy or by the overlap integral. When the masking intensity exceeds the threshold, a gain function is applied to the accompaniment track in the corresponding frequency band. The gain function adopts a smooth envelope with the beat window to avoid abruptness. At the same time, a transient enhancement module can be introduced into the starting section of the main melody track. For example, a short-time high-frequency boost can be performed on the starting section or the starting envelope can be slightly steepened. And ensure that the processing only enters and exits gradually within the beat window. As a result, the clarity of the main melody is significantly improved, especially when piano arpeggios and string undertones are present simultaneously. The flute or violin main melody is less likely to be drowned out by mid- and high-frequency overtones, while the presence of the accompaniment is still maintained because its attenuation is frequency band selective rather than overall suppression. For example, when arranged with a flute main melody, piano arpeggios, and string sustained notes undertones, masking adaptive mixing will identify the energy overlap of the piano near the critical band of the main melody and slightly attenuate that frequency band of the piano during the duration of the main melody notes, thus making the flute's main frequency and key overtones more prominent. Another example is when guitar strumming is superimposed with vocals or the main melody, narrow-band attenuation can be applied to sections where high-frequency noise from the strumming conflicts with high-frequency details in the main melody, so that the sibilance and details of the main melody are not masked.

[0039] In some cases, when multiple instrument tracks use fixed and similar initial phases at the same fundamental frequency and low-order overtones, the mixing process can result in alternating phase superposition and phase cancellation. This manifests as periodic fluctuations in energy intensity, low-frequency floating, or sudden thinning of certain notes. This problem is not noticeable on a single track but is amplified when multiple tracks are superimposed. Therefore, the initial phases of each overtone of each instrument can be considered controllable parameters. Controlled decorrelation can be introduced at the cross-instrument level for overtones of the same order, preventing statistically significant long-term in-phase or strong anti-correlation between different instruments. Simultaneously, to avoid disrupting the impact of the onset, phase changes in the onset segment need to be protected, ensuring that phase disturbances primarily affect the sustain and release segments. Based on this, in some examples, when generating single-instrument synthesis signals for at least two target instruments, the method further includes: establishing an initial phase configuration vector associated with the register for each target instrument, and applying reproducible initial phase bias and initial phase micro-drift constraints updated at the measure scale to the same-order overtone components of different target instruments, so that different target instruments meet controlled phase decorrelation conditions at the same-order overtones, thereby reducing energy fluctuations caused by phase in-phase or anti-correlation during multitrack mixing.

[0040] For example, an initial phase configuration vector can be defined for each instrument in MATLAB. This vector can depend on the register, playing style, and envelope velocity, and different initial phase values ​​can be used when generating the fundamental frequency and overtone components. For instance, a reproducible phase offset can be generated using a deterministic pseudo-random sequence, and a small phase drift term updated on a measure scale can be superimposed to avoid a fixed phase relationship for a long time. At the same time, a phase freeze strategy is used near the onset time to ensure the consistency of the onset transient and prevent the impact from being diffused. At the cross-instrument level, phase spacing constraints can be set for overtones of the same order, such as avoiding regions where the phase difference is close to zero or close to half a cycle, so that the superposition is more uniform. The technical effect is improved energy stability after multi-track mixing, more solid low frequencies without periodic voids, and no obvious pitch jitter, because the phase perturbation only changes the waveform superposition shape and does not change the fundamental frequency itself. For example, when the bass and cello play the same low-frequency root note simultaneously, if the initial phases of their low-order overtones are too similar, the resulting mixture may exhibit excessively strong low frequencies at certain times and thin low frequencies at other times. The method described above can significantly suppress this fluctuation, making the low frequencies consistently stable. Another example is when two pianos or a piano and a guitar play together on the same chord. Initial phase decorrelation can reduce phase cancellation in the same frequency band, making the chord thickness more stable and preventing a fluctuating sound.

[0041] In some cases, even if the pitch and rhythm are perfectly correct, the synthesized sound can still be mechanical. This is often not due to timbre but rather the lack of micro-timing and dynamics. In real performances, the onset may be slightly ahead or behind, the dynamics may change slowly within a phrase, and the vibrato may change with emotion and direction. Therefore, an expression curve library can be introduced. This library is not a simple random perturbation but a parametric curve derived from performance statistics or regularized expressions. It can provide controllable scheduling for the starting sampling point of each note, the subtle changes in energy during the envelope duration, and the gradual changes in vibrato depth and rate without disrupting the rhythmic grid. Furthermore, cross-measure continuity constraints ensure that musical phrases do not exhibit abrupt jumps. Based on this, in some examples, performing multitrack time axis alignment on the audio signals of each target instrument includes: while keeping the beat structure unchanged, calling a preset expression curve library, which at least includes the onset advance distribution, the energy micro-change rate of the duration, and the energy recovery curve at the end of the release segment, and performing micro-temporal offset on the starting sampling point of each note based on the expression curve library and applying a smoothing function that changes slowly with time to the envelope duration and vibrato modulation depth, so as to generate an expression-based multitrack alignment result with cross-measure continuity.

[0042] For example, curve entries can be created in MATLAB, categorized by style and instrument. These entries could include patterns such as the distribution of the pre-start amount within each beat, the slight lag in weak beats and the slight pre-start in strong beats, the slight upward or downward curves of energy in sustained sections, and the energy recovery curve at the end of a phrase. When aligning multiple tracks on the time axis, the nominal pre-start position is first determined using the beat as a framework. Then, the pre-start sampling points are fine-tuned based on the curve library, and a smoothing function is applied to the envelope parameters. Furthermore, a continuity constraint is set for the offset of adjacent notes within the same musical phrase to prevent a single note from abruptly advancing too much and disrupting the rhythm. As a result, the beat remains strictly controllable, but the listening experience has a sense of push and pull and a feeling of breathing. The main melody sounds more like a human performance, and the accompaniment has more rhythmic elasticity, especially with a significant improvement in slow, lyrical passages. For example, when the strings provide a long note and the flute plays the melody, the expression curve library can be used to allow the flute to slightly advance on the strong beat and slightly decrease on the weak beat, and to gradually increase the depth of the vibrato in the middle of the phrase and decrease it at the end of the phrase, thus forming a natural phrasing. Another example is in piano arpeggios, where the curve library can be used to make the dynamics of each arpeggio fluctuate slightly, simulating the natural differences in the touch of real fingers on the keys, avoiding the typewriter feel caused by every note being exactly the same.

[0043] In some cases, if the same instrument model uses the same set of overtone coefficients across the entire range, the high frequencies may sound thin, harsh, or lack presence, while the low frequencies may sound muddy or severely masked. When multiple instruments are layered, these defects are amplified, leading to a collapse of the overall balance. Therefore, the pitch range can be treated as a conditional variable, and an overtone energy redistribution driven by audibility indices can be introduced to make the spectral structure of different ranges more consistent with human hearing and the physical laws of instruments. Specifically, a set of target audibility indices is defined for each range and mapped to the target energy envelope of each order of overtones. Then, under amplitude constraints, the overtone coefficients are updated by minimizing the cost function, and temporal smoothing is applied to make the transitions between notes natural. Based on this, in some examples, the process of determining the target fundamental frequency based on the input pitch parameters and generating a single instrument synthesized signal further includes: identifying the octave range corresponding to the pitch parameters, and calling a preset set of auditory perception indicators for the octave range. The set of auditory perception indicators includes at least a clarity indicator, a fullness indicator, and a brightness indicator. A cost function is constructed based on the set of auditory perception indicators to characterize the target energy envelope of each order of overtones. Under the condition of satisfying the amplitude constraints of each order of overtones, the overtone coefficients are optimized and updated to achieve a range-aware redistribution of overtone energy and enhance the timbre consistency of different ranges.

[0044] For example, in MATLAB, the octave range number can be calculated based on the pitch, and the target indicators corresponding to that range, such as the weights of clarity, fullness, and brightness, can be read. These weights can then be converted into a target energy envelope. For example, in the high range, some higher-order overtones can be increased while limiting sharp peaks, and in the low range, the fundamental frequency and low-order overtones can be emphasized while suppressing excessively strong higher-order overtones to avoid muddiness. Then, the difference between the spectrum generated by the current coefficients and the target envelope can be measured using a cost function. The overtone coefficients can be updated using a constrained optimization method. Finally, the coefficients can be used to synthesize the note and the coefficients can be interpolated and smoothed between adjacent notes. As a result, the timbre is more consistent in different ranges, and the synthesized melody will not suddenly become thin or muddy when it crosses a large range. Multiple instruments can more easily maintain layering without swallowing each other up in the same harmony. For example, if the violin's high-frequency harmonics are not properly enhanced, they will sound weak and lack penetration, while excessive enhancement will be harsh. The above method, through target envelope constraint, can make the high-frequency range present without being too shrill. As another example, if the cello and bass's high-frequency harmonics are too strong in the low-frequency range, it will cause muddiness and mask other instruments. This method can suppress high-frequency components that do not consider the listening experience, making the bass cleaner and leaving space for the piano's left hand and low-frequency percussion.

[0045] In some cases, manual parameter tuning struggles to reliably reproduce a reference timbre, especially when maintaining consistency across multiple pieces or sections. Manual adjustments are not only time-consuming but also non-repeatable. Therefore, a parameterized single-instrument synthesis chain can be implemented, defining a multi-index loss function that characterizes timbre and dynamics. This allows for the inverse calculation of parameters to obtain a set of parameters that approximates the reference timbre segment in the feature space. To make optimization feasible, modules such as oscillators, envelopes, filters, and nonlinear shaping need to be expressed in differentiable or approximately differentiable forms, giving parameter updates a clear direction. Based on this, in some examples, before generating a single-instrument synthesized signal for at least two target instruments, the method further includes: acquiring a target reference timbre fragment and calculating the short-time spectral features, cepstral features, transient features, and loudness curve features of the target reference timbre fragment; constructing a multi-index loss function based on the short-time spectral features, the cepstral features, the transient features, and the loudness curve features; configuring the single-instrument synthesis chain as a differentiable parameterizable computation graph; and performing iterative inverse optimization on at least some acoustic parameters in the single-instrument synthesis chain according to the multi-index loss function to obtain an instrument acoustic parameter configuration that matches the target reference timbre fragment.

[0046] For example, a target reference timbre fragment can be read in MATLAB and multi-perspective features can be calculated. These include short-time amplitude spectra to describe the spectral envelope, cepstral features to describe the formant structure, transient indices to describe attack sharpness and energy concentration, and loudness curves to describe dynamic changes over time. A weighted loss function is then constructed, and key parameters in the synthesis chain are set as variables to be optimized. These include parameters such as overtone amplitude decay curve parameters, attack and decay times of the envelope, filter cutoff frequency and order, vibrato depth and rate, and noise texture weights. An iterative optimization method is used to update the parameters within a finite number of steps, and the obtained parameters are stored in an instrument configuration library for subsequent multitrack synthesis and reuse. This allows for the rapid acquisition of parameter configurations that are highly consistent with the reference timbre and reproducible. In multi-instrument arrangements, calibrated timbres can be established for each voice part, thus avoiding excessive differences in the listening experience of the same instrument in different sections, while significantly reducing the workload of manual parameter tuning. For example, when you want the synthesized piano to resemble the brightness and striking feel of a studio piano, you can use that recording clip as a reference and reverse engineer it to obtain a more suitable attack envelope and high-frequency attenuation parameters; when you want the synthesized flute to resemble a performer's breath texture and vibrato habits, you can reverse engineer it to obtain a more suitable noise ratio and vibrato gradient parameters; when multiple instruments are used in collaboration, this reference-based calibration can also avoid the overall style inconsistency caused by manually adjusting parameters on different tracks.

[0047] Please see Figure 2 One embodiment of the multi-instrument collaborative synthesis device based on MATLAB in this application includes: The determining unit 21 is used to set a standard sampling frequency, construct a time vector corresponding to the duration of a note based on the standard sampling frequency, establish a scale frequency mapping relationship covering a preset range, and determine the target fundamental frequency based on the input pitch parameters. Processing unit 22 is used to call the acoustic parameter configuration corresponding to the target instrument for at least two target instruments respectively, and generate a single instrument synthesized signal of the target instrument based on the target fundamental frequency; Synthesis unit 23 is used to perform multi-track time axis alignment on the audio signals of each of the target instruments, and to mix the audio signals of each of the target instruments based on the alignment results to generate a multi-instrument collaborative synthesized audio signal.

[0048] like Figure 3 As shown, this application embodiment also provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor. When the processor 320 executes the computer program 311, it implements the steps of any of the above-described methods of the multi-instrument collaborative synthesis method based on MATLAB.

[0049] Since the electronic device described in this embodiment is the device used to implement a MATLAB-based multi-instrument collaborative synthesis device in this application embodiment, those skilled in the art can understand the specific implementation method and various variations of the electronic device in this embodiment based on the method described in this application embodiment. Therefore, how the electronic device implements the method in this application embodiment will not be described in detail here. Any device used by those skilled in the art to implement the method in this application embodiment is within the scope of protection of this application.

[0050] In practical implementation, when the computer program 311 is executed by the processor, it can achieve the following: Figure 1 Any of the corresponding implementation methods in the embodiments.

Claims

1. A multi-instrument collaborative synthesis method based on MATLAB, characterized in that, include: Set a standard sampling frequency, construct a time vector corresponding to the note duration based on the standard sampling frequency, establish a scale frequency mapping relationship covering a preset range, and determine the target fundamental frequency based on the input pitch parameters; For at least two target musical instruments, the acoustic parameter configurations corresponding to the target musical instruments are called respectively, and a single-instrument synthesized signal of the target musical instrument is generated based on the target fundamental frequency; Multitrack time axis alignment is performed on the audio signals of each of the target instruments, and the audio signals of each of the target instruments are mixed based on the alignment results to generate a multi-instrument collaborative synthesized audio signal.

2. The method as described in claim 1, characterized in that, The process of generating the single-instrument synthesized signal of the target instrument based on the target fundamental frequency includes: Dynamic ADSR envelope modulation is applied to the single-instrument synthesized signal; The fundamental frequency component and multiple overtone components are superimposed based on a nonlinear overtone superposition model. Randomized vibrato modulation is introduced into the single instrument synthesized signal and natural texture noise is superimposed.

3. The method as described in claim 1, characterized in that, Before performing multi-track time axis alignment on the audio signals of each of the target instruments and mixing the audio signals of each of the target instruments based on the alignment results to generate a multi-instrument collaborative synthesized audio signal, the method further includes: Amplitude normalization is performed on the single-instrument synthesized signals of each of the target instruments to suppress distortion.

4. The method as described in claim 1, characterized in that, Also includes: Acquire and initialize a multi-instrument collaborative synthesis environment running on the MATLAB platform. The environment includes at least an interaction layer, a core algorithm layer, and a data layer. The interaction layer, the core algorithm layer, and the data layer communicate with each other through function handles and callback mechanisms. Set a standard sampling frequency and construct a time vector corresponding to the note duration based on the standard sampling frequency; establish a scale frequency mapping relationship covering a preset range and determine the target fundamental frequency based on the input pitch parameters; The interactive layer outputs the playback, visualization, and export results of the multi-instrument collaboratively synthesized audio signal.

5. The method as described in claim 1, characterized in that, The scale frequency mapping relationship covers a preset range from the low register to the super high register. The pitch parameter adopts the MIDI pitch standard, and the target fundamental frequency is accurately calibrated through the scale frequency mapping relationship.

6. The method as described in claim 2, characterized in that, The dynamic ADSR envelope modulation includes time-domain modulation of the onset period, decay period, duration period and release period, and controls the dynamic change of volume through configurable onset time, decay time, duration, release time and amplitude attenuation coefficient of the duration period. The nonlinear overtone superposition model includes non-integer order overtone terms and executes an overtone supplementation strategy according to the pitch range. In the high-pitched range, preset high-order overtones are supplemented, and preset low-order overtones are supplemented in the low-pitched range to enhance the timbre consistency and realism of different pitch ranges. The natural texture noise is the string texture noise obtained by performing a low-pass filter on white noise. The low-pass filter uses a Butterworth filter and has a configurable cutoff frequency. The amplitude of the natural texture noise meets the preset noise amplitude constraint.

7. The method as described in claim 4, characterized in that, Before performing the multi-instrument collaborative synthesis, the interaction layer also includes identity verification and permission classification processing based on the user information database, and realizes the switching of multi-instrument synthesis, audio playback, audio export and music resource navigation functions through tree-shaped function menu scheduling. The music resource navigation realizes the jump to the encapsulated music resource website link through MATLAB web call.

8. A multi-instrument collaborative synthesis device based on MATLAB, characterized in that, include: The determining unit is used to set a standard sampling frequency, construct a time vector corresponding to the note duration based on the standard sampling frequency, establish a scale frequency mapping relationship covering a preset range, and determine the target fundamental frequency based on the input pitch parameters; The processing unit is configured to, for at least two target musical instruments, respectively call the acoustic parameter configuration corresponding to the target musical instrument, and generate a single-instrument synthesized signal of the target musical instrument based on the target fundamental frequency; A synthesis unit is used to perform multi-track time axis alignment on the audio signals of each of the target instruments, and to mix the audio signals of each of the target instruments based on the alignment results to generate a multi-instrument collaborative synthesized audio signal.

9. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the computer program stored in the memory, implements the steps of the MATLAB-based multi-instrument collaborative synthesis method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the MATLAB-based multi-instrument collaborative synthesis method as described in any one of claims 1-7.