Nanyin instrumental analysis and music generation processing method and system

By constructing a set of Nanyin triplet fingerprints and mapping user intent, and combining resonator coefficients and excitation primitive sequences, the problems of insufficient reproduction of charm and texture and poor cross-sound source adaptability in Nanyin generation are solved, realizing high-fidelity generation of Nanyin audio and harmonious fusion of multiple sound sources.

CN121600889APending Publication Date: 2026-03-03FUJIAN BEITELI INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511785442.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-01
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing Nanyin generation technology cannot accurately reproduce the charm and natural texture, especially in the restoration of glissando transitions and breath resonance dynamics. Furthermore, when generating across sound sources, the adaptability of sound source characteristics is poor, resulting in a disconnect between pitch and timbre and unnatural generated audio.

Method used

By acquiring the harmonic ratio, formant shift, and transient excitation profile of the original Nanyin samples, an interpolable triplet fingerprint set is constructed, which maps the user's intent as the target fingerprint trajectory. The pitch and timbre are strongly linked through the resonator coefficient and excitation primitive sequence. The high-fidelity Nanyin audio is generated by combining fingerprint residual reports and human auditory feedback.

Benefits of technology

It achieves accurate reproduction of Nanyin pitch and timbre with strong linkage, avoids the disconnect between pitch and timbre, improves the natural texture of Nanyin generation and the harmonious integration effect of cross-sound source generation, and enhances the realism and flexibility of the generated audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600889A_ABST
    Figure CN121600889A_ABST
Patent Text Reader

Abstract

The invention discloses a Nanyin instrumental analysis and music generation processing method and system, and belongs to the technical field of audio signal processing and computer music, and the method comprises the steps: 1, obtaining an original Nanyin sample, and measuring and recording a harmonic spectrum ratio curve, a formant displacement curve and a transient excitation contour in a short frame, forming a triple fingerprint set capable of being expressed by interpolation; step 2, according to the intention of a user, mapping intention primitives into curve offset of the triple fingerprint in the step 1, and generating a target fingerprint track; step 3, mapping the target fingerprint track in the step 2 into a superposable resonator coefficient sequence and an excitation element coefficient sequence; and step 4, according to the resonator coefficient sequence and the excitation element coefficient sequence obtained through mapping in the step 3. According to the invention, the degree of freedom of creation according to the intention of a user can be met, the cavity body feeling of linkage between the Nanyin pitch and the timbre intensity can be accurately reproduced, and the unnatural problem caused by disjunction of the pitch and the timbre is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of audio signal processing and computer music technology, and more specifically, to a method and system for analyzing and generating Nanyin musical instruments. Background Technology

[0002] With the application and promotion of digital technology in the field of traditional music, the digital preservation and regeneration of Nanyin, one of the oldest existing music genres in China, has become a key focus of the industry. Existing technologies can extract macroscopic features such as the melody outline and rhythmic pattern of Nanyin music through audio analysis and synthesis algorithms, and generate audio signals with the basic musical form of Nanyin based on these features, meeting basic needs such as playing simple Nanyin excerpts and teaching demonstrations. For example, some existing solutions collect the pitch sequence and beat information of the original Nanyin audio, build a melody database, and then combine it with a preset Nanyin instrument timbre sample library to achieve melody and timbre matching and synthesis, completing the basic Nanyin audio generation.

[0003] However, existing Nanyin generation technologies fall short in restoring the charm and natural texture of Nanyin, making it difficult to reproduce its unique artistic expression. On the one hand, the charm of traditional Nanyin relies heavily on millisecond-level microscopic acoustic details. Taking a typical female Nanyin performance as an example, the vocal process involves a gentle glissando transition from a slightly lower pitch to the target pitch, accompanied by resonance vibrations caused by changes in the shape of the oral cavity and subtle friction noise generated by the flow of breath. These dynamic features on a scale of tens of milliseconds together constitute the unique vocal resonance and sense of breath in Nanyin. However, existing technologies generally use a synthesis architecture that combines pitch curves with fixed timbre templates, simulating glissando only by linearly stretching or compressing pitch curves, resulting in glissando exhibiting a mechanical, stretched electronic sound quality. At the same time, breath noise is mostly achieved by superimposing fixed noise samples, which cannot be dynamically matched with pitch changes and cavity resonance, forming blocky and separated noise signals, severely damaging the natural texture of Nanyin.

[0004] On the other hand, existing technologies suffer from poor adaptability in generating Nanyin music across different sound sources, such as human voice, flute, and erhu. The acoustic nature of different Nanyin sound sources differs significantly: human voice relies on the resonance adjustment of the oral cavity and pharyngeal cavity, and its resonant frequency shifts synchronously with pitch; flute relies on the resonance of airflow in the tube cavity, and its acoustic characteristics are determined by airflow speed and the opening and closing state of the tube cavity; erhu relies on the coupling of string vibration and resonator, and its timbre is dynamically affected by bow pressure and bow speed. However, when generating Nanyin music across sound sources, existing technologies simply reuse the same set of melody parameters without adjusting the acoustic parameters according to the differences in resonance mechanisms and excitation methods of different sound sources. This results in a stiff connection between the generated audio from different sound sources, presenting a noticeable splicing effect, and failing to reproduce the harmonious integration of various parts in a Nanyin ensemble.

[0005] In-depth analysis reveals that the shortcomings of existing technologies stem from insufficient understanding and technical implementation of the strong linkage between pitch and timbre in Nanyin. The essential acoustic characteristic of Nanyin is the dynamic coupling of pitch and timbre. During pitch changes, the shape of the vocal cavity or instrument resonance structure adjusts synchronously, leading to real-time changes in formant frequency and harmonic energy distribution. However, existing Nanyin generation systems generally employ a fragmented architecture between pitch control and timbre synthesis modules: the pitch control module only outputs pitch timing information, while the timbre synthesis module generates audio based on fixed filter parameters or timbre samples, with no dynamic linkage mechanism between the two. Even when some solutions introduce formant adjustment, real-time synchronization between formant and pitch changes is not achieved, resulting in timbre filter parameters failing to match the dynamic adjustment requirements of the vocal cavity during pitch changes. Ultimately, this fails to balance the creative freedom of Nanyin generation with the reproduction of vocal cavity texture, making it difficult to meet the high-fidelity artistic quality requirements of Nanyin digital inheritance and limiting the technical support capabilities for innovative Nanyin creation. Summary of the Invention

[0006] To address the problems existing in the prior art, the present invention aims to provide a method and system for analyzing and generating Nanyin instrumental music, which can satisfy the user's freedom to create according to their intentions, and accurately reproduce the cavity texture of Nanyin with strong linkage between pitch and timbre, avoiding the unnatural problem caused by the disconnect between pitch and timbre.

[0007] To solve the above problems, the present invention adopts the following technical solution.

[0008] Firstly, a method for analyzing and generating music from Nanyin instruments includes:

[0009] Step 1: Obtain the original Nanyin sample, and measure and record the harmonic ratio curve, formant displacement curve and transient excitation profile within a short frame to form a set of triplet fingerprints that can be interpolated.

[0010] Step 2: Based on the user's intent, map the intent primitives to a curve offset of the triplet fingerprint in Step 1 to generate the target fingerprint trajectory.

[0011] Step 3: Map the target fingerprint trajectory from Step 2 into a superimposed sequence of resonator coefficients and a sequence of excitation element coefficients;

[0012] Step 4: Based on the resonator coefficient sequence and excitation element coefficient sequence obtained in Step 3, establish a response mapping relationship with the triplet fingerprint in Step 1, and solve the pre-tuning control sequence in reverse to perform pre-morphological shaping on the transient excitation.

[0013] Step 5: Combine the pre-tuned control sequence from Step 4 with the target fingerprint trajectory from Step 2 in a hierarchical manner to construct a control package that aligns the musical phrase sequence with time, including the resonator coefficient time sequence, excitation element markers, and pre-tuned offset.

[0014] Step 6: Update the resonator coefficients and excitation elements sequentially according to the control package in chronological order, and perform phase synchronization correction. At the same time, perform a short-term similarity comparison with the fingerprint from Step 1 to generate an audio stream and fingerprint residual report.

[0015] Step 7: Based on the fingerprint residual report and human auditory feedback from Step 6, generate a set of calibration parameters and migration rules, and apply them to the adjustment of control parameters across sound sources.

[0016] Furthermore, including:

[0017] Step 21: Parse the user intent into a normalized parameter set consisting of multiple intent primitives, and form an intent timeline by time index;

[0018] Step 22: Based on the intent timeline, select template samples corresponding to the intent primitives from the predefined primitive fingerprint templates, and distort and scale the template samples in terms of time scale and amplitude to generate local target fingerprint segments.

[0019] Step 23: Calculate the cavity coupling-related compensation coefficients for the local target fingerprint segment, and generate the time-aligned compensation offset sequence of the corresponding harmonic ratio, resonant peak displacement and transient profile.

[0020] Step 24: Adhere the local target fingerprint segments in chronological order and superimpose the compensation offset sequence onto the corresponding channel to perform time consistency processing, thereby obtaining a continuous interpolable target fingerprint trajectory.

[0021] Furthermore, including:

[0022] Step 31: Divide the target fingerprint trajectory from Step 2 into multiple analysis units according to the time window, and calibrate the harmonic ratio, resonant peak displacement, and local extreme points, slope direction, and transient peak value of the transient profile within each unit.

[0023] Step 32: Take the fingerprint data in each analysis unit divided in Step 31 as input, select the matching elementary resonator template, and calculate the corresponding resonator center frequency, quality factor and gain curve coefficient, and generate the corresponding excitation elementary parameters.

[0024] Step 33: Integrate the resonator coefficient sequence generated in step 32 with the excitation element parameter sequence in sequence, and perform continuous processing on the center frequency, gain and excitation amplitude, while adjusting the transient trigger parameters.

[0025] Step 34: Combine the resonator coefficient sequence and the excitation element parameter sequence obtained in step 33 to form the final target driving sequence, and record the superposition relationship and time correspondence information of each coefficient and excitation parameter.

[0026] Furthermore, including:

[0027] Step 41: Apply the continuous resonator coefficient sequence from step 3 to the triplet fingerprint from step 1, record the response shift of the harmonic spectrum ratio, resonance peak shift, and transient profile within each time window, and derive the response delay of each time window to form a local response mapping.

[0028] Step 42: Based on the local response mapping and delay information from step 41, the pre-tuning control sequence for each time window is calculated in reverse, the driving offset is applied to the resonator coefficient sequence, and a short-term excitation is applied at the trigger position of the transient profile to form the pre-form shaping.

[0029] Furthermore, including:

[0030] Step 51: The pre-tuned control sequence of Step 4 and the target fingerprint trajectory of Step 2 are indexed in layers according to beat layer, melody layer, cavity layer and transient layer, and the time points of each layer are aligned, while key nodes are marked.

[0031] Step 52: Based on the time alignment in step 51, the resonator coefficient sequence and excitation element sequence in adjacent time windows are locally weighted and superimposed, and the parameters of each layer are combined in the same time step to form a multi-layer control package, which includes the resonator coefficient sequence, excitation element marker and pre-adjustment offset.

[0032] Step 53: Perform time series integrity verification on the control package generated in step 52, mark local high-variance segments and boundary segments, and output a hierarchical fused control package set with optimization markings.

[0033] Furthermore, including:

[0034] Step 61: The time stamp and superposition coefficient sequence in the control package generated in step 5 are used to incrementally update the resonator coefficients and excitation elements in chronological order, and the excitation elements are dynamically activated or deactivated according to the trigger mark and amplitude change to form a local incremental trigger sequence.

[0035] Step 62: Calculate the phase shift of the resonator at each time step based on the incremental sequence output in step 61, and adjust the triggering phase of the transient primitive. At the same time, apply piecewise smooth interpolation in the coefficient transition section to maintain continuity.

[0036] Step 63: Perform a short-time comparison between the cavity state data output in step 62 and the triplet fingerprint in step 1 according to the time window, and extract the residual vectors of harmonic ratio, resonant peak shift and transient profile.

[0037] Step 64: The resonator outputs of each time step in step 63 are accumulated and superimposed to form a continuous audio stream, and the residual vector is organized into a fingerprint residual report arranged in time sequence. At the same time, the audio stream and the residual report are output.

[0038] Furthermore, including:

[0039] Step 71: Receive the fingerprint residual report and human auditory feedback from Step 6, divide the harmonic ratio, resonant peak shift and transient profile deviation in the residuals into time windows, map them to subjective scores, generate weighted calibration indexes, and analyze the relationship between residuals of each channel and subjective preferences to determine the adjustment order of cavity state and transient characteristics.

[0040] Step 72: Based on the weighted calibration index output in step 71, calculate the calibration offset sequence of each resonator coefficient and excitation element, and establish a parameter mapping table for different sound source categories. Through proportional adjustment, offset correction and time window synchronization, generate calibration coefficients and cross-sound source migration rules.

[0041] Step 73: Integrate the calibration coefficients and migration rules generated in step 72 into a time-aligned multi-channel calibration parameter set, mark the applicable sound source categories and time synchronization information, and form an integrated parameter set that can be directly used to control sequence adjustment.

[0042] Furthermore, compensation coefficients related to cavity coupling are calculated for local target fingerprint segments, including:

[0043] Step 231: Divide the local target fingerprint segment into short-time windows, measure and record the transient energy change rate, frequency drift rate and local amplitude deviation of each window, and construct the cavity response offset matrix for each time window accordingly.

[0044] Step 232: The cavity response offset matrix from step 231 is decomposed into frequency, amplitude, and transient dimensions to identify various coupling offsets. Compensation weights are assigned based on their influence intensity and duration in the target fingerprint segment to generate a compensation coefficient sequence for correcting amplitude, frequency, and transient triggering.

[0045] Furthermore, the corresponding resonator center frequency, quality factor, and gain curve coefficients are calculated, and the corresponding excitation element parameters are generated, including:

[0046] Step 321: Perform differential and curvature analysis on the harmonic ratio, resonant peak displacement and transient profile of the target fingerprint segment divided by time window, extract local extreme points, slope changes and frequency drift trends, and establish a local frequency response characteristic table for each time window.

[0047] Step 322: Map the local frequency response feature table generated in step 321 to the resonator center frequency, quality factor and gain curve coefficients for each time window, and mark the transient trigger point;

[0048] Step 323: Based on the dynamic resonator coefficient sequence and transient trigger marker from step 322, perform spatiotemporal optimization on the amplitude, trigger time, and duration of the excitation element, so that the excitation element and the resonator coefficient are aligned in time and form the final parameter sequence.

[0049] Secondly, a Nanyin instrumental music analysis and music generation processing system includes:

[0050] The fingerprint acquisition module is used to acquire the original Nanyin sample and measure and record the harmonic ratio curve, formant displacement curve and transient excitation profile within a short frame to form a set of triplet fingerprints that can be interpolated.

[0051] The trajectory generation module is used to map the intent primitives to curve offsets of the triplet fingerprint in the fingerprint acquisition module according to the user's intent, thereby generating the target fingerprint trajectory.

[0052] The coefficient mapping module is used to map the target fingerprint trajectory in the trajectory generation module into a superimposed resonator coefficient sequence and an excitation element coefficient sequence.

[0053] The pre-tuning solution module is used to obtain the resonator coefficient sequence and excitation element coefficient sequence according to the mapping in the coefficient mapping module, establish a response mapping relationship with the triplet fingerprint in the fingerprint acquisition module, and solve the pre-tuning control sequence in reverse to perform pre-shape shaping of transient excitation;

[0054] The control package construction module is used to hierarchically combine the pre-tuning control sequence in the pre-tuning solution module with the target fingerprint trajectory in the trajectory generation module to construct a control package that is time-aligned with the musical phrase sequence, including the resonator coefficient time series, excitation element markers and pre-tuning offsets.

[0055] The audio generation module is used to update the resonator coefficients and excitation elements sequentially according to the control package in time order, and perform phase synchronization correction. At the same time, it performs a short-term similarity comparison with the fingerprint in the fingerprint acquisition module to generate an audio stream and a fingerprint residual report.

[0056] The calibration migration module generates a set of calibration parameters and migration rules based on the fingerprint residual report from the audio generation module and human auditory feedback, and applies them to the adjustment of control parameters across sound sources.

[0057] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0058] (1) This scheme constructs an interpolable triplet fingerprint set by measuring the harmonic ratio curve, formant displacement curve and transient excitation profile of the original Nanyin sample in short frames. It can accurately capture the details of glissando transition, oral resonance vibration and breath friction in traditional Nanyin female singing at tens of milliseconds. Compared with the existing technology that only generates pitch curves and fixed timbre templates, resulting in stretched electronic sound and block noise, the fingerprint set completely preserves the microscopic acoustic basis of the unique charm of Nanyin, effectively solving the pain point of missing charm and breathiness in Nanyin generation.

[0059] (2) This scheme maps the user intent primitives to the target fingerprint trajectory, and then maps them to the resonator coefficients and excitation primitive coefficient sequences through time window division and feature analysis. At the same time, the pitch and formant are adjusted synchronously through the pre-tuning control sequence. In view of the architectural defects of the existing system where the pitch control and timbre synthesis modules are separated and the formant is not adjusted, the scheme is designed so that the resonator parameters are linked with the cavity characteristics when the pitch changes. This not only satisfies the user's freedom to create according to intent, but also accurately reproduces the cavity texture of the strong linkage between the pitch and timbre of Nanyin, avoiding the unnatural problem caused by the disconnect between pitch and timbre.

[0060] (3) This solution combines fingerprint residual reports and human auditory feedback to establish parameter mapping tables for different sound sources. Through proportional adjustment, offset correction and time window synchronous generation of cross-sound source migration rules, this solution addresses the problem of stiff connection caused by existing technology that only reuses melodies and does not understand the resonance differences of various sound sources when generating cross-sound sources. This solution can adapt to the essential differences between flute airflow resonance, two-string bow pressure resonance and human vocal cavity resonance, so that the Nanyin audio generated across sound sources is connected naturally, eliminating the splicing feeling and improving the flexibility and integrity of multi-sound source Nanyin creation.

[0061] (4) In the process of generating audio stream, this scheme outputs fingerprint residual report by comparing the short-term similarity with the triplet fingerprint, and then generates weighted calibration index and calibration parameter set by combining human listening feedback, forming a closed-loop optimization mechanism of generation, verification and calibration. Compared with the limitations of existing technologies that lack iterative optimization capabilities, the mechanism can continuously correct the resonator coefficient and excitation element parameters, and continuously reduce the feature deviation between the generated audio and the original Nanyin. It can not only adapt to the generation needs of different styles of Nanyin, but also gradually improve the authenticity and fidelity of the generated audio, and enhance the practicality and adaptability of the scheme. Attached Figure Description

[0062] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0063] Figure 1 This is a flowchart of a method for analyzing and generating music from Nanyin instruments according to the present invention;

[0064] Figure 2 This is a flowchart illustrating the relationships between modules in a Nanyin instrumental music analysis and music generation processing system of the present invention. Detailed Implementation

[0065] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0066] Example 1:

[0067] Please see Figure 1 A method for analyzing and generating music from Nanyin instruments, comprising:

[0068] Step 1, constructing the triplet fingerprint set, the specific operations are as follows:

[0069] First, original Nanyin samples are acquired. The acquisition process needs to cover typical samples of different performance or singing styles, such as human voice, flute, and erhu. Then, preprocessing is performed, including noise reduction and normalization, to eliminate environmental interference and unify the signal amplitude reference. In order to capture the details of Nanyin at the tens of millisecond level, such as glissando transitions, resonance vibrations, and breath fricatives, the preprocessed continuous audio signal needs to be divided into short frame sequences. The frame length is usually selected to be 20 to 50 milliseconds, and the frame shift is set to 1 / 4 to 1 / 2 of the frame length to ensure time resolution while avoiding information loss.

[0070] Within each short frame, the harmonic ratio curve is first measured. The spectral distribution is obtained by performing a Fourier transform on the short frame signal, identifying the fundamental frequency and the frequency positions of each harmonic. The ratio of the energy of each harmonic to the fundamental frequency energy is calculated. The change of this ratio over time forms the harmonic ratio curve, and its relationship can be expressed as follows: ,in Let be the harmonic spectrum ratio of the nth harmonic to the fundamental frequency at time t. Let be the energy of the nth harmonic at time t. Let t be the energy of the fundamental frequency. The derivation of the formula is based on the strong correlation between timbre characteristics and harmonic energy distribution. The contribution of different harmonic components to the overall timbre is quantified by the energy ratio, thereby capturing the subtle changes in timbre caused by cavity adjustment in Nanyin.

[0071] Next, the formant displacement curves are measured. Using linear predictive coding or spectral peak detection algorithms, the center frequencies of the formants are identified from the short-frame spectrum. The first 3 to 5 main formants are selected, corresponding to the main resonance characteristics of cavities such as the oral cavity and pharynx in Nanyin music, and their frequency changes over time are tracked. Using the formant frequencies of the stable segments in the sample as a reference, the difference between the formant frequencies at each time point and the reference is calculated to form the formant displacement curves. ,in Let be the displacement of the k-th resonance peak at time t. Let be the actual frequency of the k-th resonance peak at time t. The reference frequency of the resonance peak is derived from the physical characteristic that the resonance peak frequency shifts with the change of cavity shape. The displacement quantity quantifies the influence of cavity adjustment on resonance characteristics, which is directly related to the acoustic essence of techniques such as glissando and vibrato in Nanyin.

[0072] Subsequently, transient excitation profiles are extracted. By calculating the short-time energy change rate of short-frame signals, transient intervals in the signal are identified, such as the onset of sound and sudden changes in breath. Within the transient intervals, parameters such as the start time, duration, peak energy, and energy decay rate are extracted. The distribution of these parameters over time constitutes the transient excitation profile. This process focuses on rapid dynamic events in Nanyin, such as breath friction and bow pressure changes. The temporal and energy characteristics of these events are key to forming the sense of breath.

[0073] Finally, the harmonic ratio curves, resonance peak displacement curves, and transient excitation profiles are aligned along the time axis. Each time point corresponds to a set of triplet data consisting of three curve parameters. These triplet data are processed by cubic spline interpolation to make the data form a smooth and differentiable sequence in the time dimension, ultimately forming a set of triplet fingerprints that can be represented by interpolation.

[0074] Step 2: Based on the user's intent, map the intent primitives to a curve offset of the triplet fingerprint to generate the target fingerprint trajectory. The specific operations are as follows:

[0075] In practice, user intent is usually expressed as a stylistic orientation, emotional expression, or technical requirement for Nanyin works, such as abstract descriptions like gentle glissando transitions and strong breath support. These need to be transformed into a set of quantifiable parameters through standardized mapping. These parameters cover time dimensions, such as the duration of the technique and rhythmic intervals; intensity dimensions, such as the amplitude of pitch changes and transient energy peaks; and characteristic dimensions, such as the formant drift rate and the proportion of harmonic energy distribution. They are indexed according to the order of occurrence on the timeline to form an intent timeline.

[0076] Based on the intent timeline, matching template samples need to be selected from a predefined primitive fingerprint template library. The primitive fingerprint template library contains typical techniques in Nanyin music, such as glissando, vibrato, and breath control, to correspond to triplet fingerprint segments. Each template is associated with specific intent primitive features. The selection process is achieved through feature similarity comparison. For example, the glissando amplitude of 0.5 intervals in the intent primitive is matched with the formant displacement range of the glissando segment in the template. The matched template samples need to be adjusted in terms of time scale and amplitude according to the intent parameters: time scale distortion is achieved through linear stretching or compression to match the duration requirement in the intent; amplitude scaling is adjusted proportionally to the energy proportion of the harmonic ratio, the frequency range of the formant displacement, and the energy peak of the transient profile, so that the dynamic range of the template sample fits the intent intensity, thereby generating a preliminary local target fingerprint segment.

[0077] After the local target fingerprint segment is generated, the compensation coefficient related to cavity coupling needs to be calculated to correct the deviation caused by the dynamic response of the cavity. In the singing of Nanyin, the changes in the shape of cavities such as the oral cavity of the human voice and the tube cavity of the flute will cause multi-dimensional coupling effects. For example, the enhancement of transient excitation may cause the formant frequency to drift, and rapid frequency change may be accompanied by nonlinear fluctuations in amplitude. These couplings need to be eliminated through compensation. First, the local target fingerprint segment is divided into short-time windows with the same short frame length as the original sample, with short frames of 20 to 50 milliseconds. In each window, the transient energy change rate, the increment of transient energy per unit time, the frequency drift rate are expressed as the change of formant frequency per unit time, and the local amplitude deviation is expressed as the difference between the actual amplitude and the linearly predicted amplitude. Based on these indicators, a cavity response offset matrix is ​​constructed. The matrix elements correspond to the offsets in the three dimensions of frequency, amplitude, and transient, respectively, reflecting the comprehensive influence of cavity coupling within the window.

[0078] Directional decomposition of the cavity response offset matrix can separate the coupling offset types between various dimensions, such as frequency and amplitude coupling (amplitude deviation caused by frequency changes) and transient and frequency coupling (frequency drift caused by transient excitation). Based on the influence intensity of each coupling offset in the local target fingerprint segment (intensity represented by the peak value of the offset and duration represented by the number of windows where the offset exceeds a threshold), compensation weights are assigned, with higher weights for stronger and longer-lasting coupling types. The generation of the compensation coefficient sequence is based on a linear mapping between weights and offsets. For example, the compensation coefficients for the frequency dimension need to offset the cumulative deviation caused by the frequency drift rate, ultimately forming a time-aligned compensation offset sequence that corresponds one-to-one with the three channels: harmonic ratio, resonant peak displacement, and transient profile.

[0079] Finally, the local target fingerprint segments are joined together in the order of the intent timeline. During the joining process, the overlapping areas of adjacent segments are smoothed and the boundary abrupt changes are eliminated by weighted averaging. At the same time, the compensation offset sequence is superimposed on the fingerprint segment of the corresponding channel according to time alignment. For example, the frequency compensation offset is superimposed on the formant displacement curve and the amplitude compensation offset is superimposed on the harmonic ratio curve. After superposition, time consistency processing is performed. The time interval of each segment is adjusted by interpolation to ensure that the overall sequence is continuous on the time axis and the sampling rate is uniform. Finally, a continuous interpolable target fingerprint trajectory is formed. This trajectory fully carries the user's intent and closely approximates the physical characteristics of the real Nanyin sound through cavity coupling compensation.

[0080] Step 3, mapping the target fingerprint trajectory to the resonator and the excitation element coefficient sequence, is performed as follows:

[0081] First, the generated target fingerprint trajectory is mapped to a superimposed sequence of resonator coefficients and excitation element coefficients. Through phased feature analysis and physical parameter mapping, the correspondence between the acoustic features of Nanyin and the parameters of the acoustic simulation device is established to reproduce the cavity texture and dynamic details of Nanyin.

[0082] To maintain consistency with previous short-frame processing and avoid information loss, the target fingerprint trajectory needs to be divided into multiple time windows with a short frame length of 20 to 50 milliseconds. Each time window constitutes an independent analysis unit. Within each analysis unit, key features need to be labeled for the harmonic ratio curve, the resonant peak displacement curve, and the transient excitation profile: local extreme points correspond to the inflection points of the curve shape, such as the extreme point of the harmonic ratio reflecting the abrupt change in harmonic energy distribution, and the extreme point of the resonant peak displacement marking the critical point of cavity shape adjustment; the slope direction characterizes the trend of curve change, such as a positive slope corresponding to the increase of resonant peak frequency or the enhancement of transient energy, and a negative slope corresponding to the opposite change; the transient peak value directly quantifies the intensity of transient excitation, such as the energy peak value of human voice breath friction or the intensity value of the abrupt change of two-string bow pressure.

[0083] Based on the aforementioned calibrated characteristics, further extraction of local frequency response features is needed through differential and curvature analysis. Taking the first derivative of the harmonic ratio curve yields the rate of change in harmonic energy distribution, reflecting the speed of timbre transition. Taking the first derivative of the formant displacement curve provides the drift rate of the formant frequency, corresponding to the speed of cavity adjustment, such as the rate of frequency change in Nanyin glissando. Taking the first derivative of the transient excitation profile identifies the steepness of the rising and falling edges of the transient excitation, relating to the rate of change in breath or bow pressure. Simultaneously, curvature analysis determines the trend type of each curve's change; for example, a positive curvature of the formant displacement curve indicates a gradually accelerating frequency drift rate, while a negative curvature indicates a slowing drift rate. Integrating these differential and curvature analysis results with the previously calibrated features allows for the creation of a local frequency response feature table for each time window, comprehensively recording the dynamic changes in the acoustic characteristics of Nanyin within that time window.

[0084] Next, the local frequency response characteristic table is mapped to acoustic simulation devices. The parameters of the resonator include the resonator center frequency, quality factor, and gain curve coefficients. The determination of the resonator center frequency needs to match the resonance characteristics of the Nanyin cavity; its value is equal to the sum of the predefined reference resonance peak frequency and the resonance peak displacement in the target fingerprint trajectory, which can be expressed as: This formula is derived based on the physical principle of resonator-simulated cavity resonance. The cavity resonance frequency is determined by the cavity shape, with reference to the resonance peak frequency. The resonant frequency corresponding to the cavity reference shape, and the resonant peak displacement. This is the frequency shift caused by the dynamic adjustment of the cavity shape. The sum of these two factors gives the center frequency of the resonator corresponding to the cavity resonance at the current moment. This ensures that the resonator can accurately reproduce the dynamic resonance characteristics of the cavity; the quality factor of the resonator is related to the harmonic energy attenuation characteristics of the harmonic spectrum ratio curve. The longer the higher harmonic energy in the harmonic spectrum ratio is maintained, the larger the corresponding quality factor is, in order to simulate the attenuation differences of harmonics in different cavities, such as the human vocal pharynx and the flute cavity; the gain curve coefficient is generated according to the amplitude change trend in the local frequency response characteristic table, corresponding to the amplitude dynamics of the harmonic spectrum ratio curve, to ensure that the amplitude of the resonator output is consistent with the amplitude change of the target timbre. At the same time, the transient trigger point is marked according to the peak position and rising edge start point of the transient excitation profile, providing a time reference for the parameter generation of the excitation element.

[0085] Based on the obtained resonator coefficient sequence, corresponding excitation primitive parameters need to be generated to simulate the excitation source characteristics of Nanyin singing, such as the breath excitation of the human voice and the bow friction excitation of the Erxian (two-stringed bow). The excitation primitive parameters include excitation amplitude, trigger time, and duration: the excitation amplitude directly maps to the transient peak of the transient excitation profile; the higher the peak, the greater the excitation amplitude, in order to restore the breath intensity or bow pressure; the trigger time is strictly aligned with the previously marked transient trigger point to ensure that the start of the excitation source is synchronized with the start of the cavity resonance, avoiding the abrupt disconnect between excitation and resonance; the duration is determined according to the transient excitation. The time span of the contour is determined, such as the duration of breath friction corresponding to the duration of the excitation element, to ensure that the excitation process is consistent with the dynamics of the transient contour. On this basis, the excitation element parameters also need to be optimized in time and space: in the time dimension, the change rhythm of the excitation parameters is adjusted according to the rate of change of the resonator coefficient. For example, when the center frequency of the resonator drifts slowly, the excitation amplitude also needs to change slowly to match the natural transition of the glissando. In the spatial dimension, if there are multiple resonators superimposed, such as simulating the multi-cavity resonance of human voice, the excitation energy needs to be allocated according to the gain weight of each resonator to ensure the synergy of the multi-cavity resonance.

[0086] After generating the resonator coefficient sequence and the excitation element parameter sequence, sequential integration and continuity processing are required. Sequential integration follows the order of time windows, associating the resonator coefficients and excitation element parameters corresponding to each time window to form a preliminary parameter sequence. Since time window division may cause abrupt changes in parameters between adjacent windows, continuity processing is required for the resonator center frequency, gain curve coefficients, and excitation amplitude. Cubic spline interpolation is used to smooth the transition of parameters between adjacent time windows, ensuring that the drift of the center frequency, the change of gain, and the adjustment of the excitation amplitude all present a continuous curve shape, avoiding abrupt changes resembling electronic tones. At the same time, based on the trend of parameter changes after continuity, the time interval and intensity gradient of transient trigger parameters are adjusted to ensure that the start and stop of transient excitation and parameter changes are naturally connected. For example, during glissando, the intensity of transient excitation needs to increase or decrease synchronously with the change of resonator gain.

[0087] Finally, the resonator coefficient sequence and the excitation element parameter sequence after continuous processing are integrated into the final target driving sequence, and the superposition relationship and time correspondence information of each parameter are recorded in detail. The superposition relationship includes the parallel superposition logic of multiple resonators, such as the superposition of two resonators to simulate the resonance of the human oral cavity and nasal cavity, and the correspondence between excitation elements and resonators, such as a specific excitation element driving only the resonator in the corresponding frequency band. The time correspondence information clarifies the precise time of action and duration of each coefficient and parameter on the time axis, ensuring that the resonator can adjust its parameters according to the target timing and the excitation element can be triggered at the correct time during the subsequent audio generation process, thereby accurately reproducing the Nanyin charm and dynamic details carried by the target fingerprint trajectory.

[0088] Step 4, the response mapping between the resonator and the excitation element coefficient sequence, and the solution of the pre-tuning control sequence are performed as follows:

[0089] First, a response mapping relationship is established between the resonator coefficient sequence and excitation element coefficient sequence obtained in step 3 and the triplet fingerprint from step 1. Then, the pre-tuning control sequence is solved in reverse. By quantifying the deviation between the actual resonator response and the ideal target and compensating for the physical response delay, the precise synchronization between the acoustic characteristics of Nanyin and the parameters of the analog device is achieved. Finally, the natural dynamic connection and breathing feel of Nanyin are restored through transient pre-shaping. This process requires first clarifying the actual response law under the action of device parameters, and then eliminating deviations and delays through reverse calculation to construct a control strategy that fits the physical sound production logic.

[0090] A continuous sequence of resonator coefficients is applied to the original Nanyin triplet fingerprint from step 1 to capture the true response relationship between device parameters and acoustic characteristics. The center frequency, quality factor, and gain curve coefficients in the resonator coefficient sequence correspond to the resonant frequency, resonance duration, and energy output intensity of the simulated Nanyin sound cavity, respectively. These coefficients are applied to the corresponding channels of the original triplet fingerprint in time windows: the center frequency parameter acts on the resonant peak displacement curve, inducing dynamic changes in the original resonant peak frequency by adjusting the resonant reference of the simulated cavity; the quality factor parameter acts on the harmonic ratio curve, affecting the duration of the original harmonic energy distribution by changing the attenuation rate of different harmonics in the simulated cavity; and the gain curve coefficient directly acts on the amplitude dimension of the transient excitation profile and harmonic ratio, adjusting... The overall output level of the original transient intensity and harmonic energy; in this process, the response offset of the harmonic ratio, resonant peak displacement and transient profile needs to be recorded window by window, that is, the difference between the actual output value of the original fingerprint under the action of device parameters and the corresponding value of the target fingerprint trajectory in step 2. This difference reflects the deviation between the ideal mapping of the device and the actual response. At the same time, the response delay of each time window needs to be derived: since there are physical delays in acoustic simulation, such as the propagation time of simulated sound waves in the cavity and the propagation delay of the excitation signal, by comparing the change time of the resonator coefficient sequence with the change time of the triplet fingerprint response, the time difference between the two can be calculated to obtain the response delay value of each time window. Finally, the response offset and delay value of each time window are integrated to form a local response mapping covering the entire time axis.

[0091] Based on the response offset and delay information in the local response mapping, a pre-tuned control sequence needs to be obtained through inverse calculation to eliminate deviations and synchronize response timing. For the pre-tuning compensation of the resonator coefficients, the original coefficients are corrected by driving offset, so that when the pre-tuned coefficients are applied to the triplet fingerprint, the response offset is exactly canceled out to zero. The relationship can be expressed as follows: The derivation logic of this formula utilizes the compensation principle of deviation cancellation to determine the original resonator center frequency. The response offset generated under the action is That is, the difference between the actual resonant peak output and the target value, in order to make the pre-tuned center frequency... Its response can accurately match the target and drive the offset. Must equal By adjusting in the opposite direction to offset the original deviation, the resonant output resonance characteristics are ensured to match the target fingerprint requirements. Similarly, the pre-adjustment values ​​of the quality factor and gain curve coefficients are also calculated using similar logic, and the corresponding driving offset is determined based on the response offset of the harmonic ratio and transient profile, respectively.

[0092] To address the timing misalignment caused by response delay, synchronization needs to be achieved through pre-excitation morphological shaping of the transient excitation. This is because the resonator response ratio is delayed by the input coefficient. , representing the response delay value of a certain time window. If the transient excitation is triggered at the original time, the excitation effect will lag behind the resonator response, resulting in the loss of the natural sense of breath and pitch synchronization in Nanyin music. Therefore, the excitation triggering time needs to be advanced based on the original triggering position t of the transient profile. That is, the trigger time is adjusted to At the same time, keeping the amplitude and duration of the excitation constant, this pre-triggering allows the energy release of the transient excitation and the delay response of the resonator to be precisely aligned on the time axis. For example, when simulating vocal glissando, the pre-excitation of breath can be synchronized with the delayed rise of the formant, avoiding the abrupt feeling of pitch preceding breath. Finally, through the synergy of the resonator pre-tuning coefficient and the pre-transient excitation, a complete pre-tuning control sequence is formed.

[0093] Step 5: The pre-tuned control sequence and the target fingerprint trajectory are combined in layers to construct a control package that aligns the musical phrase sequence with time. This package includes the resonator coefficient time sequence, excitation element markers, and pre-tuned offsets. The specific operations are as follows:

[0094] The process of constructing a control package involves layering and combining the pre-tuned control sequence from step 4 with the target fingerprint trajectory from step 2. Through the structured integration and temporal coordination of multi-dimensional features, the dispersed acoustic control parameters are transformed into a unified control unit precisely aligned with the musical phrase sequence. This process requires first establishing a multi-layered feature index framework, and then forming a complete control structure through temporal alignment and parameter fusion. The layered index is the foundation of the combination process, decomposing the complex features of the pre-tuned control sequence and the target fingerprint trajectory into independently controllable yet interconnected layers according to the compositional logic of Nanyin music. Specifically, the beat layer focuses on the basic units of time scale, corresponding to the rhythmic patterns and beat intervals of Nanyin. Its parameters include beat rate, beat strength distribution, and phrase breaks, providing a reference grid for overall time alignment. The melody layer, centered on pitch changes, integrates the macroscopic trend of the formant displacement curve in the target fingerprint trajectory with the pre-tuned offset of the corresponding pitch in the pre-tuned control sequence, reflecting... The text describes the development of a layered system for melody, focusing on the dynamic characteristics of cavity resonance. It examines the center frequency, quality factor, and harmonic ratio curves of the target fingerprint trajectory within the resonator coefficient sequence, reflecting the influence of different cavity shapes on timbre, such as the opening and closing of the human mouth and the fingering changes on a flute. The transient layer focuses on rapid changes in the excitation source, integrating excitation primitive markers, transient excitation profiles, and transient pre-trigger parameters in the pre-tuned control sequence to capture instantaneous dynamics such as breath friction and bow pressure abrupt changes. During the layering process, the time points of each layer need to be aligned. Using the time grid of the beat layer as a reference, the pitch change moments of the melody layer, the resonance characteristic switching moments of the cavity layer, and the excitation trigger moments of the transient layer are all mapped to a unified time axis. Key nodes of each layer are also marked, such as the pitch inflection points of the melody layer, the formant abrupt change points of the cavity layer, and the strong excitation starting points of the transient layer. These nodes are the focus of subsequent parameter fusion.

[0095] Based on time alignment, parameters within adjacent time windows need to be locally weighted and superimposed to eliminate parameter abrupt changes caused by time window division and achieve a smooth transition of features at each layer. For the resonator coefficient sequence, parameters such as the center frequency and quality factor of adjacent time windows are fused through weighted averaging, with the weights dynamically adjusted with time distance. Window parameters closer to the current time step have higher weights to ensure the continuity of parameter changes. The superposition of excitation primitive sequences needs to consider the temporal correlation of excitations. For example, if there is energy superposition between the excitation tail tone of the previous window and the excitation start tone of the current window, weight allocation is used to avoid energy superposition. The transition is made by abruptly maintaining the natural flow of breath or bow pressure. After the local fusion of adjacent windows is completed, the parameters of each layer are combined at the same time step: the resonator coefficient time sequence integrates the dynamic resonance parameters of the cavity layer with the compensation offset in the pre-tuning control sequence, the excitation primitive marker is associated with the triggering information of the transient layer and the pitch state of the melody layer, and the pre-tuning offset contains the correction amount of each layer's parameters, such as the pitch pre-tuning of the melody layer and the triggering time pre-tuning of the transient layer; these parameters form an interconnected whole within the time step, namely a multi-layer control package, and each control package corresponds to the complete control information of a time unit in the musical phrase sequence.

[0096] After the control package is generated, a time sequence integrity check is required to ensure the coherence and reliability of the overall control logic. The check includes a time continuity check to confirm that the time steps of adjacent control packages do not overlap or have gaps, and that parameter changes conform to the preset time resolution. At the same time, by analyzing the rate of change of each parameter, local high-change segments are identified. These segments usually correspond to technique transitions in Nanyin music, such as the transition from a gentle long note to a fast glissando. These segments need to be marked to indicate fine-tuning in the subsequent generation process. In addition, the boundary segments of musical phrases also need to be marked, such as the onset segment at the beginning of a phrase and the closing segment at the end. The acoustic characteristics of these segments often have special features, such as the breath support at the onset and the resonance decay at the closing, requiring targeted control strategies. After verification and marking, the output control package set achieves the organic integration of layered features, preserving the independent controllability of each layer of parameters while ensuring overall synergy through time alignment.

[0097] Step 6, generating the audio stream based on the control packet and constructing the fingerprint residual report, the specific operations are as follows:

[0098] The system updates resonator coefficients and excitation elements according to the control package time sequence, performs phase synchronization correction, and generates audio streams and fingerprint residual reports. This transforms the structured control package parameters into continuous acoustic output. Incremental parameter updates are fundamental to achieving dynamic changes in acoustic characteristics. Based on the control package's time markers and superposition coefficients, it avoids electronic sound signatures caused by abrupt parameter changes. The time markers recorded in the control package correspond to the time axis of the musical phrase sequence. Resonator coefficients, such as center frequency, quality factor, and gain, as well as excitation element parameters, such as trigger time, amplitude, and duration, must be extracted step-by-step according to this time sequence. Incremental updates do not directly replace parameters; instead, they are based on the parameter values ​​of the previous time step and the target parameter values ​​of the current control package, and are superimposed... The incremental change is calculated by adding a coefficient to smoothly transition the parameters from the current state to the target state. For example, if the center frequency of the resonator needs to increase from 1000Hz to 1100Hz in the current time step, and the superposition coefficient is 0.2, then only 20Hz will be updated in this time step, and the remaining 80Hz will be gradually completed in subsequent time steps to avoid frequency abrupt changes. At the same time, the trigger flag of the excitation element determines whether to activate a new excitation, such as the injection of new breath in the human voice or the start of a new bow section on the second string. The intensity of the activated excitation is adjusted according to the amplitude change. When the amplitude increases, the excitation energy is gradually increased, and when the amplitude decreases, the energy is slowly decreased, ultimately forming a local incremental trigger sequence to simulate the natural fluctuations of the excitation source in Nanyin, such as the gradual increase and decrease of breath and the gradual change of bow pressure.

[0099] Phase synchronization correction is crucial for resolving acoustic phase misalignment during parameter updates and ensuring sound continuity. The phase of a resonator accumulates and shifts with changes in its center frequency. Without correction, phase differences between adjacent time steps can lead to interference distortion during sound wave superposition, disrupting the smoothness of Nanyin melodies. For example, the phase corresponding to frequency changes in glissando needs to be continuously connected. The calculation of phase shift must be based on the changes in the resonator's center frequency in the incremental sequence, and the relationship can be expressed as follows: The derivation logic of this formula originates from the physical relationship between acoustic phase and frequency. Frequency is the rate of change of phase over time, while phase shift is the cumulative change of frequency relative to a reference value. Let be the phase shift at time t. Let τ be the center frequency of the resonator at time τ. This is the reference center frequency of the resonator, corresponding to the stable resonant frequency of the original Nanyin sample. The starting time of the current parameter update segment is from [time period]. To t, It is a coefficient that converts frequency into phase, obtained from calculations. The triggering start phase of the transient primitive needs to be adjusted so that the start phase of the excitation primitive is consistent with the current phase of the resonator. At the same time, piecewise smooth interpolation is applied in the transition section of the resonator coefficient, such as the interval where the center frequency changes from one stable value to another, to further eliminate phase abrupt changes and ensure continuous phase transition on the entire time axis.

[0100] Short-term similarity comparison is the step to quantify the deviation between the generated result and the original Nanyin characteristics. It needs to build a verification benchmark based on the triplet fingerprint in step 1, such as harmonic ratio, formant shift, and transient profile. The cavity state data of each time step, that is, the harmonic ratio, formant frequency and transient energy distribution corresponding to the current resonator output, need to be compared point by point with the original triplet fingerprint in the same time window in step 1: the comparison of harmonic ratio focuses on the difference in the proportion of each harmonic energy, the comparison of formant shift focuses on the offset of the current formant frequency from the original formant frequency, and the comparison of transient profile focuses on the deviation of the transient energy peak and rise time from the original profile. The deviation values ​​of these dimensions are integrated into a vector form, that is, the residual vector of the time window is obtained. Each element in the vector corresponds to the deviation of an acoustic feature dimension. The larger the element value, the more significant the difference between that dimension and the original feature.

[0101] The generation of the continuous audio stream and the fingerprint residual report is the final output of the process. The resonator output at each time step is essentially a sound wave signal with corresponding frequency and amplitude. These sound wave signals need to be accumulated and superimposed in chronological order: the tail of the sound wave in the previous time step and the onset of the sound wave in the current time step are naturally superimposed to form a seamless continuous audio stream, simulating the coherent performance or singing of Nanyin phrases. At the same time, the residual vectors of each time window are organized according to the time sequence, and the time interval, main deviation dimension and deviation value range corresponding to each residual vector are marked to form a fingerprint residual report. This report intuitively reflects in which time periods and in which acoustic feature dimensions the generated audio deviates from the original Nanyin features. For example, if the formant displacement residual of a certain glissando is too large, it indicates that the cavity simulation of that section is not accurate enough. Finally, the continuous audio stream and the fingerprint residual report are output synchronously.

[0102] Step 7: Generate the calibration parameter set and cross-sound source migration rules based on fingerprint residuals and auditory feedback. The specific operations are as follows:

[0103] First, the fingerprint residual report and human auditory feedback output from step 6 need to be received. The objective residuals and subjective feelings are then transformed into unified quantitative calibration indicators. The fingerprint residual report includes the harmonic ratio, formant shift, and transient profile residuals for each time window across the entire time axis. The temporal dimension information of the residuals must be preserved according to the original time window division method to avoid losing local deviation features, such as concentrated formant residuals in a certain glissando segment or significant transient residuals in a certain starting segment. Human auditory feedback typically includes subjective evaluations of the completeness of the timbre, the naturalness of the breathiness, and the similarity of the timbre. These qualitative descriptions need to be mapped to a subjective score of 0-10 and associated with the corresponding time window. For example, the first auditory feedback... If the glissando is stiff within 0-15 seconds, the subjective score for that time period will be lowered, and this score will be linked to the formant shift residual for that time window. Based on this, a weighted calibration index is generated: this index needs to comprehensively consider the magnitude of the objective residual and the level of the subjective score, while assigning differentiated weights to different residual channels. For example, the weight values ​​for the channel residuals of harmonic ratio, formant shift, and transient profile are determined according to the factors influencing the flavor of Nanyin music. Formant shift is directly related to changes in cavity morphology, such as cavity adjustments for glissando and vibrato, which has the greatest impact on the flavor and therefore the highest weight. Transient profile is related to breathiness and dynamic transition, and has the next highest weight. Harmonic ratio affects the fullness of timbre and has a relatively low weight. Their relationship can be expressed as follows: This formula is derived from the fact that objective residuals determine the degree of bias, subjective scoring filters key biases, and weights distinguish the importance of features: where I(t) is the weighted calibration index of the time window at time t. , , The weights for harmonic ratio, resonant shift, and transient profile residual are respectively, satisfying the following conditions: ,and , , , Let I(t) represent the absolute value of the residual for the corresponding channel at time t, and S(t) be the subjective rating mapping coefficient at time t. A higher subjective rating results in S(t) closer to 1, amplifying the impact of the residual on the index; a lower rating results in S(t) closer to 0.5, avoiding miscalibration caused by extreme ratings. By calculating I(t) for each time window, the relationship between the residual of each channel and subjective preference can be analyzed, for example... The residual of the resonance peak has the highest correlation with I(t), indicating that the resonance peak deviation is the main factor affecting the listening experience. Therefore, the cavity characteristics are associated with the resonance peak, and the adjustment order is prioritized over the transient characteristics to ensure that the calibration resources focus on the key deviations.

[0104] Based on the weighted calibration index, it is necessary to calculate the calibration offset sequence of each resonator coefficient and excitation element, and establish a parameter mapping table across sound sources. The derivation of the calibration offset sequence must follow the principle that the higher the index, the larger the adjustment amount: For each time window, the resonator coefficients within that time window, such as center frequency, quality factor, and gain, and the excitation element parameters, such as amplitude and trigger time, are determined based on the value of I(t), serving as the adjustment direction and amplitude. If I(t) originates from a large residual in the resonant peak displacement, it indicates... If the proportion is high, the calibration offset needs to be directed towards the corrected resonant peak frequency. If the current resonator center frequency is too low, resulting in a positive resonant peak residual, then the calibration offset should be negative, reducing the center frequency to the target value. If I(t) originates from the transient profile residual, it indicates... If the proportion is high, the calibration offset needs to adjust the triggering time of the excitation primitive, such as triggering earlier to enhance the feeling of breathing, or the amplitude, such as increasing the excitation amplitude to match the original transient intensity. At the same time, it is necessary to establish a parameter mapping table for different sound source categories: the acoustic characteristics of different sound sources in Nanyin are fundamentally different. The human voice relies on the resonance of the oral cavity and pharyngeal cavity, with a narrow range of resonance peak frequencies, usually concentrated in 200-3000Hz. The transient excitation is driven by breath, and the intensity changes slowly. The flute relies on the resonance of the tube cavity, with a wide range of resonance peak frequencies of 500-5000Hz. The transient excitation is the impact of airflow, and the intensity changes abruptly. The erhu relies on the resonance of the string vibration and the resonance box, with the resonance peak frequencies concentrated in 100-2000Hz. The transient excitation is the friction of bow pressure, and the intensity fluctuates with the bow speed. Based on these differences, the calibration parameters of one sound source, such as the human voice, are mapped to another sound source, such as the flute, through proportional adjustment. For example, if the calibration offset of the human voice's resonant peak is -50Hz, the corresponding flute needs to be adjusted to -100Hz according to the ratio of the cavity resonant frequency, approximately 1:2. The inherent deviation of the sound source is compensated through offset correction. For example, the bow pressure excitation of the second string naturally has a 20ms delay, so an additional 20ms advance needs to be added to the calibration offset at the trigger moment. The timing of parameter adjustment for different sound sources is matched to their response characteristics through time window synchronization. For example, the flute has a fast airflow response, so the time window can be shortened to 15ms; the second string has a slow bow pressure response, so the time window needs to be extended to 30ms. Finally, calibration coefficients and cross-sound source migration rules are generated.

[0105] Finally, the calibration coefficients and migration rules need to be integrated into a time-aligned multi-channel calibration parameter set to ensure that it can be directly used for cross-source control parameter adjustment. The integration process should be centered on the time axis, arranging the resonator calibration coefficients and excitation element calibration offsets for each time window in chronological order, while marking the applicable source category for each parameter. For example, formant calibration offset -50Hz: only applicable to human voice; excitation amplitude calibration coefficient 1.2: applicable to flute, to avoid misuse of parameters on mismatched sources. In addition, time synchronization information needs to be added, as different sources have different acoustic response delays, such as human voice response. The delay is approximately 10ms for the flute and approximately 5ms for the octave. The delay compensation amount for each sound source needs to be marked in the parameter set to ensure that the timing of the calibration parameters takes effect at the same time as the actual response of the sound source. For example, the calibration parameters for the flute need to take effect 5ms in advance. The integrated parameter set needs to pass the integrity check to ensure that each time window and each sound source category has corresponding calibration parameters, and that the migration rules are logically consistent. For example, the proportional adjustment coefficient remains stable in different time windows of the same sound source. Finally, an optimized parameter set that can be directly loaded into the control sequence adjustment module is formed to achieve the unity of the charm and the improvement of the naturalness of the Nanyin music generated across sound sources.

[0106] Example 2:

[0107] Please see Figure 2 Based on Example 1, this embodiment provides a Nanyin instrumental music analysis and music generation processing system, including:

[0108] The fingerprint acquisition module is used to acquire the original Nanyin sample and measure and record the harmonic ratio curve, formant displacement curve and transient excitation profile within a short frame to form a set of triplet fingerprints that can be interpolated.

[0109] The trajectory generation module is used to map the intent primitives to curve offsets of the triplet fingerprint in the fingerprint acquisition module according to the user's intent, thereby generating the target fingerprint trajectory.

[0110] The coefficient mapping module is used to map the target fingerprint trajectory in the trajectory generation module into a superimposed resonator coefficient sequence and an excitation element coefficient sequence.

[0111] The pre-tuning solution module is used to obtain the resonator coefficient sequence and excitation element coefficient sequence according to the mapping in the coefficient mapping module, establish a response mapping relationship with the triplet fingerprint in the fingerprint acquisition module, and solve the pre-tuning control sequence in reverse to perform pre-shape shaping of transient excitation;

[0112] The control package construction module is used to hierarchically combine the pre-tuning control sequence in the pre-tuning solution module with the target fingerprint trajectory in the trajectory generation module to construct a control package that is time-aligned with the musical phrase sequence, including the resonator coefficient time series, excitation element markers and pre-tuning offsets.

[0113] The audio generation module is used to update the resonator coefficients and excitation elements sequentially according to the control package in time order, and perform phase synchronization correction. At the same time, it performs a short-term similarity comparison with the fingerprint in the fingerprint acquisition module to generate an audio stream and a fingerprint residual report.

[0114] The calibration migration module generates a set of calibration parameters and migration rules based on the fingerprint residual report from the audio generation module and human auditory feedback, and applies them to the adjustment of control parameters across sound sources.

[0115] The above description is merely a preferred embodiment of the present invention; however, the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and its improved concepts, should be covered within the scope of protection of the present invention.

Claims

1. A method for analyzing and generating music from Nanyin instruments, characterized in that, include: Step 1: Obtain the original Nanyin sample, and measure and record the harmonic ratio curve, formant displacement curve and transient excitation profile within a short frame to form a set of triplet fingerprints that can be interpolated. Step 2: Based on the user's intent, map the intent primitives to a curve offset of the triplet fingerprint in Step 1 to generate the target fingerprint trajectory. Step 3: Map the target fingerprint trajectory from Step 2 into a superimposed sequence of resonator coefficients and a sequence of excitation element coefficients; Step 4: Based on the resonator coefficient sequence and excitation element coefficient sequence obtained in Step 3, establish a response mapping relationship with the triplet fingerprint in Step 1, and solve the pre-tuning control sequence in reverse to perform pre-morphological shaping on the transient excitation. Step 5: Combine the pre-tuned control sequence from Step 4 with the target fingerprint trajectory from Step 2 in a hierarchical manner to construct a control package that aligns the musical phrase sequence with time, including the resonator coefficient time sequence, excitation element markers, and pre-tuned offset. Step 6: Update the resonator coefficients and excitation elements sequentially according to the control package in chronological order, and perform phase synchronization correction. At the same time, perform a short-term similarity comparison with the fingerprint from Step 1 to generate an audio stream and fingerprint residual report. Step 7: Based on the fingerprint residual report and human auditory feedback from Step 6, generate a set of calibration parameters and migration rules, and apply them to the adjustment of control parameters across sound sources.

2. The method for analyzing and generating music from Nanyin instruments according to claim 1, characterized in that, include: Step 21: Parse the user intent into a normalized parameter set consisting of multiple intent primitives, and form an intent timeline by time index; Step 22: Based on the intent timeline, select template samples corresponding to the intent primitives from the predefined primitive fingerprint templates, and distort and scale the template samples in terms of time scale and amplitude to generate local target fingerprint segments. Step 23: Calculate the cavity coupling-related compensation coefficients for the local target fingerprint segment, and generate the time-aligned compensation offset sequence of the corresponding harmonic ratio, resonant peak displacement and transient profile. Step 24: Adhere the local target fingerprint segments in chronological order and superimpose the compensation offset sequence onto the corresponding channel to perform time consistency processing, thereby obtaining a continuous interpolable target fingerprint trajectory.

3. The method for analyzing and generating music from Nanyin instruments according to claim 2, characterized in that, include: Step 31: Divide the target fingerprint trajectory from Step 2 into multiple analysis units according to the time window, and calibrate the harmonic ratio, resonant peak displacement, and local extreme points, slope direction, and transient peak value of the transient profile within each unit. Step 32: Take the fingerprint data in each analysis unit divided in Step 31 as input, select the matching elementary resonator template, and calculate the corresponding resonator center frequency, quality factor and gain curve coefficient, and generate the corresponding excitation elementary parameters. Step 33: Integrate the resonator coefficient sequence generated in step 32 with the excitation element parameter sequence in sequence, and perform continuous processing on the center frequency, gain and excitation amplitude, while adjusting the transient trigger parameters. Step 34: Combine the resonator coefficient sequence and the excitation element parameter sequence obtained in step 33 to form the final target driving sequence, and record the superposition relationship and time correspondence information of each coefficient and excitation parameter.

4. The method for analyzing and generating music from Nanyin instruments according to claim 3, characterized in that, include: Step 41: Apply the continuous resonator coefficient sequence from step 3 to the triplet fingerprint from step 1, record the response shift of the harmonic spectrum ratio, resonance peak shift, and transient profile within each time window, and derive the response delay of each time window to form a local response mapping. Step 42: Based on the local response mapping and delay information from step 41, the pre-tuning control sequence for each time window is calculated in reverse, the driving offset is applied to the resonator coefficient sequence, and a short-term excitation is applied at the trigger position of the transient profile to form the pre-form shaping.

5. The method for analyzing and generating music from Nanyin instruments according to claim 4, characterized in that, include: Step 51: The pre-tuned control sequence of Step 4 and the target fingerprint trajectory of Step 2 are indexed in layers according to beat layer, melody layer, cavity layer and transient layer, and the time points of each layer are aligned, while key nodes are marked. Step 52: Based on the time alignment in step 51, the resonator coefficient sequence and excitation element sequence in adjacent time windows are locally weighted and superimposed, and the parameters of each layer are combined in the same time step to form a multi-layer control package, which includes the resonator coefficient sequence, excitation element marker and pre-adjustment offset. Step 53: Perform time series integrity verification on the control package generated in step 52, mark local high-variance segments and boundary segments, and output a hierarchical fused control package set with optimization markings.

6. The method for analyzing and generating music from Nanyin instruments according to claim 5, characterized in that, include: Step 61: The time stamp and superposition coefficient sequence in the control package generated in step 5 are used to incrementally update the resonator coefficients and excitation elements in chronological order, and the excitation elements are dynamically activated or deactivated according to the trigger mark and amplitude change to form a local incremental trigger sequence. Step 62: Calculate the phase shift of the resonator at each time step based on the incremental sequence output in step 61, and adjust the triggering phase of the transient primitive. At the same time, apply piecewise smooth interpolation in the coefficient transition section to maintain continuity. Step 63: Perform a short-time comparison between the cavity state data output in step 62 and the triplet fingerprint in step 1 according to the time window, and extract the residual vectors of harmonic ratio, resonant peak shift and transient profile. Step 64: The resonator outputs of each time step in step 63 are accumulated and superimposed to form a continuous audio stream, and the residual vector is organized into a fingerprint residual report arranged in time sequence. At the same time, the audio stream and the residual report are output.

7. The method for analyzing and generating music from Nanyin instruments according to claim 6, characterized in that, include: Step 71: Receive the fingerprint residual report and human auditory feedback from Step 6, divide the harmonic ratio, resonant peak shift and transient profile deviation in the residual into time windows, map them to subjective scores, generate weighted calibration indexes, and analyze the relationship between residuals of each channel and subjective preferences to determine the adjustment order of cavity state and transient features. Step 72: Based on the weighted calibration index output in step 71, calculate the calibration offset sequence of each resonator coefficient and excitation element, and establish a parameter mapping table for different sound source categories. Through proportional adjustment, offset correction and time window synchronization, generate calibration coefficients and cross-sound source migration rules. Step 73: Integrate the calibration coefficients and migration rules generated in step 72 into a time-aligned multi-channel calibration parameter set, mark the applicable sound source categories and time synchronization information, and form an integrated parameter set that can be directly used to control sequence adjustment.

8. The method for analyzing and generating music from Nanyin instruments according to claim 2, characterized in that, Calculate cavity coupling-related compensation coefficients for local target fingerprint segments, including: Step 231: Divide the local target fingerprint segment into short-time windows, measure and record the transient energy change rate, frequency drift rate and local amplitude deviation of each window, and construct the cavity response offset matrix for each time window accordingly. Step 232: The cavity response offset matrix from step 231 is decomposed into frequency, amplitude, and transient dimensions to identify various coupling offsets. Compensation weights are assigned based on their influence intensity and duration in the target fingerprint segment to generate a compensation coefficient sequence for correcting amplitude, frequency, and transient triggering.

9. The method for analyzing and generating music from Nanyin instruments according to claim 3, characterized in that, Calculate the corresponding resonator center frequency, quality factor, and gain curve coefficients, and simultaneously generate the corresponding excitation element parameters, including: Step 321: Perform differential and curvature analysis on the harmonic ratio, resonant peak displacement and transient profile of the target fingerprint segment divided by time window, extract local extreme points, slope changes and frequency drift trends, and establish a local frequency response characteristic table for each time window. Step 322: Map the local frequency response feature table generated in step 321 to the resonator center frequency, quality factor and gain curve coefficients for each time window, and mark the transient trigger point; Step 323: Based on the dynamic resonator coefficient sequence and transient trigger marker from step 322, perform spatiotemporal optimization on the amplitude, trigger time, and duration of the excitation element, so that the excitation element and the resonator coefficient are aligned in time and form the final parameter sequence.

10. A Nanyin instrumental music analysis and music generation processing system, applied to the Nanyin instrumental music analysis and music generation processing method according to any one of claims 1-9, characterized in that, include: The fingerprint acquisition module is used to acquire the original Nanyin sample and measure and record the harmonic ratio curve, formant displacement curve and transient excitation profile within a short frame to form a set of triplet fingerprints that can be interpolated. The trajectory generation module is used to map the intent primitives to curve offsets of the triplet fingerprint in the fingerprint acquisition module according to the user's intent, thereby generating the target fingerprint trajectory. The coefficient mapping module is used to map the target fingerprint trajectory in the trajectory generation module into a superimposed resonator coefficient sequence and an excitation element coefficient sequence. The pre-tuning solution module is used to obtain the resonator coefficient sequence and excitation element coefficient sequence according to the mapping in the coefficient mapping module, establish a response mapping relationship with the triplet fingerprint in the fingerprint acquisition module, and solve the pre-tuning control sequence in reverse to perform pre-shape shaping of transient excitation; The control package construction module is used to hierarchically combine the pre-tuning control sequence in the pre-tuning solution module with the target fingerprint trajectory in the trajectory generation module to construct a control package that is time-aligned with the musical phrase sequence, including the resonator coefficient time series, excitation element markers and pre-tuning offsets. The audio generation module is used to update the resonator coefficients and excitation elements sequentially according to the control package in time order, and perform phase synchronization correction. At the same time, it performs a short-term similarity comparison with the fingerprint in the fingerprint acquisition module to generate an audio stream and a fingerprint residual report. The calibration migration module generates a set of calibration parameters and migration rules based on the fingerprint residual report from the audio generation module and human auditory feedback, and applies them to the adjustment of control parameters across sound sources.