Multi-source input-oriented electronic musical instrument intelligent sound effect optimization method and system
By collecting multi-source data to construct a dynamic voiceprint network and generating optimized sound effect parameters, the problem of existing technologies being unable to accurately reproduce the performer's intentions and environment in sound effects is solved, thus improving the accuracy and adaptability of electronic musical instrument sound effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-03-17
AI Technical Summary
Existing intelligent sound effect optimization technology for electronic musical instruments cannot fully consider the performer's limb movement data and environmental acoustic data, resulting in the optimized sound effects failing to accurately reproduce the performer's true intentions and the actual environment, thus affecting musical expressiveness.
By collecting audio signals, performer's limb movement data, and environmental acoustic data in real time, a dynamic voiceprint network is constructed. Based on multi-dimensional feature vectors and the connection relationships between nodes, optimized sound effect parameters are generated. These parameters are then matched and corrected using a predefined sound effect template library, and the processed audio stream is output.
It achieves improved accuracy and adaptability of sound effects, ensuring that the output audio stream better matches the performer's expressive intent and environmental characteristics, thereby enhancing the overall expressiveness of the music.
Smart Images

Figure CN121686975A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and more specifically, to a method and system for intelligent sound effect optimization of electronic musical instruments for multi-source input. Background Technology
[0002] In the field of modern music composition and performance, the widespread application of electronic musical instruments has greatly enriched the expressiveness and diversity of music. Intelligent sound effect optimization for electronic instruments is crucial for enhancing the flexibility of musical composition, the emotional impact of performance, and the audience's immersion. Through intelligent sound effect optimization, creators can break through the limitations of traditional instrument sound effects, easily shaping unique and creative sound effects according to different musical styles and creative intentions, thus injecting new vitality into musical works. Simultaneously, in live performances, high-quality sound effects can better engage the audience's emotions, enhance the overall performance effect, and make the musical performance more captivating.
[0003] However, existing intelligent sound optimization technologies for electronic musical instruments still have certain shortcomings. On the one hand, most existing technologies can only process audio signals individually, and cannot comprehensively consider multi-source information such as the performer's limb movement data and environmental acoustic data. This makes it difficult for the optimized sound effects to accurately reproduce the performer's true intentions, and also makes it impossible to perfectly blend with the actual performance environment, greatly reducing the overall expressiveness of the music.
[0004] There is currently no effective solution to the above problems. Summary of the Invention
[0005] This application provides a method and system for intelligent sound effect optimization of electronic musical instruments with multi-source input to solve the above-mentioned technical problems.
[0006] This application provides an intelligent sound effect optimization method for electronic musical instruments with multi-source input, comprising: real-time acquisition of multi-source input data of electronic musical instruments, including audio signals, performer's limb movement data, and environmental acoustic data; feature extraction and encoding of the multi-source input data to generate multi-dimensional feature vectors with temporal correlation; construction of a dynamic voiceprint network based on the dynamic correlation between the multi-dimensional feature vectors; wherein the dynamic voiceprint network consists of nodes and connections between nodes, each node corresponding to a performance event, and node attributes including the multi-dimensional feature vector encoding value of the corresponding performance event; the connections between nodes are dynamically generated based on the similarity of motion trajectories and audio phase coherence between consecutive performance events, and are assigned connection weight values representing the correlation strength; matching the structural features of the dynamic voiceprint network with a predefined sound effect template library, selecting a basic sound effect parameter set based on the node distribution density and the connection weight values between nodes, and superimposing parameter correction terms based on the existence state of specific types of connections in the dynamic voiceprint network to generate optimized sound effect parameters; and outputting the processed audio stream based on the optimized sound effect parameters.
[0007] This application provides an intelligent sound effect optimization system for electronic musical instruments with multi-source input, comprising: a multi-source data acquisition module for real-time acquisition of multi-source input data of the electronic musical instrument, wherein the multi-source input data includes audio signals, performer's limb movement data, and environmental acoustic data; a feature extraction module for extracting and encoding features from the multi-source input data to generate multi-dimensional feature vectors with temporal correlation; and a voiceprint network construction module for constructing a dynamic voiceprint network based on the dynamic correlation between the multi-dimensional feature vectors; wherein the dynamic voiceprint network consists of nodes and the connections between nodes, with each node corresponding to a performance event. The system comprises several modules, each with its own attributes. Node attributes include multi-dimensional feature vector encoding values corresponding to the performance events. The connections between nodes are dynamically generated based on the similarity of motion trajectories and audio phase coherence between consecutive performance events, and are assigned connection weights representing the strength of the association. A matching optimization module matches the structural features of the dynamic voiceprint network with a predefined sound effect template library, selects a basic sound effect parameter set based on node distribution density and connection weights, and superimposes parameter correction terms based on the existence of specific types of connections in the dynamic voiceprint network to generate optimized sound effect parameters. An output module outputs the processed audio stream based on the optimized sound effect parameters.
[0008] Based on the embodiments provided in this application, by real-time acquisition of multi-source input data such as audio signals from electronic musical instruments, performer's limb movement data, and environmental acoustic data, various information affecting sound effects can be comprehensively captured, providing a richer and more comprehensive basis for subsequent sound effect optimization. Feature extraction and encoding of the multi-source input data generates multi-dimensional feature vectors with temporal correlation, and a dynamic voiceprint network is constructed based on this. Nodes correspond to performance events, and the connection relationships between nodes are dynamically generated based on the similarity of motion trajectories and audio phase coherence of consecutive performance events, and assigned correlation strength weights. The network structure features are then matched with a predefined sound effect template library to generate optimized sound effect parameters. This ensures that the generated sound effect parameters better match the actual performance and inherent correlations, thereby making the processed audio stream output more consistent with the performer's expressive intent and improving the accuracy and adaptability of electronic musical instrument sound effect optimization. Attached Figure Description
[0009] The accompanying drawings, which are included to provide a further understanding of embodiments of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0010] Figure 1 This is a flowchart of an optional intelligent sound effect optimization method for electronic musical instruments oriented towards multi-source input, according to an embodiment of this application;
[0011] Figure 2 A flowchart of another optional intelligent sound effect optimization method for electronic musical instruments oriented towards multi-source input according to an embodiment of this application;
[0012] Figure 3 This is a structural diagram of an optional intelligent sound effect optimization system for electronic musical instruments oriented towards multi-source input, according to an embodiment of this application.
[0013] Figure 4 This is a schematic diagram of the structure of an optional electronic device according to an embodiment of this application.
[0014] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0016] According to one aspect of the embodiments of this application, such as Figure 1As shown, this application provides a method for intelligent sound effect optimization of electronic musical instruments for multi-source input, including:
[0017] S101 collects multi-source input data of electronic musical instruments in real time, including audio signals, performer's limb movement data and environmental acoustic data.
[0018] The S101 breaks through the limitations of traditional sound effect optimization that relies solely on a single audio signal. By acquiring three types of multi-source inputs in real time—audio signals, performer's body motion data, and environmental acoustic data—it constructs a more comprehensive information acquisition dimension. Audio signals directly reflect the instrument's sound production itself; body motion data relates to the performer's expressive intentions and techniques; and environmental acoustic data reflects the spatial characteristics of the sound field. The combination of these three elements can completely capture all the elements affecting the sound effect presentation, providing multi-dimensional and comprehensive raw data support for subsequent optimization and avoiding the one-sidedness of sound effect optimization due to missing information.
[0019] S102, extract and encode features from multi-source input data to generate multi-dimensional feature vectors with temporal correlation;
[0020] In S102, targeted feature extraction and encoding are performed on multi-source heterogeneous data, with an emphasis on preserving temporal correlations. By transforming different types of raw data (audio, motion, environment) into unified multi-dimensional feature vectors, the problem of incompatibility between multi-source data formats is solved, achieving standardized data processing. The imposition of temporal correlations ensures that the feature vectors can reflect the dynamic continuity of the performance process (such as the sequential connection of notes and the coherent changes in movements), so that the extracted features not only contain information at a single point in time but also contain the temporal evolution logic of the performance, laying a data foundation for the subsequent construction of a network that can reflect the dynamics of the performance.
[0021] S103, Based on the dynamic correlation between multidimensional feature vectors, a dynamic voiceprint network is constructed; wherein, the dynamic voiceprint network consists of nodes and the connection relationships between nodes, each node corresponds to a performance event, and the node attributes include the multidimensional feature vector encoding value of the corresponding performance event; the connection relationships between nodes are dynamically generated according to the similarity of motion trajectories and audio phase coherence between consecutive performance events, and are assigned connection weight values that represent the strength of the correlation.
[0022] It should be explained that a performance event is a performance unit with independent significance in the process of playing an electronic musical instrument. Its recognition requires the comprehensive analysis of the characteristic changes in audio signals and limb movement data. By setting reasonable judgment criteria, it can be separated from the continuous data stream in order to accurately reflect each independent performance action of the performer.
[0023] For example, for an electronic keyboard, a performance event can be defined by detecting the start (attack) of the audio signal when a key is pressed and the end (tail) of the audio signal when the key is released. Simultaneously, by combining the player's finger movement data, a performance event is defined as the start of a finger movement from a key-free position to a key-pressed position, triggering the key action, and the audio signal energy exceeding a preset attack threshold (e.g., 0.01V); the performance event is defined as the end of a finger movement when the audio signal energy falls below the tail threshold (e.g., 0.005V). For electronic instruments requiring significant limb movement, such as electronic drums, a performance event is defined as the process from the initial acceleration of the drumstick to its final drop to zero after the strike, when the speed of the drumstick's trajectory exceeds a preset speed threshold (e.g., 1m / s) and the audio signal generated by striking the drumhead reaches a certain intensity.
[0024] In S103, scattered performance information is transformed into a structured set of nodes by assigning nodes to corresponding performance events and using node attributes to carry feature encoding values. The connections between nodes are dynamically generated based on motion trajectory similarity (associating with the continuity of performance actions) and audio phase coherence (associating with the physical characteristics of sound), and weight values are assigned to quantify the strength of the association. This allows the network to reflect both the inherent logic between performance events (the dual association of action and sound) and the tightness of the association through weights. This network structure intuitively presents the dynamic association patterns of performance, providing an analytical structural basis for subsequent sound effect parameter optimization.
[0025] It should be noted that motion trajectory similarity is used to measure the similarity of limb motion trajectories in two consecutive performance events. The coordinate sequences of the two trajectories are matched by the Dynamic Time Warping (DTW) algorithm to calculate the path cost. The smaller the cost, the higher the similarity, thereby quantifying the continuity of the performance actions.
[0026] For example, the three-dimensional coordinate sequence of two consecutive hand-raising movements of the performer is obtained, where the coordinate sequence of the first movement is [(x1,y1,z1),(x2,y2,z2),...,(x n ,y n ,z n The second is [(a1,b1,c1),(a2,b2,c2),...,(a m ,b m ,c m The DTW algorithm is used to construct an n×m distance matrix, where the matrix element d(i,j) represents (x...). i ,y i ,z i ) and (a j ,b j ,c j The Euclidean distance between ) is d(i,j) = √[(xi -a j )²+(y i -b j )²+(z i -c j Then, find a path from the top left corner to the bottom right corner of the matrix that minimizes the sum of distances along the path. This minimum sum of distances is the path cost. If the calculated path cost is 0.2, it means that the trajectories of the two hand-raising actions are highly similar; if the path cost is 1.5, the similarity is low.
[0027] Audio phase coherence is used to describe the consistency of phase between two consecutive audio signals in a specific frequency band. It is determined by calculating the peak value of the cross-correlation function of the signals in that frequency band. The higher the peak value, the more stable the phase relationship between the two signals in that frequency band, and the stronger the coherence.
[0028] For example, we can select the 200Hz-2000Hz frequency band to process the audio signals of two consecutive performance events. First, we perform Fourier transforms on the two signals to obtain their respective spectra in that frequency band. Then, we calculate their cross-correlation function, R(τ)=∫x(t)y(t+τ)dt, where x(t) and y(t) are the two audio signals.
[0029] `t` represents the time variable, used in the integration operation to traverse the entire time axis of the audio signal, covering all points in time from the start to the end of the signal. By integrating over `t`, the correspondence between the two audio signals at different times can be accumulated. `τ` represents the time offset, indicating the duration by which one of the audio signals, `y(t)`, is shifted along the time axis. When `τ=0`, the correlation between the two signals at the same moment is calculated; when `τ` is positive, the correlation between `y(t)` and `x(t)` is calculated after shifting `y(t)` backward by `τ`; when `τ` is negative, the correlation between `y(t)` and `x(t)` is calculated after shifting `y(t)` forward by the corresponding duration. By changing the value of `τ` and calculating the corresponding `R(τ)`, the time offset where the correlation between the two audio signals is strongest can be found, thus determining the peak value of phase coherence. For example, for audio signals x(t) and y(t) of two consecutive performance events, the cross-correlation function R(τ) reaches a peak value of 0.8 when τ = 0.002 seconds. This indicates that shifting y(t) backward by 0.002 seconds results in the strongest phase coherence with x(t) in that frequency band.
[0030] Find the peak value of the cross-correlation function. If the peak value is 0.8 near the 1000Hz frequency band, it indicates that the audio signals of the two consecutive performance events have strong phase coherence in the 1000Hz frequency band; if the peak value is 0.3, the coherence is weak.
[0031] It should be noted that the dynamic generation of connections is based on a joint judgment of motion trajectory similarity and audio phase coherence. By setting thresholds for both, a connection is established between corresponding nodes when both threshold conditions are met simultaneously, thus reflecting the inherent relationship between performance events. For example, the threshold for motion trajectory similarity is set to 0.7, and the threshold for audio phase coherence is set to 0.6. In a certain performance, the calculated motion trajectory similarity of two consecutive performance events is 0.8, exceeding the threshold of 0.7; the audio phase coherence is 0.7, exceeding the threshold of 0.6. In this case, it is determined that these two performance events have a strong relationship, and a connection is established between their corresponding nodes. If in another calculation, the motion trajectory similarity is 0.6 (below 0.7), even if the audio phase coherence is 0.8 (above 0.6), no connection is established; similarly, if the audio phase coherence is 0.5 (below 0.6), even if the motion trajectory similarity is 0.9 (above 0.7), no connection is established.
[0032] The connection weight value characterizes the strength of the connection between two nodes. It is obtained by weighted summation of motion trajectory similarity and audio phase coherence. Different weight coefficients are set according to the importance of each factor in the connection strength, so that the weight value can comprehensively reflect the correlation between the two aspects. For example, if the weight coefficient for motion trajectory similarity is set to 0.6 and the weight coefficient for audio phase coherence is set to 0.4, the connection weight value is calculated as: Connection weight value = 0.6 × Motion trajectory similarity + 0.4 × Audio phase coherence. If the motion trajectory similarity of two consecutive performance events is 0.8 and the audio phase coherence is 0.7, then the connection weight value = 0.6 × 0.8 + 0.4 × 0.7 = 0.48 + 0.28 = 0.76, indicating that the connection strength is 0.76; if the motion trajectory similarity is 0.5 and the audio phase coherence is 0.5, then the connection weight value = 0.6 × 0.5 + 0.4 × 0.5 = 0.5.
[0033] S104, Match the structural features of the dynamic voiceprint network with the predefined sound effect template library, select the basic sound effect parameter set according to the node distribution density and the connection weight value between nodes, and superimpose parameter correction terms according to the existence state of specific types of connection relationships in the dynamic voiceprint network to generate optimized sound effect parameters.
[0034] It should be explained that the sound effect template library is categorized and stored according to music style and node distribution density range. Each category contains a corresponding set of basic sound effect parameters. During matching, the range to which the node belongs is first determined based on the node distribution density. Then, the most matching set of basic sound effect parameters is selected from the corresponding music style based on the average value of the connection weights, thus achieving initial adaptation of the sound effect parameters.
[0035] For example, the sound effect template library is divided into music styles such as rock, classical, and jazz. Each style is further divided by node distribution density ranges (e.g., 1-3 nodes / second, 3-6 nodes / second, 6-9 nodes / second). When the node distribution density is 4 nodes / second (belonging to the 3-6 nodes / second range) and the average connection weight is 0.7, the library searches for a set of basic sound effect parameters within this density range in the rock style that have an average connection weight close to 0.7. This parameter set may include parameters such as distortion and equalizer frequency bands. If the node distribution density is 2 nodes / second and the average connection weight is 0.6, then it might match the parameter set corresponding to the 1-3 nodes / second range in the classical style.
[0036] In S104, a mapping mechanism between the dynamic voiceprint network and sound effect parameters was established. By matching the network structure features with a predefined template library, the basic sound effect parameter set was selected using node distribution density (reflecting the density of performance events, associated rhythm or intensity) and connection weight values (reflecting the strength of event association), ensuring the initial adaptability of the parameters. Furthermore, by superimposing parameter correction terms based on the existence of specific types of connection relationships, targeted adjustments were achieved for special performance scenarios (such as the association between specific actions and sounds). This combination of "basic parameters + correction terms" ensures that the generated optimized sound effect parameters not only fit the overall performance style but also respond to the needs of specific scenarios, improving the accuracy of parameter optimization.
[0037] S105 outputs the processed audio stream based on optimized sound effect parameters.
[0038] For example, optimized audio parameters can be input into the audio processing engine to output a processed audio stream.
[0039] In S105, as the closed-loop output of the entire method, the optimized sound effect parameters generated in the preceding steps are transformed into a perceptible audio stream, realizing the full-process value realization from data acquisition, feature processing, network construction to parameter optimization. This step ensures that all the technical results of the preprocessing are ultimately presented in the form of high-quality sound effects, so that the performer's expressive intention, the instrument's sound characteristics, and the acoustic characteristics of the environment can be harmoniously and uniformly reflected through the processed audio stream, achieving the core goal of improving the sound quality of electronic musical instruments.
[0040] Furthermore, such as Figure 2 As shown, after outputting the processed audio stream, the method further includes:
[0041] S201 analyzes the time-frequency characteristics of the processed audio stream in real time and maps them into a sound field particle model that includes energy value, spectral centroid coordinates, and time-domain density.
[0042] In some embodiments, the time-frequency characteristics of the audio stream are obtained through short-time Fourier transform (STFT), and then converted into particle energy values, spectral centroid coordinates, and temporal density, respectively. The energy value reflects the intensity of the sound, the spectral centroid coordinates reflect the timbre brightness of the sound, and the temporal density reflects the rate of change of the sound over time. The three together constitute particle properties to characterize the audio features.
[0043] For example, an STFT is performed on the processed audio stream with a window length of 512 points and an overlap rate of 50%. For a given audio frame, the root mean square of its signal energy is calculated, yielding an energy value of 0.5. The frequency range of this audio frame is divided into multiple frequency bands, and the energy of each band is calculated. A weighted average is then used to obtain the centroid coordinates of the spectrum at 1000Hz. The zero-crossing points of the waveform within this frame are detected, and the standard deviation of the interval between adjacent zero-crossing points is calculated as the dispersion, yielding a dispersion of 0.2. The time-domain density is then 1 / 0.2 = 5.
[0044] Temporal density is obtained by calculating the reciprocal of the dispersion of the zero-crossing intervals of the waveform within a frame. First, zero-crossings are detected and adjacent intervals are calculated. Then, the standard deviation of these intervals is taken as the dispersion. Finally, the reciprocal of the dispersion is the temporal density, used to reflect the variation of sound in the time dimension. For example, in a 20ms audio frame, zero-crossing times are detected at 2ms, 7ms, 12ms, and 17ms. The intervals between adjacent zero-crossings are calculated as follows: 7-2=5ms, 12-7=5ms, 17-12=5ms. The average of these intervals is 5ms, and the standard deviation is 0. Therefore, the dispersion is 0, and the temporal density is 1 / 0 = infinity (indicating that the audio segment changes very regularly in the time domain). In another audio frame, the zero-crossing times are 3ms, 8ms, 15ms, and 20ms, with intervals of 5ms, 7ms, and 5ms. The average is 5.67ms, the standard deviation is approximately 1.15ms, the dispersion is 1.15, and the temporal density is approximately 0.87.
[0045] The sound field particle model comprehensively characterizes sound field properties through three attributes: energy value, spectral centroid coordinates, and temporal density. Energy value corresponds to sound intensity, spectral centroid coordinates correspond to timbre, and temporal density corresponds to time-varying characteristics. These three attributes describe the sound field from different dimensions, and their combination comprehensively reflects the sound field status of the audio stream, providing a reliable basis for subsequent sound field balance analysis. For example, in a noisy performance environment, a high energy value (e.g., 0.8) indicates a strong sound; a spectral centroid coordinate of 3000Hz indicates a bright timbre; and a temporal density of 2 indicates a moderate rate of sound change. These three attributes clearly reveal the sound field characteristics of the audio segment. If the energy value suddenly changes to 0.3, the spectral centroid coordinate drops to 1000Hz, and the temporal density becomes 5, it indicates a significant change in the sound field, possibly a decrease in performance intensity, a darkening of the timbre, and a faster rate of change.
[0046] S202, based on the preset sound field balance rules, performs correction operations on sound field particles that deviate from the balance state;
[0047] S203, when the number of consecutive failures of the correction operation exceeds the preset threshold, the dynamic voiceprint network reconstruction mechanism is triggered; wherein, the reconstruction mechanism includes: resetting the connection weight values between nodes, updating the network structure of the dynamic voiceprint network, and re-executing the steps of matching the sound effect template and generating the sound effect parameters.
[0048] The preset number of times threshold may include, but is not limited to, 5 times, 10 times, etc.
[0049] Based on the embodiments provided in this application, a closed-loop control for sound effect optimization is formed by introducing a real-time analysis and dynamic adjustment mechanism after the output processed audio stream. The time-frequency characteristics of the processed audio stream are mapped to a sound field particle model containing energy values, spectral centroid coordinates, and time-domain density, making the abstract audio features concrete into quantifiable and analyzable particle units, facilitating accurate identification of sound effect deviations from the equilibrium state. When particles deviate from the equilibrium, they are first directly adjusted through correction operations. When continuous correction failures exceed a threshold, dynamic voiceprint network reconstruction is triggered, fundamentally optimizing the sound effect generation logic through operations such as resetting connection weights and updating the network structure. This dual-layer adjustment mode of real-time correction plus deep reconstruction can effectively cope with sudden and continuous sound effect deviations during performance, ensuring that the sound effects always remain adapted to the performance scene and intention, thus improving the robustness of sound effect optimization.
[0050] Furthermore, specific types of connections refer to dynamic feedback channels triggered by environmental reverberation interference and performance transient events;
[0051] When two consecutive performance events meet the following conditions: the rate of change of audio phase coherence exceeds the transient judgment threshold, and the energy ratio of ambient direct sound to reverberant sound is lower than the reverberation interference threshold, a feedback channel is established between the corresponding nodes of the two performance events.
[0052] It should be explained that the rate of change of audio phase coherence is used to determine the degree of change in phase correlation between two consecutive performance events. It is obtained by dividing the difference in phase coherence between the current event and the previous event by the phase coherence value of the previous event, thereby determining whether a significant transient has occurred in the performance.
[0053] Its calculation formula is: ;in, The rate of change of audio phase coherence. The audio phase coherence of the current performance event (value range 0-1). The audio phase coherence of the previous performance event.
[0054] It should be noted that the impulse response based on environmental acoustic data separates the direct sound and reverberant sound energy. By setting a time threshold, the early energy (direct sound) and the later energy (reverberant sound) of the impulse response are divided, and then the ratio of the environmental direct sound energy to the reverberant sound energy is calculated.
[0055] Among them, the transient judgment threshold and the reverberation interference threshold are both configurable parameters. The empirical range is set in combination with the type of electronic instrument (such as electric guitar, electronic drum) and common performance environment (such as indoor and outdoor) to take into account both versatility and scene adaptability.
[0056] For example, the transient detection threshold ranges from 0.2 to 0.4 (absolute value). For instance, a threshold of 0.3 is used in electric guitar playing because rapid strumming easily produces transients; a threshold of 0.2 is suitable for electronic keyboard playing, as transients are smoother. The reverberation interference threshold ranges from 0.4 to 0.6 (direct sound / reverberation energy ratio). For example, a threshold of 0.5 is used in halls with strong acoustic reflections, while a threshold of 0.6 is used in well-insulated rooms to avoid excessive reverberation suppression.
[0057] The existence status of a specific type of connection is determined by the number of active feedback channels in the dynamic voiceprint network; wherein, the feedback channel remains active for a preset duration after its establishment and is automatically deactivated after the timeout.
[0058] The preset duration is a fixed value, starting from the moment the channel is established, to ensure continuous response to short-term interference. It automatically deactivates after the timeout to avoid unnecessary resource consumption. For example, the preset duration is set to 2 seconds (based on the average interval of performance events). When the feedback channel is established at t=10s, the active state lasts until t=12s, after which it automatically deactivates. The timing is implemented through the system clock and synchronized with the timestamp of the performance event.
[0059] The superposition parameter correction items include: counting the number of feedback channels in the active state; when the number of feedback channels in the active state does not exceed the preset number, reducing the reverberation decay time to a first proportion of the reference value, increasing the first gain value in the preset high frequency band, and reducing the dynamic compression start threshold by a first amount; when the number of feedback channels in the active state exceeds the preset number, reducing the reverberation decay time to a second proportion of the reference value, increasing the second gain value in the preset high frequency band, and reducing the dynamic compression start threshold by a second amount.
[0060] The preset high-frequency band refers to a range of high frequencies pre-defined based on the type of electronic instrument and common sound effect requirements. It is used to specifically enhance high-frequency components and improve sound clarity during environmental reverberation interference or transient changes in performance. The division of the preset high-frequency band is based on the instrument's sound characteristics and the human ear's sensitivity range to high-frequency details (usually 2kHz-8kHz). It can be dynamically adjusted according to the type of instrument (such as electric guitar, electronic synthesizer) to ensure that the enhanced high frequencies match the instrument's own tonal characteristics.
[0061] For example, the default preset high frequency band is 2kHz-5kHz (covering the area of human ear sensitive to speech and musical instrument overtones).
[0062] Instrument-specific adaptation: For electric guitar playing, expand to 3kHz-7kHz (enhancing high-frequency harmonics of distorted tones); for electronic piano playing, reduce to 2kHz-4kHz (avoiding excessively strong high frequencies that can be harsh). Combined with parameter correction: When the number of active feedback channels exceeds a preset number (e.g., 3), add a second gain of 6dB in the 2kHz-5kHz frequency band to compensate for high-frequency attenuation caused by reverberation, ensuring clear harmonics in guitar solos or the high register of the piano.
[0063] High frequencies are easily attenuated by ambient reverberation (reverberation absorbs high frequencies more strongly), and the details of transient performance changes (such as rapid plucking and keystrokes) are mainly reflected in high-frequency changes. By preset high-frequency bands and targeted gain, high-frequency losses in these scenarios can be accurately compensated, avoiding low-frequency muddiness caused by global gain.
[0064] It should be explained that the correction parameters are set in stages according to the number of activated feedback channels. The more channels there are (the more severe the interference), the greater the correction magnitude. The interference is offset by the coordinated adjustment of the ratio, gain, and compression threshold.
[0065] For example, the first ratio = 80%, the second ratio = 50% (the second ratio < the first ratio, to enhance reverberation attenuation); the first gain value = 3dB, the second gain value = 6dB (the second gain > the first gain, to enhance high-frequency penetration); the first amplitude = 2dB, the second amplitude = 5dB (the second amplitude > the first amplitude, to reduce the compression threshold to suppress burst noise).
[0066] For example, with 4 active channels (more than the preset number of 3), the reverberation decay time is reduced to 50% of the baseline value, the high-frequency gain (2-5kHz) is increased by 6dB, and the compression trigger threshold is reduced by 5dB. The number of active channels reflects the superposition intensity of ambient reverberation and performance transients; a larger number indicates more complex interference, requiring more aggressive parameter adjustments to maintain sound clarity.
[0067] Based on the embodiments provided in this application, by explicitly defining specific types of connection relationships as dynamic feedback channels triggered by environmental reverberation interference and performance transient events, and by superimposing differentiated parameter correction terms based on their existence states, a precise response to complex performance scenarios is achieved. The ingenuity lies in quantifying and associating environmental interference (reverberation) and performance dynamics (transient events) through the network structure feature of the feedback channels, transforming abstract interference factors into statistically measurable network parameters (the number of active feedback channels). Simultaneously, by setting different proportions of reverberation decay time, gain value, and compression threshold according to the number of channels, the adjustment of sound effect parameters dynamically matches the interference intensity. Parameters are fine-tuned when interference is small, and the adjustment range is increased when interference is large, thus maintaining the clarity and expressiveness of the sound effect even under complex acoustic environments and changing performance dynamics.
[0068] Furthermore, the reconstruction mechanism updates the network structure of the dynamic voiceprint network, including:
[0069] Scan all connections and remove connections with weight values lower than the preset failure threshold based on the reset connection weight values between nodes.
[0070] It's important to clarify that resetting the connection weights between nodes means recalculating the weights based on the latest performance event features, rather than resetting them to zero or randomly initializing them. This ensures that the weights still reflect the true correlation between events while eliminating historical accumulated errors. For example, during the reset, the similarity of motion trajectories and audio phase coherence are recalculated for each pair of nodes.
[0071] The preset failure threshold is set based on the statistical distribution of connection weights to filter weak connections and simplify the network structure. The value is typically 1 / 3 to 1 / 2 of the average weight. For example, if the historical average connection weight is 0.6, the preset failure threshold is set to 0.2 (approximately 1 / 3 of the average). If a connection's weight is 0.15 (<0.2) after a reset, it is considered failed and removed.
[0072] Identify newly generated performance events after the most recent construction or update of the dynamic voiceprint network, and the performance event does not yet have a corresponding node in the dynamic voiceprint network; create a new node for the identified performance event, and calculate the multidimensional feature vector encoding value of the node as a node attribute;
[0073] New performance events are identified by comparing timestamps with features. The timestamp of the most recent network update is recorded. Events that occur after this timestamp and whose features (such as audio fingerprints or motion trajectories) do not match those of existing nodes are identified as new performance events. For example, if the most recent network update time is t=20s, and a performance event is detected at t=21s, and the zero-crossing phase feature of its audio signal matches the features of existing nodes with a degree of <0.5 (preset matching threshold), then it is identified as a new performance event, and a node is created for it.
[0074] Based on the similarity of motion trajectories and audio phase coherence between the newly added nodes and existing nodes, new connections are established and assigned initial weight values.
[0075] Based on the embodiments provided in this application, a three-step operation of "removing low-weight connections - adding nodes - establishing new connections" ensures that the reconstructed network can both streamline redundant information and incorporate the latest performance events. In other words, connections with weights below the failure threshold are removed to prevent invalid information from interfering with network analysis; nodes are created and features are calculated for new performance events to ensure the network captures the latest performance dynamics; and new connections are established based on motion trajectory similarity and audio phase coherence, enabling new nodes to form meaningful associations with existing nodes. This structured update method keeps the reconstructed network in a "streamlined and complete" state, reducing computational burden while accurately reflecting the internal logic of the current performance, thus improving the network's adaptability to continuously changing performance processes.
[0076] Furthermore, the time-frequency characteristics of the processed audio stream are analyzed in real time and mapped to a sound field particle model that includes energy values, spectral centroid coordinates, and time-domain density, including:
[0077] The processed audio stream is divided into continuous analysis frames with a fixed duration, and each frame corresponds to a basic particle unit.
[0078] The fixed duration needs to balance time resolution and computational efficiency, and is typically selected to be 20-30ms. A 50% overlap rate is also set to avoid information breaks between frames. For example, the fixed duration can be set to 20ms, with a 10ms start time interval for each frame (10ms overlap). For instance, the first frame covers 0-20ms, and the second frame covers 10-30ms, ensuring a smooth transition of continuous audio.
[0079] For each basic particle unit, feature quantization is performed, including: calculating the root mean square value of the signal energy within the frame as the particle energy value; dividing the energy distribution of the critical frequency band and solving the weighted average frequency of the energy in each frequency band as the centroid of the particle spectrum; and calculating the inverse of the dispersion of the zero-crossing interval of the waveform within the frame as the particle temporal density.
[0080] In some embodiments, the Bark scale is used to divide the critical frequency bands because the Bark scale is highly matched with the frequency selectivity of the human auditory system, and can more accurately reflect the human ear's perception of energy in different frequency bands. For example, the 20Hz-20kHz range is divided into 24 critical frequency bands according to the Bark scale, with the first band corresponding to 20-100Hz, the tenth band corresponding to 1.5-1.7kHz, and so on, and the signal energy within each frequency band is calculated.
[0081] In some embodiments, the centroid of the spectrum is obtained by weighted summation of the center frequency of the frequency band and the corresponding energy, and then divided by the total energy, thus quantifying the "brightness" of the sound (the higher the proportion of high-frequency energy, the more the centroid is to the right).
[0082] In one specific implementation, the spectral centroid Determined based on the following formula:
[0083]
[0084] in, For the first The center frequency of each frequency band For the first Energy of each frequency band This represents the total number of frequency bands.
[0085] The process involves calculating the interval by detecting zero-crossing moments, and using the standard deviation of the interval to characterize the dispersion. The reciprocal of this standard deviation reflects the density of time-domain changes (the smaller the dispersion, the larger the reciprocal, indicating more regular and dense signal changes). For example, zero-crossing detection involves traversing the audio waveform within a frame and recording the moments when the signal crosses zero (e.g., t1=5ms, t2=10ms, t3=15ms). Interval calculation: The interval between adjacent zero-crossing moments is Δt1=t2-t1=5ms, Δt2=t3-t2=5ms. Dispersion (standard deviation): The standard deviation of the interval sequence is calculated; the reciprocal of the dispersion is 1 / σ (maximum value when σ=0, e.g., 1000). A larger value indicates a more stable zero-crossing interval and a more dense change in the signal in the time domain (e.g., rapidly repeating drumbeats).
[0086] A unique particle identifier is generated for each basic particle unit, and its energy value, spectral centroid coordinates, and time-domain density are bound to it to form an initial particle set.
[0087] It should be noted that the expression "spectral centroid coordinates" in this application does not contradict the essence of "spectral centroid being a single-frequency value," but rather is a concrete description based on the "sound field particle model." Here, "coordinates" do not refer to two-dimensional or three-dimensional coordinates in geometric space, but rather to analogizing the frequency axis to a one-dimensional coordinate axis. The specific frequency value of the spectral centroid (e.g., 1500Hz) corresponds to the "position coordinates" of the particle on this frequency axis. This expression aims to make the abstract audio frequency characteristics more closely fit the metaphor of the "particle" model, facilitating the understanding of the particle's distribution in the frequency dimension. Just as physical particles have position coordinates in space, audio particles also have their corresponding "coordinate" positions on the frequency axis.
[0088] From a technical perspective, the calculation of the spectral centroid is always a single-frequency value obtained by weighted averaging of the center frequency of the frequency band and the energy. Its physical meaning is the "center of gravity" of the audio signal energy on the frequency axis. Calling it a "coordinate" is only to adapt to the construction logic of the "sound field particle model" so that the particle attributes (energy value, spectral centroid coordinates, temporal density) form a unified "particle characteristic system". This makes it convenient to judge whether the spectral distribution is balanced by the "coordinate difference" between particles. For example, comparing whether the difference between the spectral centroid coordinates of a certain particle and the dominant particle exceeds the dynamic radius threshold.
[0089] The core purpose of this approach is to enhance the intuitiveness and operability of the model, rather than altering the essence of the spectral centroid. Whether it's a "single frequency point value" or "spectral centroid coordinates," they both refer to the same physical quantity—the key frequency value characterizing the frequency distribution of the audio signal. This unified approach accurately reflects the technical principles and makes the analytical logic of the "sound field particle model" (such as particle clustering and balance determination) easier to understand and implement, avoiding the hindrance that abstract concepts can pose to the practical application of technical solutions.
[0090] Based on the numerical continuity of particle temporal density, adjacent initial particles with a density difference less than a preset density threshold are aggregated into a particle cluster.
[0091] The preset density threshold is set based on the statistical distribution of temporal density and is used to aggregate particles with similar characteristics. The threshold is 10%-20% of the average density. For example, if the average temporal density of historical particles is 2.0, the preset density threshold is set to 0.3 (approximately 15% of the average). If particle A has a density of 1.8 and particle B has a density of 2.0, the difference 0.2 < 0.3, and they will aggregate into a cluster.
[0092] Within each particle cluster, the particle with the highest energy value is selected as the dominant particle;
[0093] It should be noted that the particle with the highest energy best represents the overall intensity characteristics of the particle cluster. Using this as the dominant particle can simplify the subsequent sound field balance analysis and ensure that adjustments are made based on the strongest energy characteristics.
[0094] For example, if a particle cluster contains three particles with energies of 0.5, 0.8, and 0.6, the particle with energy of 0.8 is selected as the dominant particle, and its spectral centroid and temporal density are used as representative features of the cluster.
[0095] Output the set of all particle clusters and their dominant particles as the sound field particle model.
[0096] Based on the embodiments provided in this application, an accurate sound field particle model is constructed by segmenting the audio stream into analysis frames, quantizing features to generate particles, clustering them into clusters, and selecting dominant particles. The core ingenuity lies in transforming the continuous audio stream into a discrete model with a hierarchical structure (particle-particle cluster-dominant particle): the basic particle unit retains the detailed features of each frame of audio (energy, spectral centroid, temporal density); particle clusters aggregate similar particles through temporal density continuity, reducing redundancy; and the dominant particle extracts the core features within the cluster. This modeling approach preserves the microscopic features of the audio while achieving macroscopic-level pattern extraction through clustering, enabling subsequent sound field analysis and adjustment to be accurate down to the frame level and grasp the overall trend, providing a precise and efficient analytical foundation for sound field balance control.
[0097] Furthermore, the preset sound field balance rules include: taking the spectral centroid of the dominant particle in the particle cluster as the reference value, the absolute difference between the spectral centroid of other particles in the cluster and the reference value shall not exceed the dynamic radius threshold.
[0098] The initial value of the dynamic radius threshold is set based on the common frequency distribution range of the audio signal, which is used to define the normal fluctuation range of the spectral centroid within the particle cluster; the preset ratio adjusts the threshold according to environmental characteristics (strong reflection or high noise) so that the sound field balance judgment standard can adapt to different environments.
[0099] For example, the initial dynamic radius threshold is set to 200Hz, which takes into account the natural fluctuation range of the spectral centroid during the performance of most electronic musical instruments. In a strong reflection environment, the preset ratio is set to 80%, that is, the dynamic radius threshold is reduced to 200 × 80% = 160Hz, in order to strictly control the spectral shift and reduce the timbre deviation caused by reflection interference; in a high noise environment, the preset ratio is set to 120%, that is, the dynamic radius threshold is expanded to 200 × 120% = 240Hz, allowing for a wider range of spectral fluctuations and avoiding misjudgment of imbalance due to noise interference.
[0100] Based on preset sound field balance rules, correction operations are performed on sound field particles that deviate from the equilibrium state, including:
[0101] Traverse all particle clusters in the sound field particle model; calculate the absolute difference between the centroid of the spectrum of each particle within the cluster and the reference value;
[0102] If the absolute difference exceeds the dynamic radius threshold, a compensatory equalization filter parameter is generated based on the offset direction and applied to the original audio frame of the particle; with the dominant particle energy value as the target, dynamic range compression is applied to the original audio frame of the particle.
[0103] It needs to be explained that the generation of the compensation equalization filter parameters includes: determining the adjustment direction of the equalization filter based on the offset direction (too high or too low) between the particle spectrum centroid and the dominant particle spectrum centroid, and then calculating specific parameters (center frequency, bandwidth, gain) in combination with the magnitude of the offset to achieve targeted spectrum correction.
[0104] For example, if the centroid of the particle's spectrum is 300Hz higher than that of the dominant particle (the offset direction is high frequency), the generated equalization filter parameters are as follows: the center frequency is set to the centroid value of the particle's spectrum (e.g., if the dominant particle's frequency is 1000Hz and this particle's is 1300Hz, then the center frequency is 1300Hz), the bandwidth is set to 200Hz based on the offset (the larger the offset, the larger the bandwidth should be to cover more offset frequency bands), and the gain is set to -3dB (a negative value indicates attenuation of high frequencies). If the offset direction is low frequency (e.g., the centroid of the particle's spectrum is 800Hz and the dominant particle's is 1000Hz), then the center frequency is 800Hz, the bandwidth is 200Hz, and the gain is set to +2dB (a positive value indicates enhancement of low frequencies).
[0105] The application of dynamic range compression includes: targeting the energy value of the dominant particle, by setting parameters such as the compressor threshold, ratio, start time, and release time, compressing the energy of the deviating particles to a level close to the energy of the dominant particle, thus ensuring the consistency of energy within the particle cluster.
[0106] For example, if the dominant particle's energy value is 0.6 (root mean square energy), then the compressor's threshold is set to 0.6, the ratio is set to 2:1 (i.e., energy exceeding the threshold is compressed by half), the start-up time is set to 5ms (for rapid response to energy changes), and the release time is set to 50ms (to avoid sudden energy changes). When a particle's energy value is 0.8 (above the threshold), after compression, its energy value becomes 0.6 + (0.8 - 0.6) / 2 = 0.7, which is closer to the dominant particle's energy. If the particle's energy value is 0.4 (below the threshold), compression is not initiated, preserving the original energy to avoid overprocessing.
[0107] The dynamic radius threshold is updated based on environmental acoustic data, including: if in a strong reflection environment, the dynamic radius threshold is reduced to a preset proportion of the original radius threshold; if in a high noise environment, the dynamic radius threshold is expanded to a preset proportion of the original radius threshold.
[0108] Specifically, when the measured ambient reverberation time exceeds the first set threshold, it is determined to be a strong reflection environment; when the background noise sound pressure level exceeds the second set threshold, it is determined to be a high noise environment.
[0109] The first threshold (strong reflection judgment) is set based on the common range of environmental reverberation time. Exceeding this value indicates that the environmental reflection is strong. The second threshold (high noise judgment) is set based on the human ear's perception threshold of background noise. Exceeding this value indicates that the environmental noise has a significant impact on the sound effect.
[0110] For example, the first threshold is set at 1.2 seconds. When the measured reverberation time is 1.5 seconds (exceeding 1.2 seconds), it is determined to be a strong reflection environment. The second threshold is set at 60 dBSPL (sound pressure level). When the background noise sound pressure level is 65 dBSPL (exceeding 60 dBSPL), it is determined to be a high-noise environment. These thresholds refer to typical classification standards for acoustic environments to ensure the rationality of environmental judgments.
[0111] Based on the embodiments provided in this application, by setting a dynamic radius threshold based on the spectral centroid as the sound field balance rule, and dynamically adjusting the threshold in conjunction with environmental acoustic data, adaptive balance of sound effects in different environments is achieved. The principle is as follows: using the spectral centroid of the dominant particle as a benchmark, the consistency of the sound field is ensured by controlling the spectral offset range of other particles within the cluster; when the offset exceeds the limit, targeted correction is performed through equalization filtering and dynamic compression; simultaneously, depending on whether the environment has strong reflections (reverberation time exceeding the threshold) or high noise (background sound pressure level exceeding the threshold), the dynamic radius threshold is reduced or increased accordingly to match the balance standard with environmental characteristics. This combined strategy of "benchmark control + dynamic threshold + environmental adaptation" allows the sound effects to maintain inherent consistency in complex acoustic environments while adapting to environmental interference, thus improving the environmental adaptability of the sound effects.
[0112] Furthermore, feature extraction and encoding are performed on the multi-source input data to generate multi-dimensional feature vectors with temporal correlation, including:
[0113] Zero-crossing phase-locked analysis is performed on the audio signal to generate a first-dimensional feature characterizing the transient accuracy of the attack signal;
[0114] The zero-crossing phase-locking analysis includes: comparing the expected attack time of the note triggered by the performer's movement with the actual zero-crossing time of the audio signal, calculating the statistical characteristics of the time difference (such as mean deviation and standard deviation), and generating features characterizing the transient accuracy of the attack, reflecting the synchronization between the performance action and the sound production. For example, based on the performer's limb movement data (such as key presses), the expected attack time is predicted to be t0, and the corresponding zero-crossing times in the audio signal are detected as t1, t2, ..., tn. The time difference Δt = ti - t0 is calculated. If the average value of Δt measured multiple times is 0.002 seconds and the standard deviation is 0.0005 seconds, then (average value + 2 × standard deviation) = 0.003 seconds is taken as the characteristic value of the transient accuracy of the attack. The smaller this value, the more accurate the transient accuracy of the attack.
[0115] The curvature radius of the three-dimensional trajectory of the performer's limb movement data is calculated to generate a second-dimensional feature representing the complexity of the movement.
[0116] The calculation of the radius of curvature of the three-dimensional trajectory includes: calculating the instantaneous tangent vector and normal vector based on continuous three-dimensional coordinate points of limb movement, and then obtaining the instantaneous curvature. Its reciprocal is the instantaneous radius of curvature, used to characterize the degree of bending of the movement and reflect its complexity. For example, obtaining three continuous three-dimensional points P1(x1,y1,z1), P2(x2,y2,z2), and P3(x3,y3,z3) of limb movement, and calculating vectors P1P2=(x2-x1,y2-y1,z2-z1) and P2P3=(x3-x2,y3-y2,z3-z2). The normal vector is calculated through the cross product of the vectors, and then the curvature formula K=|P1P2×P2P3| / |P1P2| is applied. 3 We obtain the curvature K, and the instantaneous radius of curvature R = 1 / K. If we calculate K = 0.5 / m, then R = 2m. The smaller the radius of curvature, the more complex the action (such as a rapid turn).
[0117] Autocorrelation attenuation slope is extracted from environmental acoustic data to generate a third-dimensional feature characterizing spatial reflection intensity.
[0118] The extraction of autocorrelation attenuation slope includes: performing autocorrelation analysis on the impulse response of environmental acoustic data to obtain the envelope of the autocorrelation function; taking the logarithm of the envelope and then performing linear fitting; the slope of the fitted line is the autocorrelation attenuation slope; the larger the absolute value of the slope, the stronger the spatial reflection intensity.
[0119] Based on the embodiments provided in this application, targeted feature extraction dimensions are designed for multi-source input data. These dimensions extract the accuracy of the attack transient from the audio signal (first dimension), the complexity of movement from limb motion data (second dimension), and the spatial reflection intensity from environmental acoustic data (third dimension). The ingenuity lies in the fact that each dimension accurately corresponds to a key factor affecting the sound effect: the attack transient directly relates to the initial clarity of the sound effect, the complexity of movement reflects the intensity of the performer's expressive intent, and the spatial reflection intensity determines the degree of environmental influence on the sound effect. This ensures that the generated multi-dimensional feature vectors can comprehensively and deeply depict the core information of the performance scene, providing high-quality foundational data for the subsequent construction of a dynamic voiceprint network and enhancing the feature vectors' ability to represent the performance scene.
[0120] Furthermore, the method also includes:
[0121] If the dynamic voiceprint network reconstruction mechanism is triggered more than three times within a preset time period, the node creation, connection relationship update and weight value reset operations will be stopped, and the audio processing engine will be switched to the pre-stored static audio effect parameter set.
[0122] The preset duration is set based on a reasonable judgment cycle for continuous anomalies in the performance scenario. It is used to define the time range within which the reconstruction mechanism is frequently triggered in a short period of time, so as to avoid the system wasting resources in a continuous abnormal state. For example, if the preset duration is set to 30 seconds, if the dynamic voiceprint network reconstruction mechanism is triggered more than three times in a row within 30 seconds, it means that the current dynamic optimization logic can no longer adapt to the scenario, and the system switches to the static sound effect parameter set to ensure basic sound effect output.
[0123] The static sound effect parameter set comes from factory default settings (optimized based on common performance scenarios) and user-defined presets. It contains core parameters that ensure basic sound quality and serves as a backup in case dynamic optimization fails. For example, the sources include three default parameter sets pre-stored at the factory (corresponding to indoor, outdoor, and small stage scenarios respectively), as well as user-defined parameters saved through the device interface. The content covers basic parameters such as reverb type (e.g., hall reverb, room reverb), reverb time (e.g., 1.0 second), equalizer band gain (e.g., 100Hz +1dB, 1kHz 0dB, 10kHz -1dB), and dynamic compression threshold (e.g., -12dB).
[0124] Cyclic redundancy check codes are embedded when generating the multidimensional feature vector corresponding to each performance event; if the check fails, the original input data of the current performance event is discarded, and the creation of nodes in the dynamic voiceprint network for that performance event is prohibited.
[0125] In some embodiments, a checksum is generated using a standard CRC algorithm and embedded into a specific position of a multidimensional feature vector to detect errors in the feature vector during transmission or processing, ensuring data integrity and preventing erroneous data from affecting the dynamic voiceprint network.
[0126] For example, using the CRC-32 algorithm, a checksum is calculated on the numerical sequence of a multi-dimensional feature vector (e.g., 0.7 for the first dimension, 0.6 for the second, and 0.5 for the third), resulting in a 32-bit binary number (e.g., 0x12345678). This checksum is appended to the end of the feature vector (e.g., the vector becomes [0.7, 0.6, 0.5, 0x12345678]). During verification, the CRC-32 value of the main feature vector is recalculated. If it does not match the appended checksum, the verification fails. Verification failure indicates that the data may be corrupted due to transmission interference or hardware failure. Discarding the data and prohibiting node creation can prevent erroneous information from entering the dynamic voiceprint network and prevent network model distortion.
[0127] Based on the embodiments provided in this application, the stability and reliability of the system are enhanced by setting a dual guarantee mechanism of "continuous reconstruction of static parameters more than three times" and "feature vector embedding of verification codes". When continuous reconstruction occurs more than three times, it indicates that the current dynamic optimization mechanism may not be suitable for the scenario. Switching to the pre-stored static parameter set can prevent the system from continuously failing and ensure basic sound effect output. Embedding verification codes in the feature vector can promptly detect invalid or erroneous performance event data. By discarding data and prohibiting the creation of nodes, erroneous information is prevented from polluting the dynamic voiceprint network. This "fault switching + data verification" design reduces the interference of abnormal situations on sound effect optimization from both system fault tolerance and data source perspectives, and improves the system's stable operation capability under complex working conditions.
[0128] Furthermore, the node distribution density is obtained by calculating the number of valid nodes within a unit time window, where the criteria for determining a valid node is that the multidimensional feature vector encoding value corresponding to the node exceeds a preset dynamic threshold.
[0129] The preset dynamic threshold is set based on the statistical distribution of historical multidimensional feature vector encoding values. It is used to distinguish between valid performance events (significant actions and sounds) and invalid interference (such as minor accidental touches or environmental noise), ensuring the accuracy of node distribution density calculation.
[0130] For example, the multidimensional feature vector encoding values of the past 1000 performance events are statistically analyzed, yielding an average of 0.5 and a standard deviation of 0.15. A preset dynamic threshold is set to the average plus one standard deviation, i.e., 0.5 + 0.15 = 0.65. If the multidimensional feature vector encoding value of a node is 0.7 (above 0.65), it is considered a valid node; if the encoding value is 0.5 (below 0.65), it is considered an invalid node and is not included in the distribution density statistics.
[0131] Based on the embodiments provided in this application, node distribution density is defined by the "number of valid nodes whose feature encoding values exceed a preset dynamic threshold within a unit time window," enabling this parameter to accurately reflect the density of meaningful performance events per unit time. In other words, only nodes whose feature values exceed the dynamic threshold are considered valid nodes, excluding interference from weak or meaningless performance events; through unit time window statistics, the density value reflects changes in the rhythm or intensity of the performance. This quantification method allows node distribution density to accurately reflect the activity level of performance events while avoiding interference from invalid information, providing an accurate quantitative basis for selecting the basic sound effect parameter set when matching with the sound effect template library, thus improving the accuracy of parameter selection.
[0132] According to another aspect of the embodiments of this application, an intelligent sound effect optimization system for electronic musical instruments oriented towards multi-source input is also provided. For example... Figure 3 As shown, the system includes:
[0133] The multi-source data acquisition module 301 is used to acquire multi-source input data of electronic musical instruments in real time. The multi-source input data includes audio signals, performer's limb movement data and environmental acoustic data.
[0134] The feature extraction module 302 is used to extract and encode features from multi-source input data to generate multi-dimensional feature vectors with temporal correlation.
[0135] The voiceprint network construction module 303 is used to construct a dynamic voiceprint network based on the dynamic correlation between multidimensional feature vectors. The dynamic voiceprint network consists of nodes and the connection relationships between nodes. Each node corresponds to a performance event, and the node attributes include the multidimensional feature vector encoding value of the corresponding performance event. The connection relationships between nodes are dynamically generated based on the similarity of motion trajectories and audio phase coherence between consecutive performance events, and are assigned connection weight values that represent the strength of the correlation.
[0136] The matching optimization module 304 is used to match the structural features of the dynamic voiceprint network with the predefined sound effect template library, select the basic sound effect parameter set according to the node distribution density and the connection weight value between nodes, and superimpose parameter correction terms according to the existence state of specific types of connection relationships in the dynamic voiceprint network to generate optimized sound effect parameters.
[0137] Output module 305 is used to output the processed audio stream based on optimized sound effect parameters.
[0138] It should be noted that the embodiments implemented in this application for the intelligent sound effect optimization system for electronic musical instruments with multi-source input can be referenced in conjunction with the embodiments implemented in the intelligent sound effect optimization method for electronic musical instruments with multi-source input, and will not be described in detail here.
[0139] According to another aspect of the embodiments of this application, an electronic device for implementing the above-described intelligent sound effect optimization method for electronic musical instruments oriented towards multi-source input is also provided. This electronic device may be... Figure 4 The terminal device or server shown. This embodiment uses this electronic device as an example of a server. Figure 4 As shown, the electronic device includes a memory 402, a processor 404, and a transmission device 406. The memory 402 stores a computer program, and the processor 404 is configured to execute the steps in any of the above method embodiments through the computer program.
[0140] Optionally, in this embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.
[0141] Optionally, the transmission device 406 is used to receive or send data via a network. Specific examples of the network described above may include wired and wireless networks. In one example, the transmission device 406 includes a Network Interface Controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In another example, the transmission device 406 is a Radio Frequency (RF) module used to communicate with the Internet wirelessly.
[0142] In addition, the aforementioned electronic device also includes: a display 408 for displaying target identification characters contained in the identity identifier of the identified target object; and a connection bus 410 for connecting various module components in the aforementioned electronic device.
[0143] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A method for intelligent sound effect optimization of electronic musical instruments oriented towards multi-source input, characterized in that, include: Real-time acquisition of multi-source input data from electronic musical instruments, including audio signals, performer's limb movement data, and environmental acoustic data; Feature extraction and encoding are performed on the multi-source input data to generate a multi-dimensional feature vector with temporal correlation; A dynamic voiceprint network is constructed based on the dynamic correlation between the multidimensional feature vectors. The dynamic voiceprint network consists of nodes and the connection relationships between nodes. Each node corresponds to a performance event, and the node attributes include the multidimensional feature vector encoding value of the corresponding performance event. The connection relationships between nodes are dynamically generated based on the similarity of motion trajectories and audio phase coherence between consecutive performance events, and are assigned connection weight values that represent the strength of the correlation. The structural features of the dynamic voiceprint network are matched with a predefined sound effect template library. The basic sound effect parameter set is selected based on the node distribution density and the connection weight values between nodes. Parameter correction terms are superimposed based on the existence state of specific types of connection relationships in the dynamic voiceprint network to generate optimized sound effect parameters. The processed audio stream is output based on the optimized sound effect parameters.
2. The intelligent sound effect optimization method for electronic musical instruments oriented towards multi-source input as described in claim 1, characterized in that, Following the output of the processed audio stream, the method further includes: The time-frequency characteristics of the processed audio stream are analyzed in real time and mapped to a sound field particle model that includes energy value, spectral centroid coordinates, and time-domain density. Based on the preset sound field balance rules, correction operations are performed on sound field particles that deviate from the balance state; When the number of consecutive failures in the correction operation exceeds a preset threshold, the dynamic voiceprint network reconstruction mechanism is triggered. The reconstruction mechanism includes: resetting the connection weight values between nodes, updating the network structure of the dynamic voiceprint network, and re-executing the steps of matching sound effect templates and generating sound effect parameters.
3. The intelligent sound effect optimization method for electronic musical instruments oriented towards multi-source input as described in claim 1, characterized in that, The specific type of connection relationship refers to the dynamic feedback channel triggered by environmental reverberation interference and performance transient events; When two consecutive performance events meet the following conditions: the rate of change of audio phase coherence exceeds the transient judgment threshold, and the energy ratio of ambient direct sound to reverberant sound is lower than the reverberation interference threshold, a feedback channel is established between the corresponding nodes of the two performance events. The existence state of the specific type of connection is determined by the number of feedback channels in the dynamic voiceprint network that are active; wherein, the feedback channel remains active for a preset duration after being established, and is automatically deactivated after the timeout. The superposition parameter correction items include: counting the number of feedback channels in the active state; when the number of feedback channels in the active state does not exceed a preset number, reducing the reverberation decay time to a first proportion of the reference value, increasing the first gain value in the preset high-frequency band, and reducing the dynamic compression start threshold by a first magnitude; when the number of feedback channels in the active state exceeds the preset number, reducing the reverberation decay time to a second proportion of the reference value, increasing the second gain value in the preset high-frequency band, and reducing the dynamic compression start threshold by a second magnitude.
4. The intelligent sound effect optimization method for electronic musical instruments oriented towards multi-source input as described in claim 2, characterized in that, The reconstruction mechanism involves updating the network structure of the dynamic voiceprint network, including: Scan all connections and remove connections with weight values lower than the preset failure threshold based on the reset connection weight values between nodes. Identify newly generated performance events after the most recent construction or update of the dynamic voiceprint network, and the performance event does not yet have a corresponding node in the dynamic voiceprint network; create a new node for the identified performance event, and calculate the multidimensional feature vector encoding value of the node as a node attribute; Based on the similarity of motion trajectories and audio phase coherence between the newly added nodes and existing nodes, new connections are established and assigned initial weight values.
5. The intelligent sound effect optimization method for electronic musical instruments oriented towards multi-source input as described in claim 2, characterized in that, The time-frequency characteristics of the audio stream after real-time analysis and processing are mapped to a sound field particle model including energy value, spectral centroid coordinates, and time-domain density, including: The processed audio stream is divided into continuous analysis frames with a fixed duration, and each frame corresponds to a basic particle unit. For each basic particle unit, feature quantization is performed, including: calculating the root mean square value of the signal energy within the frame as the particle energy value; dividing the energy distribution of the critical frequency band and solving the weighted average frequency of the energy in each frequency band as the centroid of the particle spectrum; and calculating the inverse of the dispersion of the zero-crossing interval of the waveform within the frame as the particle temporal density. A unique particle identifier is generated for each basic particle unit, and its energy value, spectral centroid coordinates, and time-domain density are bound to it to form an initial particle set. Based on the numerical continuity of particle temporal density, adjacent initial particles with a density difference less than a preset density threshold are aggregated into a particle cluster. Within each particle cluster, the particle with the highest energy value is selected as the dominant particle; Output the set of all particle clusters and their dominant particles as the sound field particle model.
6. The intelligent sound effect optimization method for electronic musical instruments oriented towards multi-source input as described in claim 5, characterized in that, The preset sound field balance rules include: taking the spectral centroid of the dominant particle in the particle cluster as the reference value, the absolute difference between the spectral centroid of other particles in the cluster and the reference value shall not exceed the dynamic radius threshold. The method of correcting sound field particles that deviate from the equilibrium state based on preset sound field balance rules includes: Traverse all particle clusters in the sound field particle model; calculate the absolute difference between the centroid of the spectrum of each particle within the cluster and the reference value; If the absolute difference exceeds the dynamic radius threshold, a compensatory equalization filter parameter is generated according to the offset direction and applied to the original audio frame of the particle; with the dominant particle energy value as the target, dynamic range compression is applied to the original audio frame of the particle. Updating the dynamic radius threshold based on environmental acoustic data includes: if in a strong reflection environment, reducing the dynamic radius threshold to a preset proportion of the original radius threshold; if in a high noise environment, expanding the dynamic radius threshold to a preset proportion of the original radius threshold. Specifically, when the measured ambient reverberation time exceeds the first set threshold, it is determined to be a strong reflection environment; when the background noise sound pressure level exceeds the second set threshold, it is determined to be a high noise environment.
7. The intelligent sound effect optimization method for electronic musical instruments oriented towards multi-source input as described in claim 1, characterized in that, The step of extracting and encoding features from the multi-source input data to generate a multi-dimensional feature vector with temporal correlation includes: Zero-crossing phase-locked analysis is performed on the audio signal to generate a first-dimensional feature characterizing the transient accuracy of the attack signal; The three-dimensional trajectory curvature radius of the performer's limb movement data is calculated to generate a second-dimensional feature representing the complexity of the movement. The environmental acoustic data is subjected to autocorrelation attenuation slope extraction to generate a third-dimensional feature characterizing the spatial reflection intensity.
8. The intelligent sound effect optimization method for electronic musical instruments oriented towards multi-source input as described in claim 2, characterized in that, The method further includes: If the dynamic voiceprint network reconstruction mechanism is triggered more than three times within a preset time period, the node creation, connection relationship update and weight value reset operations will be stopped, and the audio processing engine will be switched to the pre-stored static audio effect parameter set. Cyclic redundancy check codes are embedded when generating the multidimensional feature vector corresponding to each performance event; if the check fails, the original input data of the current performance event is discarded, and the creation of nodes in the dynamic voiceprint network for that performance event is prohibited.
9. The intelligent sound effect optimization method for electronic musical instruments oriented towards multi-source input as described in claim 1, characterized in that, The node distribution density is obtained by calculating the number of valid nodes within a unit time window. The condition for determining a valid node is that the multidimensional feature vector encoding value corresponding to the node exceeds a preset dynamic threshold.
10. An intelligent sound effect optimization system for electronic musical instruments with multi-source input, characterized in that, include: A multi-source data acquisition module is used to acquire multi-source input data of electronic musical instruments in real time. The multi-source input data includes audio signals, performer's limb movement data, and environmental acoustic data. The feature extraction module is used to extract and encode features from the multi-source input data to generate a multi-dimensional feature vector with temporal correlation. The voiceprint network construction module is used to construct a dynamic voiceprint network based on the dynamic correlation between the multidimensional feature vectors. The dynamic voiceprint network consists of nodes and the connection relationships between nodes. Each node corresponds to a performance event, and the node attributes include the multidimensional feature vector encoding value of the corresponding performance event. The connection relationships between nodes are dynamically generated based on the similarity of motion trajectories and audio phase coherence between consecutive performance events, and are assigned connection weight values that represent the strength of the correlation. The matching optimization module is used to match the structural features of the dynamic voiceprint network with a predefined sound effect template library, select a basic sound effect parameter set based on the node distribution density and the connection weight values between nodes, and superimpose parameter correction terms based on the existence state of specific types of connection relationships in the dynamic voiceprint network to generate optimized sound effect parameters. The output module is used to output the processed audio stream based on the optimized sound effect parameters.