Choir balance guidance method and system based on voiceprint recognition
By extracting acoustic energy and proprioceptive voiceprint features using voiceprint recognition technology, and combining cross-correlation ranging and nonlinear dissipation models, the problem of energy assessment distortion caused by physical crosstalk interference on open stages was solved, and the accuracy of reconstructing pure acoustic energy and balancing guidance was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YUANBANG (SHENYANG) ENERGY TECH CO LTD
- Filing Date
- 2026-06-18
- Publication Date
- 2026-07-24
AI Technical Summary
In the sound pickup environment of an open stage, existing technologies cannot effectively identify and eliminate physical crosstalk interference, resulting in energy assessment distortion and balance guidance distortion, which affects the overall listening experience of the band.
Acoustic energy features and ontological voiceprint matching features are extracted using voiceprint recognition technology. The temporal change rate is calculated and orthogonal state deviation is determined. A spectral dissipation filter is constructed using cross-correlation ranging and a nonlinear acoustic dissipation model to generate pure acoustic energy parameters and perform targeted cancellation processing.
It accurately identifies and removes physical crosstalk interference, reconstructs pure acoustic energy, improves system robustness under extreme acoustic aliasing conditions, and ensures the accuracy of balance guidance and the presentation of band timbre characteristics.
Smart Images

Figure CN122454986A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital audio signal processing technology, specifically to a band balance guidance method and system based on voiceprint recognition. Background Technology
[0002] In modern music performances and productions, the mixing balance of the live sound reinforcement (FOH) system directly determines the final listening experience. Traditional band balance instruction methods mainly rely on multi-track microphones to collect audio, use blind source separation algorithms to extract the independent audio streams of each instrument, calculate their short-time energy (RMS), and compare them with a preset energy balance template to output adjustment instructions.
[0003] However, in the actual sound pickup conditions of an open stage, due to the omnidirectional sound wave diffusion of real physical sound sources such as drum kits and speaker cabinets, the microphones of each track will inevitably pick up the sound of adjacent instruments. This unavoidable physical crosstalk will have its frequency domain characteristics distorted after complex spatial attenuation and reflection. Existing linear blind source separation algorithms often cannot effectively remove these distorted crosstalks, and instead will incorrectly attribute them to the current track's intrinsic energy or spatial reverberation. When strong crosstalk interference occurs (such as guitar strumming crossing into the snare drum microphone), this incorrect energy attribution will cause the calculated energy of the target track to spike abnormally. Since the system lacks secondary physical verification of the separation results, it will misjudge this false energy surge caused by crosstalk as the instrument's own volume exceeding the limit, and then continuously output incorrect reverse guidance instructions to reduce the volume.
[0004] Therefore, how to solve the energy assessment distortion and balance guidance distortion caused by the inability to identify and remove physical crosstalk interference in open radio environments has become a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a band balance guidance method and system based on voiceprint recognition, which can identify and eliminate energy assessment distortion and balance guidance distortion caused by physical crosstalk interference.
[0006] To achieve the above objectives, the technical solution of the present invention is as follows:
[0007] In a first aspect, the present invention discloses a band balance guidance method based on voiceprint recognition, comprising the following steps:
[0008] Acquire initial audio data containing multiple pickup channels, and extract the acoustic energy features and ontology voiceprint matching features of the initial audio data within the current processing time window;
[0009] Calculate the first temporal rate of change of acoustic energy features relative to the historical processing time window, and the second temporal rate of change of ontology voiceprint matching features relative to the historical processing time window;
[0010] Based on the first and second time-series change rates, an orthogonal state divergence determination is performed. When the determination result meets the preset triggering conditions, the pickup channel that diverges is identified as the target pickup channel affected by external crosstalk interference.
[0011] Within the preset time delay constraint range, the audio signals of the target pickup channel and the preset candidate pickup channels are extracted and cross-correlation ranging processing is performed to determine the crosstalk emission source channel and the corresponding absolute propagation time delay parameter;
[0012] The spatial physical propagation distance is inverted based on the absolute propagation delay parameter, and a spatial spectrum dissipation filter is constructed based on the spatial physical propagation distance and a preset nonlinear acoustic dissipation model.
[0013] Phase shift processing is performed on the frequency domain signal of the crosstalk source channel using the absolute propagation delay parameter, and frequency response deformation constraint processing is performed on the frequency domain signal of the crosstalk source channel using the spatial spectrum dissipation filter to generate a physical deformation crosstalk prediction signal.
[0014] The frequency domain signal of the target pickup channel is subjected to targeted cancellation processing with the physical deformation crosstalk prediction signal to reconstruct the pure acoustic energy parameters of the target pickup channel, and the target balance guidance command is generated based on the pure acoustic energy parameters.
[0015] Secondly, this invention discloses a band balancing guidance system based on voiceprint recognition, comprising:
[0016] The data acquisition module is used to acquire initial audio data containing multiple pickup channels, and extract the acoustic energy features and body voiceprint matching features of the initial audio data within the current processing time window.
[0017] The data processing module is used to calculate the first temporal rate of change of acoustic energy features relative to the historical processing time window, and the second temporal rate of change of the ontology voiceprint matching features relative to the historical processing time window.
[0018] The target channel determination module is used to perform orthogonal state deviation determination based on the first time change rate and the second time change rate. When the determination result meets the preset triggering condition, the pickup channel that has deviated is determined as the target pickup channel affected by external crosstalk interference.
[0019] The crosstalk source determination module is used to extract the audio signals of the target pickup channel and the preset candidate pickup channels within the preset time delay constraint range, perform cross-correlation ranging processing, and determine the crosstalk emission source channel and the corresponding absolute propagation time delay parameter;
[0020] The filter construction module is used to invert the spatial physical propagation distance based on the absolute propagation time delay parameter, and to construct a spatial spectrum dissipation filter based on the spatial physical propagation distance and a preset nonlinear acoustic dissipation model.
[0021] The crosstalk signal generation module is used to perform phase shift processing on the frequency domain signal of the crosstalk source channel using the absolute propagation delay parameter, and to perform frequency response deformation constraint processing on the frequency domain signal of the crosstalk source channel using the spatial spectrum dissipation filter, thereby generating a physical deformation crosstalk prediction signal.
[0022] The pure parameter reconstruction module is used to perform targeted cancellation processing on the frequency domain signal of the target pickup channel and the physical deformation crosstalk prediction signal to reconstruct the pure acoustic energy parameters of the target pickup channel, and generate target balance guidance instructions based on the pure acoustic energy parameters.
[0023] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0024] 1. By extracting acoustic energy and matching features with the body's voiceprint and calculating its temporal rate of change, orthogonal state deviation judgment is performed to anchor the real physical crosstalk interference. On this basis, the physical spatial distance is inverted through restricted cross-correlation ranging, and a spatial spectrum dissipation filter is constructed in combination with a nonlinear acoustic dissipation model. Crosstalk prediction signals with equivalent real physical deformation are generated in the complex frequency domain for targeted cancellation. This scheme removes heterogeneous spatial crosstalk from the physical propagation mechanism of sound waves, reconstructs pure background energy, and solves the problems of energy assessment distortion and balance guidance distortion caused by crosstalk aliasing in open radio environments.
[0025] 2. A background purity index based on the weighted fusion of signal-to-noise ratio and voiceprint confidence was constructed. The system scientifically quantifies and dynamically sorts the degree of contamination of each channel through the purity index, always locking the highest purity channel as the absolute emission source anchor point, and performing asymmetric cleaning that propagates from high purity to low purity step by step. This effectively blocks the cascading amplification of secondary crosstalk contamination in the multi-channel network and greatly improves the system robustness under extreme acoustic aliasing conditions. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is an overall block diagram of the method in Embodiment 1 of the present invention;
[0028] Figure 2 This is a timing diagram of the overall data flow of the method in Embodiment 1 of the present invention;
[0029] Figure 3 This is an overall block diagram of the system in Embodiment 2 of the present invention. Detailed Implementation
[0030] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0031] In the field of modern live music performance sound reinforcement and multitrack mixing, the precise balance of dynamic energy between the various parts of the band is regarded as a key indicator for measuring the overall auditory atmosphere and mixing quality. The generation of this high-quality balance guide is essentially a process of precise decoupling of multi-source sound field energy at the physical and signal processing level. That is, the independent sound of each instrument is picked up by the microphone array, the target sound track is separated by the audio separation algorithm, and the pure short-time electrical energy is losslessly mapped to the actual acoustic contribution of the instrument in the overall mix, thereby outputting a tuning trajectory that conforms to the dynamic characteristics of a specific musical style.
[0032] However, existing technologies lack a mechanism to verify the causal consistency between the physical space sound wave propagation deformation and the digital audio separation products. This makes it impossible to accurately identify the legitimate dynamic bursts and crosstalk issues commonly found in regular band ensemble performances. Legitimate dynamic bursts are characterized by musicians significantly increasing their playing volume due to heightened emotions, actually providing pure high-energy sound. Crosstalk, on the other hand, is characterized by high-volume sound waves from other instruments propagating across the air medium. Due to nonlinear dissipation and high-frequency absorption, these waves undergo frequency response deformation and then penetrate the target microphone, causing the target channel to exhibit a false surge in energy. Consequently, a strict orthogonal correspondence cannot be established between the channel's electrical energy mutation and the actual physical sound source's output state. This leads to misjudgment or ambiguous feedback of abnormal signals by the system, thus affecting the accurate assessment of the target instrument's acoustic energy parameters and the targeted nature of balance guidance.
[0033] If the above problems are not addressed, the balance guidance system will continue to lose its ability to objectively judge the true sound state. In particular, the failure to identify crosstalk from different spatial sources will cause the system to rely excessively on a single energy evaluation dimension, leading to a deviation of the sound field evaluation path from the principle of pure, ontological guidance, thereby weakening the overall dynamic tension of the band. At the same time, the failure to correct the misjudgment of legitimate dynamic bursts will cause the overtone details of the target instruments to be unnecessarily suppressed and swallowed up, making it impossible for the mixing results to present the true timbre characteristics of the instruments, ultimately causing the live sound field to lose its proper sense of layering and explosiveness. Therefore, the inaccuracy of balance guidance feedback will systematically hinder sound engineers from grasping the core state of energy decoupling in the true sound field, seriously affecting the achievement of high-fidelity live sound reinforcement and high-quality mixing production goals.
[0034] Example 1:
[0035] like Figures 1-2 As shown, the band balance guidance method based on voiceprint recognition includes the following steps:
[0036] To facilitate understanding, we will use a four-channel sound reinforcement scenario of a live band performance as an example. Assume the band's stage pickup network contains four pickup channels: channel 1 for the kick drum, channel 2 for the snare drum, channel 3 for the electric guitar amp, and channel 4 for the lead vocals. Due to physical space constraints, the sound sources of each instrument are relatively close together. The current system's audio sampling rate is set to 48,000 Hz, the processing time window length (i.e., the short-time analysis window length) is configured to 2048 sampling points, and the step offset between adjacent time windows is 1024 sampling points.
[0037] Before implementing the band balance guidance method, the system needs to pre-construct and configure the target voiceprint reference template feature vector offline. In this pre-construction step, the system collects the dry sound signals of each instrument in an ideally pure state (such as in a separate audition) as the data source for the prior knowledge base. For these dry sound signals, the system extracts their Mel-frequency cepstral coefficient features and uses mean aggregation or clustering algorithms to generate centroid vectors in a high-dimensional feature space, which are then stored as the target voiceprint reference template feature vectors corresponding to each pickup channel.
[0038] Specifically, the system segments the collected clean dry sound signal into frames and extracts a large number of MFCC sample points. The K-Means unsupervised clustering algorithm is used to divide the sample points into a single main cluster. After removing outliers (such as performance errors or transient noise) that deviate from the center of the main cluster by a certain Mahalanobis distance, the arithmetic mean position of the retained core sample set is established as the centroid vector of the pickup channel, thereby ensuring that the voiceprint reference template is absolutely pure.
[0039] Specifically, the target voiceprint baseline template feature vector is constructed in the offline stage. The complete steps are as follows: During the independent sound check before the orchestra's performance, the dry sound signals of each instrument are individually acquired under ideal, pure, and crosstalk-free conditions for each pickup channel. The dry sound signals are then framed and a large number of Mel-frequency cepstral coefficient feature vectors are extracted in batches as the original training sample points, forming the training sample set. Where J is the total number of sample points, Represents the L-dimensional feature sample vector extracted at the j-th time. ).
[0040] Run the K-Means unsupervised clustering algorithm and set the cluster number parameter. The initial global centroid vector is obtained by setting it to 1 (i.e., assigning samples to a single main cluster) and iteratively updating it by minimizing the sum of squared Euclidean distances from sample points to cluster centers. To eliminate feature distortion caused by occasional performance errors or transient external noise, the feature sample vector is calculated. to the initial centroid vector Mahalanobis distance :
[0041] ;
[0042] In the formula, Let the Mahalanobis distance be the j-th sample point. Let the current initial centroid vector be... For training sample set The inverse of the covariance matrix.
[0043] Introduce a preset distance threshold Iterate through all sample points and determine if the following conditions are met. If the sample point is identified as an outlier, it will be removed, thus obtaining a pure core sample set free of interference. For the pure core sample set All retained sample vectors are subjected to an arithmetic mean calculation to finally generate the target voiceprint reference template feature vector specific to this pickup channel. : ;
[0044] In the formula, The feature vector of the solidified target voiceprint reference template. This is the core sample set retained after removing outliers. For the sample vectors in the core sample set, This represents the total number of valid samples within the core sample set. This multi-step convergence and cleaning process ensures that the baseline template faithfully reflects the inherent physical acoustic signature of the musical instrument.
[0045] The reason for choosing Mel frequency cepstral coefficients instead of simple spectral amplitude is that these coefficients map the nonlinear perception mechanism of human hearing, can accurately capture the formant envelope and timbre skeleton of the instrument's sound, and have a natural anti-bias capability against absolute volume. This allows the system to have an objective anchor point in subsequent processing that is strongly correlated with the instrument's "physical identity" and decoupled from the "playing intensity".
[0046] Using the example of the band performance mentioned above, during the sound check before the performance, the musicians play the snare drum individually. The system collects this pure snare drum audio, extracts its timbre envelope features, and solidifies it into a "snare drum target soundprint baseline template feature vector" exclusive to Channel 2 in the knowledge base. This vector will serve as a "genetic map" to verify the purity of the sound in Channel 2 during subsequent real-time performances.
[0047] Based on the above pre-constructed foundation, in step S1, initial audio data containing multiple pickup channels is obtained, and the acoustic energy features and body voiceprint matching features of the initial audio data within the current processing time window are extracted respectively.
[0048] In this step, the initial audio data refers to the continuous discrete electrical signal acquired in real time by the stage microphone array and converted from analog to digital. To meet the stringent requirements of backpropagation delay localization and frequency domain phase cancellation that may be triggered in subsequent steps, the system pushes the continuous data stream into a pre-established deep time-domain circular buffer queue in real time after acquiring the initial audio data from each pickup channel. The length of this buffer queue is set to be greater than the total number of sampling points corresponding to the maximum physical delay of the stage, thereby ensuring that the underlying hardware has enough historical data at any time to support the backtracking source on the timeline.
[0049] Specifically, for the extraction of acoustic energy features, the system performs frame segmentation and windowing processing on the initial audio data within the current processing time window, and calculates the effective electrical energy based on the preset short-time energy integration rules, thereby generating the current energy scalar corresponding to the acoustic energy features.
[0050] Because continuous audio signals are non-stationary, processing them directly on the global time axis leads to the loss of local transient features. Therefore, the system segments the continuous signal into short audio frames using a sliding time window. At the frame-segmentation edges, to suppress potential spectral leakage during frequency domain mapping, the system applies a smoothing window function (such as a Hanning or Hamming window) to each frame for initial and final attenuation. After windowing, the system calculates the root mean square energy of the frame signal based on Pasvald's theorem in physical electroacoustics or directly in the time domain based on the short-time energy integration rule. The specific calculation formula is as follows:
[0051] ;
[0052] In the formula, The current energy scalar calculated for the i-th pickup channel within the m-th processing time window; The discrete amplitude values of the initial audio data; The total number of sampling points within the current processing time window; n is the local sampling point index within the window; m is the global incrementing frame index of the time window; H is the step offset between adjacent time windows; denoted by the applied window function. This formula, in a physical sense, characterizes the effective value of the total macroscopic sound field energy captured by the pickup channel during this brief instant.
[0053] Using the previous band performance example, at a specific moment (e.g., frame m), the drummer strikes the snare drum loudly, while the electric guitarist next to him is also playing a loud strumming game. The microphone in channel 2 (snare drum channel) not only picks up the direct sound of the snare drum but also inevitably picks up the high-decibel electric guitar sound. At this moment, the current energy scalar of channel 2, calculated using the aforementioned short-time energy integration rule... This will result in an extremely high value. In a traditional balance guidance system, this high value would directly trigger a false alarm of "snare drum volume is too loud"; however, in this embodiment, this value is only temporarily stored as one of the dimensions.
[0054] To overcome the blind spots of a single energy dimension, the system further extracts ontological voiceprint matching features in parallel. Specifically, the system converts the initial audio data into frequency domain features and extracts the Mel frequency cepstral coefficient feature vector. It then calculates the similarity between this Mel frequency cepstral coefficient feature vector and the preset target voiceprint reference template feature vector, generating the current confidence scalar corresponding to the ontological voiceprint matching features.
[0055] In this processing chain, the system performs a short-time Fourier transform on the windowed time-domain signal, mapping it to the complex frequency domain. Based on this, it performs filtering and discrete cosine transform using a Mel filter bank to extract the Mel frequency cepstral coefficient feature vector reflecting the actual timbre composition within the current time window. The specific processing steps are as follows:
[0056] First, the initial audio data is processed by framing and windowing to obtain short-time frame signals in the time domain. Then, a short-time Fourier transform (STFT) is performed on each frame signal to map it from the time domain to the complex frequency domain, resulting in the corresponding complex spectrum signal. Where b is the discrete frequency band index ( (This refers to the number of points in the short-time Fourier transform, which is set to 2048 in this embodiment). Let b be the actual physical frequency mapped to the b-th discrete frequency band, and its calculation formula is: ,in The system sampling rate scalar (in this embodiment, it is...) ), where m is the globally incrementing frame index of the processing time window.
[0057] Subsequently, the energy spectrum of the complex spectral signal is calculated. The energy spectrum is then input into a preset Mel filter bank for frequency domain energy integration to obtain the short-time energy output of the k-th Mel filter channel. :
[0058] ;
[0059] In the formula, This represents the energy output of the k-th Mel filter channel of the i-th pickup channel in the m-th frame. Let k be the frequency response function of the k-th Mel filter, which has an equal bandwidth triangular distribution on the Mel scale, and k is the channel index of the Mel filter ( K is the total number of channels in the Mel filter, which is set to 26 in this embodiment.
[0060] Furthermore, regarding energy output Taking the logarithm yields the logarithmic Mel energy. The log-Mel energy is subjected to a discrete cosine transform (DCT) to remove correlations between multiple channels, the first L-dimensional low-frequency coefficients are extracted and the 0th-dimensional DC component is excluded, thereby combining to generate the Mel frequency cepstral coefficient eigenvector. :
[0061] ;
[0062] In the formula, This is the feature vector of the Mel frequency cepstral coefficients extracted in real time. This represents the cepstral coefficient of the l-th dimension, where L is the preset cepstral coefficient dimension (set to 13 in this embodiment, i.e., ...). ).
[0063] Subsequently, the system introduces a cosine similarity measurement mechanism to calculate the spatial angle correlation between the real-time extracted feature vector and the pre-fixed benchmark template vector in the knowledge base. The specific calculation formula is as follows:
[0064] In the formula, The current confidence scalar of the i-th pickup channel in the m-th processing time window is normalized to a valid range of 0 to 1; This is the feature vector of the Mel frequency cepstral coefficients extracted in real time; The feature vector of the target voiceprint reference template specific to this channel is invoked; This indicates the norm of the vector. The confidence scalar output by this formula objectively characterizes the purity of the components belonging to its "legal instrument" in the sound captured by the current channel. Let's continue with the previous example of the snare drum and guitar playing simultaneously. In frame m, although the current energy scalar of channel 2 (snare drum channel)... The aliasing of the guitar's intrusion is abnormally high, but due to the significant topological differences between the distorted tone of the electric guitar and the crisp percussion tone of the snare drum in the Mel frequency cepstral feature space, this mixing of heterogeneous signals severely disrupts the original tonal envelope of channel 2. Therefore, the current confidence scalar calculated using the above formula... A significant drop will occur (e.g., from 0.95 in its pure state to 0.42).
[0065] Through the aforementioned data extraction and feature decoupling scheme, the system constructs a time-strictly aligned dual-parameter data stream of "energy-confidence" for each pickup channel in parallel within an extremely short time delay. This underlying design, which abandons the traditional one-way thinking of "emphasizing energy and neglecting timbre" and orthogonally extracts the energy amplitude in the electrical dimension and the timbre purity in the semantic dimension, not only depicts the transient physical slice of the stage sound field, but also lays an indispensable quantitative data foundation for revealing the hidden mechanisms of complex crosstalk interference, implementing state deviation judgment, and fault-tolerant reconstruction in subsequent steps.
[0066] Following the successful construction of the time-aligned "energy-confidence" dual-parameter data stream in step S1, the system needs to perform longitudinal time-axis operations on these transient physical slices to capture the dynamic evolution trend of sound features. In step S2, the first temporal rate of change of the acoustic energy feature relative to the historical processing time window and the second temporal rate of change of the ontological voiceprint matching feature relative to the historical processing time window are calculated.
[0067] In real-world live performance environments, due to the inherent acoustic characteristics of different instruments such as bass drums, snare drums, and guitars, as well as the drastically different gain settings of the preamps on each channel of the mixing console, there is a significant gap in the absolute energy scalar values extracted by each pickup channel. If only static absolute thresholds are used to determine whether the energy exceeds the limit or whether the sound signature is distorted, the system will easily produce serious false alarms when musicians perform normal, dynamic playing (such as a drummer's powerful strikes during an emotional outburst). To eliminate this judgment bias caused by differences in absolute physical magnitudes and hardware gain, this embodiment uses a dynamic gradient in the relative time dimension for measurement, aiming to focus on the relative abrupt changes in sound characteristics within extremely short adjacent time intervals.
[0068] Specifically, to obtain accurate and robust trend data, the system extracts the historical energy scalar and historical confidence scalar corresponding to the historical processing time window adjacent to the current processing time window (e.g., the immediately preceding audio frame) from an internally established temporal buffer queue. It is worth noting that in real digital audio systems, during the background noise phase before a performance or during a rest in an instrumental piece, the short-term energy of the pickup channel may approach absolute zero. If the current energy scalar is directly divided by the historical energy scalar under such extreme silence, it will inevitably trigger a divide-by-zero overflow anomaly at the underlying digital signal processing level, causing the entire audio processing process to crash instantly. To completely solve this engineering dead end, the system forcibly introduces a preset anti-overflow constant factor when performing ratio calculations. Using this constant factor, the system performs ratio calculations on the current energy scalar and the historical energy scalar, and on the current confidence scalar and the historical confidence scalar, respectively, to obtain the corresponding first and second temporal change rates. The specific calculation formulas are as follows:
[0069] ;
[0070] ;
[0071] In the formula, Let i be the first timing change rate of the i-th pickup channel; The second timing change rate of the i-th pickup channel; and These are the current energy scalar and the current confidence scalar, respectively, for the current processing time window; and These are the historical energy scalar and historical confidence scalar for the historical processing time window, respectively. This is the preset anti-collapse overflow constant factor. In actual engineering deployments, this factor is usually set to an extremely small fixed value (e.g., ).
[0072] Continuing with the previous example of a performance where the snare drum and guitar sound simultaneously, let's assume that within the historical processing time window (i.e., the...) At frame ( ), when the snare drum is in a state of weak drumhead reverberation or normal ambient noise, the system extracts the historical energy scalar of channel 2 (snare drum channel). Extremely low, only 0.01; however, due to the relatively pure recording environment, the extracted historical confidence scalar representing the purity of the snare drum's timbre is quite high. Up to 0.95. When time instantly advances to the current processing time window (i.e., frame m), the electric guitarist next to him suddenly bursts into a loud strumming performance. The powerful guitar sound waves instantly overcome the spatial barrier and enter the snare drum microphone, causing the current energy scalar of channel 2 to... It surged to 0.80. However, due to the intrusive guitar distortion tone deviating drastically from the snare drum's acoustic signature envelope, the current confidence scalar for this frame decreased. It dropped sharply to 0.40.
[0073] At this point, the system calls the above-mentioned time series rate of change formula (setting a constant factor). The calculation is performed. For the energy dimension, the first time-series rate of change is calculated. This result characterizes a dramatic 80-fold increase in channel energy between two frames; simultaneously, for the timbre dimension, the second temporal rate of change was calculated. This indicates that the timbre purity of the currently acquired signal has dropped to less than half of its original value.
[0074] Based on the aforementioned dynamic gradient calculation scheme, the system successfully transforms physical parameters, which are constrained by the hardware gain scale of each channel and the absolute acoustic characteristics of the instrument, into a unified, dimensionless relative mutation ratio. This enables the system to keenly perceive the anomalies in the multidimensional features of all pickup channels with a unified scale, thus providing a robust underlying logical criterion for feature cross-validation and multi-channel orthogonal state deviation determination in subsequent steps.
[0075] After calculating the first and second timing change rates in step S2, the system has grasped the transient change characteristics of each pickup channel in terms of energy and timbre. Next, the system needs to make a final judgment on whether the channel is affected by external crosstalk based on these quantified characteristics. In step S3, an orthogonal state deviation judgment is performed based on the first and second timing change rates. When the judgment result meets the preset triggering conditions, the pickup channel that deviates is identified as the target pickup channel affected by external crosstalk interference.
[0076] In traditional single-dimensional detection algorithms, relying solely on abnormal energy spikes can easily misjudge a musician's energetic playing (a legitimate surge in their own energy) as crosstalk interference. To prevent this, this embodiment introduces the logic of "orthogonal state deviation judgment." "Orthogonal state deviation" means that in a real acoustic physical scenario, when a microphone picks up the sound of its own instrument, regardless of the force with which the musician strikes / plays, causing a surge in energy, the extracted voiceprint features should always maintain a high degree of similarity to its target voiceprint reference template (i.e., the confidence scalar remains stable at a high level, or even slightly increases due to the improved signal-to-noise ratio). In other words, under legitimate vocalization conditions, the rate of energy change and the rate of confidence change should either be "in the same direction" or "decoupled" in their evolutionary trends. Conversely, only when an external, heterogeneous sound (crosstalk) invades with overwhelming energy will the total energy surge, and the superposition of heterogeneous timbre components completely destroy the original sound signature structure, causing a precipitous drop in the confidence scalar. This phenomenon of simultaneous "energy surge" and "timbre collapse" with completely opposite trends is called "orthogonal state divergence".
[0077] Specifically, to perform this determination, the system pre-sets two key boundary parameters in the determination logic gating: a first variation threshold (used to define the baseline for abnormal energy surges) and a second variation threshold (used to define the red line for severe drops in timbre purity). The system will use the first temporal change rate output in step S2 ( ) and the second time series rate of change ( The input is fed into the logic gate function, where a logical AND operation is performed. The specific formula for this operation is as follows:
[0078] ;
[0079] In the formula, This is the determination trigger result variable for the i-th pickup channel. When the value is 1, it means that the trigger condition is met and the system determines that serious crosstalk pollution has occurred; when the value is 0, it means that the signal is within the normal performance dynamics or safe noise range. The first time series rate of change; The preset first change threshold (in actual engineering configurations, it is usually set to a value greater than 1, such as 1.5 or 2.0, which means that the energy has increased by at least 50% to 100%). This is the second time-series rate of change; The second threshold for change is preset (usually set to a value less than 1, such as 0.7 or 0.5, representing a drop in confidence level to below 70% or 50% of the original level). This represents logic and operation. The system will officially identify the locked i-th pickup channel as the "target pickup channel affected by external crosstalk interference" that urgently needs to be cleaned if and only if the output of this logic gate is 1, and will send an error correction and reconstruction request instruction to subsequent steps.
[0080] Returning to the previous example of a performance where the snare drum and guitar sound simultaneously, suppose the system engineer sets the first variation threshold based on the band's musical style and dynamic range. Set to 2.0 (meaning a surge is only considered a doubling of energy), and adjust the second threshold for variation. The value is set to 0.6 (meaning that a 40% drop in tone purity is considered distortion).
[0081] In the first scenario: the drummer suddenly performs an extremely forceful "rimshot." The energy scalar of channel 2 (snare drum channel) may surge from the usual 0.1 to 0.5 instantaneously, and the system calculates the first temporal rate of change. Clearly, 5.0 > 2.0 (satisfying this condition). (Conditions). However, since this is still the sound produced by the snare drum itself, its Mel frequency cepstral envelope has not fundamentally changed, and the calculated second time-series rate of change... It might be 0.98 (remaining extremely stable). Since 0.98 is not less than 0.6 (not satisfying the condition). (conditions), logic and gating decision conditions breaking. The output is 0. The system correctly recognizes this as a legitimate burst of dynamics in the performance and will not mark it as crosstalk. The "fader down (volume reduction)" warning command that was originally intended to be output to the mixing console will not be incorrectly blocked.
[0082] In the second scenario: a guitarist's loud strumming is interrupted by a snare drum microphone. Based on the calculation results of step S2, the first timing variation rate of channel 2... Second time series rate of change At this point, the system performs a joint analysis: the first time-series change rate (80) is significantly greater than the first change threshold (2.0), and the second time-series change rate (0.42) is significantly less than the second change threshold (0.6). Both conditions are met simultaneously, and the logic gating is breached. The output is 1. The system immediately becomes alert and determines that there is serious foreign source contamination inside channel 2 (snare drum channel), officially locking it as the "target pickup channel".
[0083] Through this orthogonal state deviation judgment mechanism, the system accurately separates the fuzzy zone between "legitimate performance dynamics" and "illegal crosstalk pollution" in an extremely complex acoustic aliasing environment, successfully avoiding the false alarms that frequently occur in traditional single-variable threshold systems, and laying the foundation for subsequent accurate location of the true culprit of crosstalk and implementation of physical-level deformation cancellation.
[0084] In step S3, the system successfully identified the target pickup channel affected by crosstalk interference from the complex stage sound field through orthogonal state deviation determination. However, to achieve subsequent physical-level cancellation, the system must accurately know "who contaminated it" and "how far away the contamination source is." In step S4, within a preset time delay constraint range, the audio signals of the target pickup channel and preset candidate pickup channels are extracted and cross-correlation ranging processing is performed to determine the crosstalk emission source channel and its corresponding absolute propagation time delay parameter.
[0085] Traditional blind source separation algorithms often treat crosstalk as an abstract statistical noise, attempting to blindly search for a separation matrix in the entire mathematical space. This approach is prone to getting trapped in local optima when dealing with complex music signals and involves enormous computational costs, making it unsuitable for the low-latency requirements of live performances. This embodiment introduces rigid constraints in three-dimensional physical space. Crosstalk on stage is essentially the physical process of sound waves propagating from one instrument (the source of contamination) across the air medium to another instrument's microphone (the contaminated party). As long as physical distance exists, there will inevitably be a time delay in sound wave propagation.
[0086] Specifically, to avoid meaningless roaming calculations across the entire mathematical domain, the system performs a dimensionality reduction constraint operation before executing cross-correlation ranging. The system obtains a preset maximum physical pickup distance scalar (e.g., set to 10 meters based on the actual size of the stage), the current ambient sound velocity scalar (typically around 343 m / s), and the system's sampling rate scalar (e.g., 48000 Hz). The system calculates the maximum number of discrete sampling points required for the sound to travel across the entire stage by dividing the maximum physical pickup distance by the ambient sound velocity and multiplying by the system sampling rate. This number is then rounded up to generate the theoretical maximum sampling point delay truncation boundary. The corresponding truncation boundary generation formula is as follows:
[0087] ;
[0088] In the formula, The theoretical maximum sampling point delay truncation boundary; This is the scalar value for the maximum physical pickup distance; For the environmental sound speed scalar; The system sampling rate scalar; This represents the floor function. The system rounds up from zero to the nearest integer. The interval is strictly limited to a preset latency constraint range. This operation defines a very small physical isolation zone, and any latency results exceeding this range will be discarded.
[0089] After establishing the time delay constraint range, the system officially enters the cross-correlation source-finding phase. The system synchronously extracts the audio signals from the target pickup channel (e.g., channel 2 snare drum in the previous example) and all candidate pickup channels (e.g., channel 1 bass drum, channel 3 guitar, channel 4 vocals) from the circular buffer queue. Within the preset time delay constraint range, the system calculates the discrete cross-correlation function between the target pickup channel signal and each candidate pickup channel signal. In physical acoustics, the peak value of the cross-correlation function precisely indicates the point of highest similarity between two signal waveforms, and the offset of this peak value on the time axis represents the absolute time difference between the arrival of the same sound source at the two microphones. The system finds the global maximum peak value among all cross-correlation calculation results by comparison, and directly identifies the candidate channel that generates this peak value as the true crosstalk emission source channel. Simultaneously, the sampling point offset corresponding to this maximum peak value is locked as the absolute propagation time delay parameter.
[0090] Continuing with the previous snare drum and guitar example, in step S3, channel 2 (snare drum) triggers an alarm. At this point, the system quickly calculates the maximum stage size (e.g., 10 meters). The system extracts the contaminated audio segment from channel 2 and compares it with the signals from channels 1, 3, and 4 respectively. (Assuming a sampling rate of 48000Hz, the system uses 1 sampling point.) Sliding cross-correlation calculations are performed within a window of sampling points. Since the kick drum and vocals are completely uncorrelated with the currently intrusive guitar waveform, their cross-correlation with channel 2 will result in a flat noise floor curve. However, when the system calculates the cross-correlation between channel 2 and channel 3 (electric guitar channel), the offset... At each sampling point, the cross-correlation function suddenly spikes with an extremely high peak. This peak indicates that the anomalous energy surge in channel 2 originates from the electric guitar signal in channel 3. Simultaneously, the system accurately extracts... This is the absolute propagation delay parameter.
[0091] By introducing macroscopic physical boundary constraints and microscopic digital computation, step S4 abandons the time-consuming global blind search and completes the accurate location and time delay calibration of the pollution source in the complex stage network, thus determining the key physical parameters for the subsequent construction of the deformation cancellation model.
[0092] In step S4, the system locks the crosstalk source channel within a preset time delay constraint range by introducing rigid constraints in three-dimensional physical space and extracts its absolute propagation time delay parameter. However, in order to achieve true physical-level cancellation, in step S5, the spatial physical propagation distance is inverted based on the absolute propagation time delay parameter, and a spatial spectral dissipation filter is constructed based on the spatial physical propagation distance and a preset nonlinear acoustic dissipation model.
[0093] Traditional crosstalk cancellation methods often employ linear scalar attenuation models, assuming that crosstalk is simply a scaled-down version of the original sound (e.g., simply multiplying the crosstalk signal by a fixed attenuation factor). However, according to the principles of aerodynamics and acoustic propagation (such as the ISO 9613-1 standard), air is not only a propagation medium but also a nonlinear low-pass filter. When sound waves propagate through air, their energy dissipation is proportional to the square of the frequency; in other words, high-frequency components attenuate much faster than low-frequency components. Using a simple linear scalar model for cancellation inevitably leads to the possibility of low frequencies being canceled out while high frequencies are over-cancelled, thus introducing harsh metallic artifacts into the target pickup channel.
[0094] This embodiment uses "distance" as the core dimension and introduces a nonlinear acoustic dissipation model. First, the system needs to convert the microscopic digital time delay obtained in step S4 into a macroscopic absolute physical propagation distance. Based on the ambient sound speed scalar and the system sampling rate scalar, the system performs the following inversion calculation:
[0095] ;
[0096] In the formula, d is the spatial physical propagation distance calculated by inversion (in meters). This is the absolute propagation delay parameter (i.e., the number of sampling points corresponding to the cross-correlation peak value). The system sampling rate scalar; The ambient sound speed is a scalar. This inversion process successfully maps the time delay at the digital algorithm level back precisely to the physical space of the real three-dimensional stage.
[0097] After obtaining the precise spatial physical propagation distance d, the system begins to construct the core "space spectrum dissipation filter". The construction of this filter consists of two parts: geometric diffusion attenuation and nonlinear air absorption attenuation.
[0098] The first part addresses the linear attenuation caused by the spherical diffusion of sound waves in three-dimensional space, but special attention needs to be paid to the near-field singularity problem. If a model inversely proportional to distance (1 / d) is simply used, the attenuation factor will tend towards infinity when the distance d approaches 0 (e.g., the microphone is extremely close to the sound source), causing the system gain to run rampant. To curb this anomalous near-field gain, the system introduces a sum-of-squares model incorporating a reference distance parameter, generating a "soft-knee" regularization factor. Its corresponding construction formula is as follows:
[0099] ;
[0100] In the formula, is the soft knee regularization factor; d is the spatial physical propagation distance; This is a preset reference distance parameter (usually set to the microphone's near-field reference distance, such as 1 meter). This formula ensures that at extremely close distances, the gain is safely clamped to 1, while at long distances, it smoothly transitions to an attenuation curve that conforms to the inverse square law.
[0101] The second part addresses the nonlinear absorption of sound waves of different frequencies by the air medium. The system extracts the corresponding standard air absorption coefficients for each discrete frequency band contained in the frequency domain signal. (The unit is usually dB / m, which can be obtained by looking up a table or calculating based on ambient temperature and humidity). Since this coefficient is expressed in logarithmic units (decibels), the system must strictly map it to a linear amplitude multiplier and multiply it by the aforementioned soft-knee regularization factor. The corresponding filter construction formula is as follows:
[0102] ;
[0103] In the formula, This is the spatial spectrum dissipation filter corresponding to the b-th discrete frequency band; This represents the actual physical frequency mapped to the b-th discrete frequency band. This is the standard air absorption coefficient corresponding to this physical frequency. The formula outputs a dimensionless gain coefficient that varies non-linearly with frequency, characterizing the "spectral tilt" phenomenon produced when sound waves pass through the air.
[0104] In the specific implementation of constructing this filter, the discrete frequency band index b and the corresponding physical frequency are involved. air absorption coefficient With reference distance parameters The specific ways to obtain the value are as follows:
[0105] Regarding discrete frequency band index b and physical frequency The mapping, the discrete granularity of the filter in the frequency domain and the short-time Fourier transform used in step S1. The point frequency domain mapping is strictly aligned to correspond to the physical frequency. Determined by the following formula: ;
[0106] In the formula, Let b be the physical frequency (in Hz) corresponding to the b-th discrete frequency band, where b is the discrete frequency band index. The sampling rate scalar of the aforementioned system. The number of points in the short-time Fourier transform (in this embodiment) =2048, },thus It is uniformly distributed in steps of approximately 23.4 Hz within the range of 0 Hz to 24000 Hz.
[0107] Regarding air absorption coefficient The data is obtained based on the airborne sound attenuation calculation model given in ISO 9613-1 standard, combined with the ambient temperature obtained in advance or configured on-site. (Unit: degrees Celsius), relative humidity (RH) (unit: percentage), and atmospheric pressure (Unit: kPa), according to the standard calculation procedure for the pure tone absorption coefficient, for each discrete frequency band... Calculate the corresponding absorption coefficients respectively. (Unit: dB / m). To reduce the real-time calculation load on site, the system initializes under typical operating conditions (in this embodiment, we take...) , The absorption coefficients for all discrete frequency bands are calculated at once and stored as a lookup table. During runtime, the absorption coefficients can be retrieved directly from the table using the frequency band index b, without repeated calculations.
[0108] Regarding reference distance parameters The principle of value selection, The physical meaning of this value lies in defining the distance boundary of the microphone entering the near-field zone. The system sets this value to the near-field reference distance of a typical dynamic or condenser microphone; in this embodiment... Take 1 meter as an example, and set the soft knee regularization factor accordingly. When d is much less than 1 meter, the gain is clamped within a safe range close to 1. When d is much greater than 1 meter, it smoothly transitions to a curve that conforms to the inverse square attenuation law of spherical waves. This parameter allows for a one-time configuration based on the actual stage microphone model and placement habits during the system deployment phase, and remains fixed during operation.
[0109] Continuing with the example of simultaneous snare drum and guitar sounding, in step S4, the system identifies the guitar (channel 3) as the crosstalk emission source and extracts the absolute propagation delay. There are 10000 sampling points. Assuming a sampling rate of 48000Hz and a sound speed of 343m / s, the system first determines the physical distance from the guitar amplifier to the snare drum microphone. Meters. Subsequently, the system begins constructing a spatial spectrum dissipation filter. For low-frequency bands (e.g., 100Hz), the air absorption coefficient... Minimal, filter The attenuation is mainly due to geometric diffusion. The air absorption coefficient is determined; however, for high-frequency bands (such as 10kHz), it is determined. Significantly increased, leading to filter At high frequencies, it exhibits an exponential and rapid decay. By substituting the physical distance d=3 meters into the above formula, the system obtained a frequency domain dissipation curve that closely matches the actual physical deformation.
[0110] By using this mapping design of "using macroscopic physical space parameters to forcibly constrain the shape of microscopic frequency domain filters", step S5 abandons the blindness of pure mathematical fitting and lays the foundation for subsequent physical deformation cancellation.
[0111] In step S5, the system constructs a spatial spectrum dissipation filter by introducing a nonlinear acoustic dissipation model. This is equivalent to restoring the environmental resistance encountered by sound waves when crossing physical space. However, to achieve final physical-level cancellation, it is also necessary to align the time axis and convert the original crosstalk signal into a predicted signal after physical deformation. In step S6, phase shift processing is performed on the frequency domain signal of the crosstalk emission source channel using the absolute propagation delay parameter, and frequency response deformation constraint processing is performed on the frequency domain signal of the crosstalk emission source channel using the spatial spectrum dissipation filter to generate a physically deformed crosstalk prediction signal.
[0112] In traditional time-domain cancellation schemes, the system often directly subtracts the delayed source signal from the contaminated signal in the time domain. However, due to the limitations of the sampling rate of digital systems, the absolute propagation delay parameter measured in step S4 often contains a fractional part (i.e., residual sub-sampling point delay). Forcibly performing integer sampling point offset alignment in the time domain will result in microsecond-level phase misalignment, thus introducing severe phase jumps (click artifacts) and comb filtering effects into the final cancellation result.
[0113] To completely solve this problem, this embodiment adopts an architecture of "integer / fractional delay decoupling" and "time-frequency hybrid alignment". First, the system decouples the absolute propagation delay parameter in floating-point form obtained in step S4 into two parts: the first part is the integer sampling point delay (used for coarse alignment in the time domain); the second part is the residual fractional delay (used for fine-tuning phase shift in the frequency domain).
[0114] Specifically, the system uses integer sampling point delays to perform index backoff offsets on the historical time-domain buffer queue of the crosstalk emission source channel, extracting the coarsely aligned time-domain signal. Subsequently, the system performs a short-time Fourier transform on this coarsely aligned signal, converting it into a frequency-domain signal matrix. Next, the system converts the residual fractional delay into an absolute time deflection scalar and performs complex exponential phase shift compensation on the frequency-domain signal matrix in the complex frequency domain. Based on the time-shifting characteristics of the Fourier transform, the complex exponential phase shift in the frequency domain is equivalent to an infinitely precise continuous-time shift in the time domain. This operation compensates for the minute errors caused by the integer sampling point offset, achieving seamless signal time alignment.
[0115] After aligning the time axis, the system must reshape the crosstalk signal through "physical deformation." In the complex frequency domain, the system performs a synchronously coupled complex multiplication operation on the aligned frequency domain signal matrix, the complex exponential phase shift factor corresponding to the complex exponential phase shift compensation processing, and the spatial spectrum dissipation filter constructed in step S5, thereby generating a physically deformed crosstalk prediction signal. The specific processing formula is as follows:
[0116] ;
[0117] In the formula, The generated physical deformation crosstalk prediction signal corresponds to the m-th frame; This is the frequency domain signal matrix of the crosstalk emission source channel after coarse time-domain alignment and frequency-domain transformation; It is the absolute time scalar obtained by converting the residual decimal time delay; That is, the complex exponential phase shift factor for performing complex exponential phase shift compensation processing; This refers to the spatial spectrum dissipation filter constructed in step S5; is the actual physical frequency mapped to the b-th discrete frequency band.
[0118] Continuing with the four-channel band performance example, in step S4, the system identifies channel 3 (guitar) as the crosstalk emission source and extracts the absolute propagation delay. The system first decouples this time delay: the integer part consists of 420 sample points, and the residual fractional part consists of 0.3 sample points. The system then directs the pointer to move backward 420 sample points in the historical time-domain buffer queue of channel 3, extracting this guitar audio segment, and performing a short-time Fourier transform to obtain the frequency domain signal matrix. This completes a rough alignment in the time domain.
[0119] The system then converts the remaining 0.3 sampling points into absolute time (assuming a sampling rate of 48000Hz). (seconds), and use this to calculate the complex exponential phase shift factor for each discrete frequency band. .
[0120] Finally, the system performs a synchronous complex multiplication operation on the coarsely aligned guitar spectrum, the complex exponential phase shift factor, and the spatial spectrum dissipation filter constructed in step S5, which includes distance dissipation and high-frequency absorption characteristics of air.
[0121] After this series of calculations, the originally pure, full-range guitar signal was forcibly delayed by 0.3 sampling points at the microsecond level, while the high-frequency components were attenuated, ultimately transforming into a guitar sound that was slightly muffled and phase-shifted. Physically, this sound is equivalent to the actual crosstalk that traversed the 3-meter stage space and finally entered the snare drum microphone in Channel 2. The system successfully predicted and reproduced this physical deformation process, preparing for the final cancellation step.
[0122] After completing the physical deformation prediction of the crosstalk signal in step S6, in step S7, the frequency domain signal of the target pickup channel and the physical deformation crosstalk prediction signal are subjected to targeted cancellation processing to reconstruct the pure acoustic energy parameters of the target pickup channel, and the target balance guidance command is generated based on the pure acoustic energy parameters.
[0123] In traditional audio noise reduction or crosstalk processing, simple noise gates or dynamic EQs are often used for compression. This not only irreversibly swallows the original overtone details of the target instrument but also results in a dry and unnatural sound in the mix. This embodiment uses complex frequency domain subtraction based on the principle of wave field interference. Specifically, since the predicted signal generated in step S6 is perfectly equivalent to the real physical crosstalk in terms of time phase and frequency response envelope, the system only needs to perform point-to-point complex subtraction between the contaminated frequency domain signal of the target pickup channel and the predicted signal of the physical deformation crosstalk in the complex frequency domain. This operation is equivalent to creating an "antimatter wave" that is completely out of phase with the interfering sound wave, achieving targeted cancellation. The corresponding calculation formula is as follows:
[0124] ;
[0125] In the formula, This corresponds to the clean frequency domain signal after cancellation in the m-th frame; The target pickup channel's frequency domain signal is contaminated before processing; This is the physical deformation crosstalk prediction signal generated in step S6. The subtraction here is performed in the complex domain (including the real and imaginary parts), which not only eliminates the amplitude energy of the crosstalk, but also more accurately removes the complex phase interference caused by the crosstalk in the target microphone.
[0126] After obtaining a completely clean frequency domain signal matrix, the system needs to convert it back into macroscopic energy parameters that a mixing engineer can intuitively understand. Based on Parseval's theorem in physical electroacoustics, the total energy of a time-domain signal is equal to the sum of the energies of its frequency components in the frequency domain. Therefore, the system does not need to perform an inverse Fourier transform to return the signal to the time domain. Instead, it directly performs a discrete-band summation operation on the squared amplitudes of the clean frequency domain signal after cancellation of all discrete frequency bands, thereby reconstructing the most realistic clean acoustic energy parameters. The corresponding reconstruction formula is as follows:
[0127] ;
[0128] In the formula, The pure acoustic energy parameters extracted for the final reconstruction of the system; This represents the total number of sampling points within the current processing time window. This is the square of the amplitude of the clean frequency domain signal after crosstalk removal in the b-th discrete frequency band. This formula, when directly integrated in the frequency domain, restores the true physical sound pressure level after crosstalk removal. Finally, the system uses this clean acoustic energy parameter... The balance guidance engine injected into the master control unit compares the health status with the preset music style energy template, and then generates accurate target balance guidance instructions (such as outputting a green normal indicator on the mixing console UI, or giving precise decibel push and pull suggestions).
[0129] Specifically, in the actual implementation of targeted cancellation processing, frame alignment, complex subtraction, and inter-frame stitching are completed sequentially in the following steps:
[0130] First, using the globally incrementing frame index m of the current processing time window of the target audio pickup channel as a reference, the starting position of the sample that the crosstalk source channel should back-fetch in the time-domain buffer queue is determined according to the previously obtained integer sampling point delay. This allows the frequency domain signal matrix obtained after the crosstalk source channel back-fetches and performs a short-time Fourier transform to be... ), and the frequency domain signal of the target pickup channel in the m-th frame. There is a strict one-to-one correspondence at frame index m.
[0131] Then, in the same discrete frequency band corresponding to the m-th frame The system will output the frequency domain signal of the target pickup channel. Crosstalk prediction signal with physical deformation Decompose the complex number into real and imaginary components, and perform complex number subtraction using the following formula:
[0132] ;
[0133] ;
[0134] In the formula, and These represent the operators for extracting the real and imaginary parts of the complex number within the parentheses, respectively. This corresponds to the clean frequency domain signal after cancellation in the m-th frame; The target pickup channel's frequency domain signal is contaminated before processing; This is the physical deformation crosstalk prediction signal generated in step S6. The system sequentially performs this complex subtraction operation on all discrete frequency bands b (ranging from 0 to N / 2) to obtain the clean frequency domain signal matrix after full-band cancellation for the current frame.
[0135] Then directly use the aforementioned The calculation formula sums the squares of the pure frequency domain signal amplitudes of all discrete frequency bands to obtain the pure acoustic energy parameter corresponding to the m-th frame. This parameter is continuously output as the frame index m advances, without the need for inverse short-time Fourier transform to return to the time domain, thus significantly reducing the processing delay of the system in on-site sound reinforcement scenarios.
[0136] Finally, considering the overlapping region between adjacent processing time windows determined by the aforementioned step offset H (in the example above, the processing time window length) (Using 2048, step offset H is 1024, and adjacent frame overlap rate is 50%), the system outputs the pure acoustic energy parameters of two adjacent frames. and Perform smooth transition processing: When the difference in energy parameters between two adjacent frames exceeds the preset energy jump threshold, the output value of the current frame and the output value of the previous frame are weighted and fused according to the preset smoothing factor to avoid step jumps in energy parameters between adjacent frames caused by transient analysis window boundary effects; the smoothed pure acoustic energy parameters are then injected into the balance guidance engine of the master control end, and the health is compared with the preset music style energy template to output the target balance guidance command.
[0137] In the processing flow of steps S1 to S7 above, the system has solved the problem of precise location and physical cancellation of single crosstalk interference. However, in the complex stage environment of reality, it is often not just a single channel that is interfered with. When there are large dynamic bursts such as a full ensemble of instruments (Tutti) on stage, it is very easy for multiple pickup channels to simultaneously meet the aforementioned orthogonal state deviation triggering conditions. If only a unidirectional parallel cancellation logic is adopted, the system may fall into a topological deadlock due to secondary pollution. For example, if the sound of channel B is crosstalked into channel A, and channel B itself is severely polluted by channel C, if channel B containing impurities is used directly to clean channel A, the characteristics of channel C will be incorrectly imprinted into channel A, causing a cascading amplification of errors. In order to completely solve this multi-source aliasing problem, after reconstructing the pure acoustic energy parameters, the system further introduces a multi-dimensional purity scoring and a multi-channel iterative topological cascading cancellation architecture.
[0138] Specifically, when the system detects that multiple pickup channels simultaneously meet preset trigger conditions, the primary task is to accurately assess which of these alarm channels is the least contaminated "relatively pure source." To this end, the system controls a pointer to backtrack, extracting the baseline background noise scalar recorded by these multiple pickup channels within a preset historical silent time window (e.g., the standby period before a band performance or during a musical rest). Subsequently, the system combines the acoustic energy characteristics within the current processing time window with this baseline background noise scalar to calculate the current instantaneous signal-to-noise ratio (SNR) characteristics of each pickup channel. To comprehensively and holistically characterize the true uncontaminated level of the data, the system performs a deep weighted fusion calculation of this electrical-level instantaneous SNR characteristic and the semantic-level ontological voiceprint matching characteristic to generate a baseline purity index. The specific calculation formula is as follows:
[0139] ;
[0140] ;
[0141] In the formula, To calculate the instantaneous signal-to-noise ratio characteristic of the i-th pickup channel, it is based on the current energy. Compared with the reference background noise scalar The ratio is generated by taking the logarithm; is the background purity index of the i-th pickup channel; The current confidence scalar corresponding to the ontology voiceprint matching features extracted in the preceding steps; and These are the first and second preset weights, which are pre-trained through machine learning or calibrated empirically. The physical meaning of this index is that it comprehensively considers whether a channel's sound is loud enough (interference resistance) and whether its timbre is pure enough (degree of interference), providing a quantitative basis for subsequent ranking.
[0142] To eliminate the weight imbalance caused by the absolute dimensional difference between cosine similarity and decibel signal-to-noise ratio (which has a very large range of values), the system uses the Sigmoid function to normalize and map the instantaneous signal-to-noise ratio features before performing weighted fusion: In the formula, The normalized signal-to-noise ratio feature (values strictly falling within) ), This is the sensitivity coefficient. This is the preset nominal signal-to-noise ratio threshold. Subsequently, the system performs fusion calculations under the same dimensions: This ensures the accuracy of multidimensional feature fusion.
[0143] For the first preset weight With the second preset weight In the machine learning training process, the system pre-constructs a multi-track audio verification dataset. The input feature samples of this dataset are two-dimensional feature sets corresponding to a large number of known clean or contaminated audio slices collected from historical performances. The supervision labels employ binary one-hot codes (1 representing extremely clean data, 0 representing severe crosstalk) verified manually or through multi-track direct recording. The system uses a Logistic Regression model or Support Vector Machine (SVM) to construct a linear classification hyperplane, using binary cross-entropy as the loss function. Gradient descent optimization is performed on the training set to allow the model to automatically find the optimal classification hyperplane in the feature space; the normal vector components of this hyperplane are the final determined classification hyperplane. and Through this supervised learning mechanism, the system quantifies and transforms the expert's auditory experience into a generalizable combination of high-order weights.
[0144] After successfully generating the background purity index for all alarm channels, the system first constructs a cleaning queue for multiple pickup channels that meet the trigger conditions. The channels in the queue are then sorted in real time, and the channel with the highest background purity index value is precisely selected and locked as the highest purity anchor channel. Since this anchor channel has the highest score, the system logically assumes that its current level of external contamination is extremely low, and its frequency domain signal is sufficient to perfectly represent the true characteristics of the instrument. Furthermore, the system configures this highest purity anchor channel as the absolute crosstalk emission source channel, and performs the aforementioned cross-correlation ranging processing and targeted cancellation processing based on spatial spectrum dissipation filters on all other target pickup channels in the cleaning queue, completing the first round of global cleaning.
[0145] It is worth noting that after the first round of cancellation cleaning, the most significant interference source has been removed from the previously heavily contaminated channels, inevitably altering their sound field purity. Based on this, the system then re-samples the target pickup channels that have undergone targeted cancellation processing and recalculates their updated background purity index. Subsequently, the system selects the channel with the highest updated value from the remaining channels, locks it as the secondary reference anchor channel, and configures it as a new crosstalk emission source channel for the second round of cancellation processing. The system iteratively executes this topology cancellation operation, progressively advancing from high purity to low purity, until all target pickup channels in the cleaning queue have been completely cancelled and reconstructed, ultimately generating target balance guidance instructions.
[0146] To ensure the system's robustness in real-time control under extremely complex aliasing environments and avoid iteration traps, the system incorporates dual truncation protection within the iteration loop, including a maximum number of iterations and a minimum purification gradient threshold. If the number of cleaning iterations exceeds the preset fault tolerance limit (e.g., N times for the total number of channels), or if the improvement in the background purity index resulting from two consecutive cleaning rounds is less than the minimum purification gradient threshold, the system will forcibly exit the iteration loop. At this point, the system triggers a degradation fault tolerance mechanism, forcibly outputting the currently obtained best cancellation result as the final state, thereby ensuring that low-latency sound reinforcement guidance for live performances is not interrupted.
[0147] To make this cascading logic more intuitive, let's upgrade the aforementioned four-channel example to a scenario of cascading crosstalk. Suppose that during the climax of a piece of music, the electric guitar (channel 3) is frantically strumming and howling right up against the amp, the snare drum (channel 2) is hitting with extremely strong, continuous force, and the vocalist (channel 4) is singing passionately. At this moment, the powerful guitar sound waves not only crosstalk into the snare drum microphone but also bypass the drum kit and crosstalk into the vocal microphone; simultaneously, the loud snare drum sound also crosstalks into the vocal microphone. The system's crosstalk detection logic is instantly broken, and channels 2 (snare drum) and 4 (vocal) simultaneously trigger the highest level of crosstalk alarms. Faced with this chaotic situation, the system doesn't rush to blindly subtract, but immediately calculates the background purity index of the three sound channels. The evaluation results show that the guitar (channel 3), due to its extremely high volume and close proximity to the amp, has a high crosstalk intensity. The index is as high as 95 points; although the snare drum (channel 2) is affected by the guitar, it still has sufficient energy. The score is 60 points; vocals (channel 4) are the weakest, suffering from the double attack of guitar and snare drum. The index has only 30 points left.
[0148] Based on this ranking, in the first iteration, the system locks the guitar (channel 3) as the highest purity anchor channel. The system first uses the guitar signal to clean the snare drum (channel 2) and vocals (channel 4), accurately removing guitar interference noise from the two microphones.
[0149] After the first round of cleaning, the guitar impurities in the snare drum (channel 2) were completely removed, and its tone became pure. The system recalculated and found that its purity index was... It jumped to 85 points, while the voice channel Although there was an improvement, the snare drum noise still remained. Therefore, in the second iteration, the system locked the cleaned snare drum (channel 2) as the secondary reference anchor channel and used it to precisely cancel the snare drum crosstalk remaining in the vocals (channel 4).
[0150] Through this asymmetric topology cascaded network, the system blocks the logical paradox of signals containing contaminants being misused as a cancellation reference, peeling away the complex crosstalk layers stacked in the vocal channel. Ultimately, what the sound engineer sees on the screen is no longer chaotic parameters with frequent alarms, but pure balance guidance data that reflects the instrument's sound production state, cleansed by the algorithm.
[0151] In summary, this embodiment utilizes the dual-dimensional divergence between acoustic energy and the original acoustic signature to accurately anchor contamination events. It extracts the absolute physical distance using a restricted cross-correlation search and time-delay inversion mechanism, and combines a nonlinear spatial spectrum dissipation filter and time-delay decoupling technology to generate a high-precision deformation prediction signal, achieving targeted cancellation in the complex dimensions of the frequency domain. After superimposing a multi-channel purity iterative topology cascade architecture, the system effectively cuts off the secondary contamination propagation chain in complex aliasing scenarios. Therefore, this method overcomes the inherent limitations of traditional purely statistical blind source separation algorithms, which are prone to getting trapped in local optima and causing phase distortion. While preserving the original overtones and dynamic structure of the target instrument, it removes spatial crosstalk interference and reconstructs the true background acoustic energy parameters, thus providing a balanced guideline for high-fidelity live sound reinforcement and automated mixing in complex stage environments.
[0152] Example 2:
[0153] like Figure 3 As shown, the band balance guidance system based on voiceprint recognition includes:
[0154] The data acquisition module is used to acquire initial audio data containing multiple pickup channels, and extract the acoustic energy features and body voiceprint matching features of the initial audio data within the current processing time window.
[0155] The data processing module is used to calculate the first temporal rate of change of acoustic energy features relative to the historical processing time window, and the second temporal rate of change of the ontology voiceprint matching features relative to the historical processing time window.
[0156] The target channel determination module is used to perform orthogonal state deviation determination based on the first time change rate and the second time change rate. When the determination result meets the preset triggering condition, the pickup channel that has deviated is determined as the target pickup channel affected by external crosstalk interference.
[0157] The crosstalk source determination module is used to extract the audio signals of the target pickup channel and the preset candidate pickup channels within the preset time delay constraint range, perform cross-correlation ranging processing, and determine the crosstalk emission source channel and the corresponding absolute propagation time delay parameter;
[0158] The filter construction module is used to invert the spatial physical propagation distance based on the absolute propagation time delay parameter, and to construct a spatial spectrum dissipation filter based on the spatial physical propagation distance and a preset nonlinear acoustic dissipation model.
[0159] The crosstalk signal generation module is used to perform phase shift processing on the frequency domain signal of the crosstalk source channel using the absolute propagation delay parameter, and to perform frequency response deformation constraint processing on the frequency domain signal of the crosstalk source channel using the spatial spectrum dissipation filter, thereby generating a physical deformation crosstalk prediction signal.
[0160] The pure parameter reconstruction module is used to perform targeted cancellation processing on the frequency domain signal of the target pickup channel and the physical deformation crosstalk prediction signal to reconstruct the pure acoustic energy parameters of the target pickup channel, and generate target balance guidance instructions based on the pure acoustic energy parameters.
[0161] The above description is merely an example and illustration of the structure of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the structure of the invention or exceed the scope defined in the claims, all of which should fall within the protection scope of the present invention.
[0162] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0163] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A band balance guidance method based on voiceprint recognition, characterized in that, Includes the following steps: Acquire initial audio data containing multiple pickup channels, and extract the acoustic energy features and body voiceprint matching features of the initial audio data within the current processing time window respectively; Calculate the first temporal rate of change of the acoustic energy feature relative to the historical processing time window, and the second temporal rate of change of the body voiceprint matching feature relative to the historical processing time window; Based on the first time-series change rate and the second time-series change rate, an orthogonal state divergence determination is performed. When the determination result meets the preset triggering condition, the pickup channel that diverges is determined as the target pickup channel affected by external crosstalk interference. Within a preset time delay constraint range, the audio signals of the target pickup channel and the preset candidate pickup channel are extracted and cross-correlation ranging processing is performed to determine the crosstalk emission source channel and the corresponding absolute propagation time delay parameter. The spatial physical propagation distance is inverted based on the absolute propagation delay parameter, and a spatial spectrum dissipation filter is constructed based on the spatial physical propagation distance and a preset nonlinear acoustic dissipation model. The frequency domain signal of the crosstalk source channel is subjected to phase shift processing using the absolute propagation delay parameter, and the frequency response deformation constraint processing of the frequency domain signal of the crosstalk source channel is subjected to the spatial spectrum dissipation filter to generate a physical deformation crosstalk prediction signal. The frequency domain signal of the target pickup channel and the physical deformation crosstalk prediction signal are subjected to targeted cancellation processing to reconstruct the pure acoustic energy parameters of the target pickup channel, and a target balance guidance command is generated based on the pure acoustic energy parameters.
2. The band balance guidance method based on voiceprint recognition according to claim 1, characterized in that: Extracting the acoustic energy features and ontological voiceprint matching features of the initial audio data within the current processing time window includes: The initial audio data is subjected to frame segmentation and windowing processing, and the effective electrical energy is calculated based on the preset short-time energy integration rule to generate the current energy scalar corresponding to the acoustic energy feature; The initial audio data is converted into frequency domain features and the Mel frequency cepstral coefficient feature vector is extracted. The similarity between the Mel frequency cepstral coefficient feature vector and the preset target voiceprint reference template feature vector is calculated, and the current confidence scalar corresponding to the ontological voiceprint matching feature is generated.
3. The band balance guidance method based on voiceprint recognition according to claim 2, characterized in that: Based on the first and second time-series change rates, an orthogonal state divergence determination is performed. When the determination result meets a preset trigger condition, the microphone channel where the divergence occurs is identified as the target microphone channel affected by external crosstalk interference, including: Obtain the historical energy scalar and historical confidence scalar corresponding to the historical processing time window; Using a preset anti-collapse overflow constant factor, the ratio calculation is performed between the current energy scalar and the historical energy scalar, and the ratio calculation is performed between the current confidence scalar and the historical confidence scalar, respectively, to obtain the corresponding first time series change rate and second time series change rate; When the first timing change rate is greater than a preset first change threshold and the second timing change rate is less than a preset second change threshold, the triggering condition is determined to be met, and the pickup channel that deviates is identified as the target pickup channel.
4. The band balance guidance method based on voiceprint recognition according to claim 1, characterized in that: Before performing cross-correlation ranging processing, the method further includes: Obtain the preset maximum physical pickup distance, ambient sound velocity scalar, and system sampling rate scalar; The theoretical maximum sampling point delay cutoff boundary is generated by dividing the maximum physical pickup distance by the ambient sound speed scalar and multiplying the resulting quotient by the system sampling rate scalar. The interval from zero to the theoretical maximum sampling point delay truncation boundary is determined as the preset time delay constraint range.
5. The band balance guidance method based on voiceprint recognition according to claim 1, characterized in that: After determining the crosstalk source channel and the corresponding absolute propagation delay parameter, the method further includes: The absolute propagation delay parameter is decoupled into integer sampling point delay and residual fractional delay; Performing phase shift processing on the frequency domain signal of the crosstalk emission source channel using the absolute propagation delay parameter includes: The audio signal of the crosstalk emission source channel is stored in the time-domain buffer queue, and the time-domain buffer queue is indexed backoff offset using the integer sampling point delay to extract the time-domain coarse alignment signal and convert it into a frequency-domain signal matrix. The residual fractional delay is used to perform complex exponential phase shift compensation processing on the frequency domain signal matrix in the complex frequency domain.
6. The band balance guidance method based on voiceprint recognition according to claim 5, characterized in that: A spatial spectrum dissipation filter is constructed based on the aforementioned spatial physical propagation distance and a preset nonlinear acoustic dissipation model, including: Based on the spatial physical propagation distance and the preset reference distance parameters, a soft knee regularization factor for clamping near-field gain is constructed and generated. For each discrete frequency band contained in the frequency domain signal, the corresponding standard air absorption coefficient is extracted, and a logarithmic-to-linear mapping transformation is performed in combination with the spatial physical propagation distance. This transformation is then multiplied by the soft-knee regularization factor to construct a spatial spectral dissipation filter for the discrete frequency band dimension. The calculation formula is as follows: ; ; in, As the soft knee regularization factor, For spatial physical propagation distance, For reference distance parameters, For the corresponding number Spatial spectrum dissipation filter for discrete frequency bands For the first The physical frequency mapped to a discrete frequency band. This is the standard air absorption coefficient corresponding to the physical frequency.
7. The band balance guidance method based on voiceprint recognition according to claim 6, characterized in that: Generate physical deformation crosstalk prediction signals and perform targeted cancellation processing, including: The frequency domain signal matrix obtained after the coarse time-domain signal conversion is used to generate the physical deformation crosstalk prediction signal by synchronously coupling a complex exponential phase shift factor with the spatial spectrum dissipation filter through complex multiplication; wherein, the complex exponential phase shift factor is determined by the residual fractional time delay conversion. In the complex frequency domain, the physical deformation crosstalk prediction signal is subtracted from the frequency domain signal of the target pickup channel to obtain the canceled pure frequency domain signal; The pure acoustic energy parameters are reconstructed by performing discrete-band summation on the squared amplitudes of the clean frequency domain signals after cancellation across all discrete frequency bands; the formulas for targeted cancellation and reconstruction are as follows: ; ; in, For the corresponding to the first The clean frequency domain signal after frame cancellation. The frequency domain signal of the target pickup channel. For physical deformation crosstalk prediction signals, For pure acoustic energy parameters, This represents the total number of sampling points within the current processing time window.
8. The band balance guidance method based on voiceprint recognition according to claim 1, characterized in that: When multiple pickup channels simultaneously meet the preset triggering conditions, before performing cross-correlation ranging processing, the method further includes: Extract the baseline background noise scalar of multiple pickup channels within a preset historical silence time window; Based on the baseline background noise scalar and the acoustic energy characteristics within the current processing time window, the instantaneous signal-to-noise ratio characteristics of each pickup channel are calculated. A weighted fusion calculation is performed based on the instantaneous signal-to-noise ratio features and the ontological voiceprint matching features to generate a background purity index that characterizes the degree of data uncontamination. The channel with the highest background purity index value is selected from multiple pickup channels and designated as the reference anchor channel. The reference anchor channel is then configured as the crosstalk emission source channel. The cross-correlation ranging process and the target cancellation process are then performed on the remaining pickup channels one by one.
9. The band balance guidance method based on voiceprint recognition according to claim 8, characterized in that: Perform the cross-correlation ranging process and the target cancellation process on the remaining pickup channels one by one, including: Multiple pickup channels that meet the triggering conditions are constructed into a queue to be cleaned, the reference anchor point channel is configured as a crosstalk emission source channel, and the cross-correlation ranging processing and target cancellation processing are performed on the remaining target pickup channels in the queue to be cleaned one by one. The background purity index corresponding to the target pickup channel after the targeted cancellation process is recalculated and updated. The channel with the largest updated value is selected and locked as the secondary reference anchor channel and configured as the new crosstalk emission source channel. The second round of cancellation process is performed on the remaining channels to be cleaned. The second round of cancellation processing is performed iteratively until all target pickup channels in the queue to be cleaned have completed cancellation and reconstruction.
10. A band balance guidance system based on voiceprint recognition, characterized in that: Using the band balance guidance method based on voiceprint recognition as described in any one of claims 1-9, comprising: The data acquisition module is used to acquire initial audio data containing multiple pickup channels, and extract the acoustic energy features and body voiceprint matching features of the initial audio data within the current processing time window. The data processing module is used to calculate the first temporal rate of change of the acoustic energy features relative to the historical processing time window, and the second temporal rate of change of the body voiceprint matching features relative to the historical processing time window. The target channel determination module is used to perform orthogonal state deviation determination based on the first time-series change rate and the second time-series change rate. When the determination result meets the preset triggering condition, the pickup channel that has deviated is determined as the target pickup channel affected by external crosstalk interference. The crosstalk source determination module is used to extract the audio signals of the target pickup channel and the preset candidate pickup channel within a preset time delay constraint range, perform cross-correlation ranging processing, and determine the crosstalk emission source channel and the corresponding absolute propagation time delay parameter. The filter construction module is used to invert the spatial physical propagation distance based on the absolute propagation delay parameter, and to construct a spatial spectrum dissipation filter based on the spatial physical propagation distance and a preset nonlinear acoustic dissipation model. The crosstalk signal generation module is used to perform phase shift processing on the frequency domain signal of the crosstalk emission source channel using the absolute propagation delay parameter, and to perform frequency response deformation constraint processing on the frequency domain signal of the crosstalk emission source channel using the spatial spectrum dissipation filter, thereby generating a physical deformation crosstalk prediction signal. The pure parameter reconstruction module is used to perform targeted cancellation processing on the frequency domain signal of the target pickup channel and the physical deformation crosstalk prediction signal to reconstruct the pure acoustic energy parameters of the target pickup channel, and generate target balance guidance instructions based on the pure acoustic energy parameters.