Mixing console digital audio signal processing method, system, product and medium

Through time-frequency domain conversion and the establishment of a timbre feature fingerprint library, an audio feature coordination index system was constructed, which solved the problem of insufficient audio signal processing caused by fixed parameter presets in the mixing console, achieved adaptive optimization of audio signals, and improved sound quality.

CN120319263BActive Publication Date: 2025-09-26SHENZHEN SOUNDFIT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510814995.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-26
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

The audio signal processing of existing mixing consoles uses a fixed parameter preset method, which makes it difficult to accurately analyze the overall coordination, layering and spatial balance between multi-channel audio signals, resulting in poor overall sound quality.

Method used

The time-frequency feature map of the audio signal is obtained through time-frequency domain conversion, a timbre feature fingerprint library is established, an audio feature coordination index system is constructed, the audio processing parameters are calculated and dynamically adjusted in real time, and adaptive processing is performed in combination with the timbre feature fingerprint library.

Benefits of technology

It achieves accurate recognition and adaptive processing of multi-channel audio signals, improves overall coordination, layering and spatial balance, and ensures the optimization effect of audio signal processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120319263B_ABST
    Figure CN120319263B_ABST
Patent Text Reader

Abstract

A method, system, product, and medium for processing digital audio signals for a mixing console relate to the field of speech or sound processing. The method comprises: obtaining multiple audio input signals and performing time-frequency domain conversion, analyzing their frequency distribution patterns, extracting fundamental frequency and overtone structural features, and establishing a timbre feature fingerprint library. A preset standard audio signal is then obtained and a multidimensional feature vector is extracted, which is then matched with the fingerprint library to obtain a set of feature parameters. Based on this, a coordination index system is constructed, including overall sound coordination, part level clarity, and spatial sound and image balance. Adjustment values ​​are obtained by calculating the coordination index of the input signal in real time and comparing it with the standard signal. Finally, a target audio processing parameter combination, including equalization, dynamics, and spatial parameters, is calculated in combination with the fingerprint library matching features. Implementing this method can improve the sound quality of multiple audio signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of voice or sound processing, and in particular to a method, system, product and medium for processing digital audio signals of a mixing console. Background Art

[0002] With the rapid development of modern music production and live performance technologies, mixing consoles, as core equipment for audio signal processing, are playing an increasingly important role in music production, live performances, broadcasting, and television. Digital mixing consoles digitize multiple audio input signals, enabling precise adjustment and optimization of sound to meet the professional audio processing needs of diverse scenarios.

[0003] In related technologies, audio signal processing in mixing consoles uses fixed parameter presets, processing each channel using a pre-set parameter template. This solution uses pre-configured audio processing modules such as equalizers, compressors, and reverberators to process the input audio signal according to fixed parameter combinations, thereby achieving basic audio adjustment functions.

[0004] However, due to the differences in timbre characteristics between different audio signals, the preset fixed parameters are difficult to accurately analyze the overall coordination, layering and spatial balance between multi-channel audio signals, resulting in poor overall sound quality. Summary of the Invention

[0005] The present application provides a method, system, product and medium for processing digital audio signals of a mixing console, which are used to improve the sound quality of multi-channel audio signals.

[0006] In a first aspect, the present application provides a method for processing digital audio signals of a mixing console, which is applied to a digital audio signal processing system of a mixing console. The method comprises: obtaining multiple audio input signals, performing time-frequency domain conversion on the audio input signals to obtain a time-frequency feature graph; analyzing the frequency distribution law of each audio input signal based on the time-frequency feature graph, extracting the fundamental frequency and overtone structure characteristics of each audio input signal, and establishing a timbre feature fingerprint library based on the fundamental frequency and overtone structure characteristics; obtaining a preset standard audio signal, extracting a multidimensional audio feature vector from the preset standard audio signal, and matching the multidimensional audio feature vector with the timbre feature fingerprint library. An audio feature parameter group is obtained; an audio feature coordination index system is constructed based on the audio feature parameter group, and the audio feature coordination index system includes overall sound coordination, voice layer clarity, and spatial sound image balance; based on the audio feature coordination index system, the audio feature coordination index of multiple audio input signals is calculated in real time, and the audio feature coordination index is compared with the coordination index of a preset standard audio signal to obtain a signal adjustment value; based on the signal adjustment value, a target audio processing parameter combination is calculated in combination with a timbre feature fingerprint library, and the target audio processing parameter combination includes equalization parameters, dynamic parameters, and spatial parameters.

[0007] In the above embodiment, accurate extraction and identification of the unique timbre features of each audio signal are achieved through time-frequency domain conversion and the establishment of a timbre feature fingerprint library. Based on the feature matching of the timbre feature fingerprint library and the preset standard audio signal, an audio feature coordination index system is constructed, which includes the overall coordination of the sound, the clarity of the voice layer, and the spatial sound and image balance, thus breaking through the limitations of fixed parameter presets. The coordination index calculated in real time is dynamically compared with the standard signal, and the target parameter combination is calculated in combination with the matching features of the timbre feature fingerprint library, thereby achieving adaptive processing of audio signals with different timbre features, ensuring the optimization effect of multi-channel audio signals in terms of overall coordination, layering, and spatial balance.

[0008] In combination with some embodiments of the first aspect, in some embodiments, the audio feature coordination index system is constructed based on the audio feature parameter group, and the index system includes the steps of overall sound coordination, part level clarity and spatial sound image balance, specifically including: calculating the energy values ​​of different frequency bands in the audio feature parameter group, obtaining the energy ratio between adjacent frequency bands, and calculating the overall sound coordination index based on the energy ratio; extracting the part loudness value and frequency value in the audio feature parameter group, and calculating the part level clarity index corresponding to the overall sound coordination index; calculating the phase difference and amplitude difference of the channel signals based on the left and right channel signal values ​​in the audio feature parameter group, determining the sound image position distribution in combination with the part level clarity index, and obtaining the spatial sound image balance index.

[0009] In the above embodiment, the overall sound harmony index is calculated by calculating the energy ratios of different frequency bands. This index reflects the balance of spectral energy distribution. The degree of differentiation between different voices is analyzed by combining voice loudness and frequency values, and the voice layer clarity index is calculated to reflect the sense of hierarchy between voices. Furthermore, based on the phase and amplitude differences between the left and right channel signals, the sound image position distribution is determined in combination with the voice layer clarity, ultimately forming a complete spatial sound image balance index, which provides comprehensive coordination assessment capabilities for audio signal processing.

[0010] In combination with some embodiments of the first aspect, in some embodiments, the step of calculating the audio feature coordination index of multiple audio input signals in real time based on the audio feature coordination index system, and comparing the audio feature coordination index with the coordination index of a preset standard audio signal, specifically includes: according to the overall sound coordination calculation rules, the voice layer clarity calculation rules and the spatial sound image balance calculation rules in the audio feature coordination index system, the multiple audio input signals are segmented according to time intervals, and the coordination index of each signal is calculated; the coordination index of the corresponding time period is extracted from the preset standard audio signal, and the numerical comparison is performed with the coordination index of the multiple audio input signals to obtain a numerical comparison result; the signal adjustment value is determined according to the numerical comparison result and the matching characteristics of the timbre feature fingerprint library.

[0011] In the above embodiment, multiple audio input signals are processed in time intervals, and a coordination index is calculated for each signal segment, establishing a dynamic evaluation mechanism for audio signal processing. By comparing the coordination index values ​​with those of corresponding time segments in a preset standard audio signal and combining them with matching features from a timbre fingerprint library to determine signal adjustment values, this allows for precise adjustment of audio signal processing parameters, ensuring the consistency and accuracy of the processing results.

[0012] In combination with some embodiments of the first aspect, in some embodiments, after the step of calculating the target audio processing parameter combination based on the signal adjustment value and in combination with the timbre feature fingerprint library, the method also includes: extracting the characteristic change law of the preset standard audio signal and establishing a characteristic change record table; analyzing the adjustment direction of the target parameter combination according to the characteristic change record table and generating a parameter adjustment step; adjusting the target audio processing parameter combination according to the parameter adjustment step.

[0013] In the above embodiment, the characteristic variation patterns of a preset standard audio signal are extracted and a characteristic variation record table is established, providing a dynamic reference for parameter adjustment. The characteristic variation record table is used to analyze the adjustment direction of the target parameter combination and generate a parameter adjustment step size, enabling precise control of parameter adjustment. Using this parameter adjustment step size to numerically adjust the target audio processing parameter combination makes the audio processing parameter adjustment process smoother and avoids fluctuations in sound quality caused by sudden parameter changes.

[0014] In combination with some embodiments of the first aspect, in some embodiments, after the steps of obtaining multiple audio input signals and performing time-frequency domain conversion on the audio input signals to obtain a time-frequency feature graph, the method further includes: analyzing the timbre feature distribution in the time-frequency feature graph, and identifying the signal characteristics of different musical instruments and human voices based on the timbre feature distribution; separating the multiple audio input signals into independent audio tracks based on the signal characteristics; extracting the fundamental frequency and overtone structure characteristics of the independent audio tracks, and performing feature matching with the timbre feature fingerprint library; calculating the audio feature coordination index of the independent audio tracks, and obtaining the target parameter combination of each audio track.

[0015] In the above embodiment, the signal characteristics of different instruments and vocals are distinguished based on the timbre feature distribution in the time-frequency feature graph, and multiple audio input signals are separated into independent tracks. By extracting the fundamental frequency and overtone structure characteristics of the independent tracks for timbre feature fingerprint matching, and calculating the audio feature coordination index for each track separately, this achieves accurate recognition and independent processing of different types of audio signals, improving the level of sophistication of audio processing.

[0016] In combination with some embodiments of the first aspect, in some embodiments, after the step of calculating the target audio processing parameter combination based on the signal adjustment value and in combination with the timbre feature fingerprint library, the method also includes: extracting the frequency change characteristics and energy change characteristics in the time-frequency feature graph, and analyzing the change law of the characteristics; predicting the fundamental frequency change trend and overtone structure change trend of the next time period based on the change law and historical data of the audio feature parameter group; adjusting the values ​​of various indicators in the audio feature coordination index system based on the fundamental frequency change trend and overtone structure change trend; substituting the adjusted values ​​of various indicators into the calculation process of the target audio processing parameter combination to obtain a pre-adjusted parameter group including equalization parameters, dynamic parameters and spatial parameters.

[0017] In the above embodiment, the frequency and energy variations in the time-frequency feature graph are analyzed, and the fundamental frequency and overtone structure variation trends for the next time period are predicted in combination with historical data from the audio feature parameter group. Based on the prediction results, the values ​​of various indicators in the audio feature coordination index system are dynamically adjusted, and the adjusted values ​​of each indicator are applied to the parameter combination calculation, thereby enabling pre-adjustment of audio processing parameters and enhancing the real-time and continuity of audio processing.

[0018] In combination with some embodiments of the first aspect, in some embodiments, after the step of calculating the target audio processing parameter combination based on the signal adjustment value and in combination with the timbre feature fingerprint library, the method also includes: obtaining an ambient noise signal, converting the ambient noise signal into a noise time-frequency feature graph, and extracting the frequency distribution characteristics and energy distribution characteristics in the noise time-frequency feature graph; comparing the frequency distribution characteristics and energy distribution characteristics of the noise time-frequency feature graph with the time-frequency feature graph of the audio input signal to determine the noise frequency band and noise energy; constructing a noise suppression index based on the noise frequency band and noise energy, and adding the noise suppression index to the audio feature coordination index system; and adjusting the equalization parameters and dynamic parameters in the target audio processing parameter combination based on the value of the noise suppression index.

[0019] In the above-mentioned embodiment, an ambient noise signal is acquired and converted into a noise time-frequency feature map. Its frequency distribution characteristics and energy distribution characteristics are compared with the time-frequency feature map of the audio input signal to accurately identify the noise frequency band and noise energy. Based on the identification results, a noise suppression index is constructed and incorporated into the audio feature coordination index system. The equalization parameters and dynamic parameters in the target audio processing parameter combination are adjusted in a targeted manner, achieving accurate identification and suppression of ambient noise and improving the anti-interference ability of audio signal processing in complex environments.

[0020] In a second aspect, an embodiment of the present application provides a mixing console digital audio signal processing system, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the mixing console digital audio signal processing system to execute the method described in the first aspect and any possible implementation method of the first aspect.

[0021] In a third aspect, an embodiment of the present application provides a computer program product comprising instructions. When the above-mentioned computer program product is run on a mixing console digital audio signal processing system, the above-mentioned mixing console digital audio signal processing system executes the method described in the first aspect and any possible implementation method of the first aspect.

[0022] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium comprising instructions. When the instructions are executed on a mixing console digital audio signal processing system, the mixing console digital audio signal processing system executes the method described in the first aspect and any possible implementation of the first aspect.

[0023] It is understood that the mixing console digital audio signal processing system provided in the second aspect, the computer program product provided in the third aspect, and the computer storage medium provided in the fourth aspect are all used to execute the methods provided in the embodiments of this application. Therefore, the beneficial effects achieved by these methods can be referenced to the beneficial effects of the corresponding methods and will not be further elaborated here.

[0024] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0025] 1. This application realizes the precise extraction and identification of the unique timbre characteristics of each audio signal through time-frequency domain conversion and the establishment of a timbre feature fingerprint library. Based on the feature matching of the timbre feature fingerprint library and the preset standard audio signal, an audio feature coordination index system is constructed, which includes the overall coordination of the sound, the clarity of the voice layer, and the spatial sound and image balance, thus breaking through the limitations of fixed parameter presets. The coordination index calculated in real time is dynamically compared with the standard signal, and the target parameter combination is calculated in combination with the matching characteristics of the timbre feature fingerprint library, thereby realizing the adaptive processing of audio signals with different timbre characteristics, and ensuring the optimization effect of multi-channel audio signals in terms of overall coordination, layering and spatial balance.

[0026] 2. This application calculates the energy ratio of different frequency bands to obtain an overall sound coordination index, which reflects the balance of spectral energy distribution. The degree of differentiation between different parts is analyzed by combining the loudness and frequency values ​​of the parts, and the part level clarity index is calculated to reflect the sense of hierarchy between the parts. Furthermore, based on the phase difference and amplitude difference between the left and right channel signals, the sound image position distribution is determined in combination with the part level clarity, ultimately forming a complete spatial sound image balance index, which enables audio signal processing to have comprehensive coordination evaluation capabilities.

[0027] 3. This application establishes a dynamic evaluation mechanism for audio signal processing by segmenting multi-channel audio input signals into time intervals and calculating a coordination index for each signal segment. By comparing the coordination index values ​​with those of corresponding time segments in a preset standard audio signal and combining them with matching features from a timbre fingerprint library to determine signal adjustment values, this allows for precise adjustment of audio signal processing parameters, ensuring the continuity and accuracy of the processing results. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 This is a flow chart of a method for processing digital audio signals of a mixing console according to an embodiment of the present application;

[0029] Figure 2 is another flowchart of the method for processing digital audio signals of a mixing console in an embodiment of the present application;

[0030] Figure 3This is a schematic diagram of the physical device structure of the digital audio signal processing system of the mixing console in the embodiment of the present application. DETAILED DESCRIPTION

[0031] The terms used in the following examples of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of the present application, the singular expressions "a", "an", "above", "the", and "this" are intended to include plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to any or all possible combinations of one or more of the listed items.

[0032] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of the technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.

[0033] For ease of understanding, the application scenarios of the embodiments of the present application are introduced below.

[0034] Live broadcasts from large stadiums present a variety of complex audio processing requirements. Pre-event rehearsal recordings were conducted, resulting in a standard audio signal under ideal conditions. This signal includes clear commentary (-20dBFS, 200Hz-8kHz), moderate audience background noise (-35dBFS, 300Hz-3kHz), and live music (-25dBFS, 50Hz-16kHz). This standard audio signal has a smooth frequency response, a dynamic range of 90dB, a signal-to-noise ratio greater than 60dB, and a uniform spatial sound field. However, in actual live broadcasts, the ambient noise is complex and variable, ranging from -20dBFS to -80dBFS and distributed over a wide frequency range (20Hz-20kHz). Traditional fixed-parameter audio processing solutions struggle to adjust the live audio to a level close to the standard audio. Furthermore, spectral overlap between different sound sources, such as the mid-frequency overlap between audience noise and commentary, complicates the separation and enhancement of the audio signals. At the same time, the live audio system also needs to deal with the sound fuzziness caused by the venue's reverberation (RT60 reaches 2.5 seconds) and the uneven frequency response caused by the spatial acoustic characteristics.

[0035] A studio used a traditional audio processing system for live program production. Before the show, the production team collected standard audio from a standard host voice (-18dBFS, 200Hz-8kHz) and background music (-30dBFS, 50Hz-16kHz) in a non-interference-free environment. The audio exhibited excellent frequency response and a 90dB dynamic range. The system used a fixed 31-band graphic equalizer for frequency response correction, standard compressor settings (2:1 ratio, -18dB threshold) for dynamic range control, and preset reverb parameters (80% room size, 30ms pre-delay) for spatial effects. Midway through the show, low-frequency noise (100-300Hz, -45dBFS) generated by the studio's air conditioning system increased. The traditional system was unable to optimize its parameters based on the standard audio characteristics, resulting in the low-frequency noise masking the host's voice. Furthermore, the fixed compressor settings, which failed to adapt to the noise level, resulted in a pumping effect. Furthermore, the system's preset reverberation parameters couldn't be adjusted to the actual acoustic environment, resulting in an increased sense of muddiness in noisy environments. The system also experienced processing latency (approximately 50ms), which resulted in a distortion in the on-site monitoring experience. Ultimately, the program's audio quality fell significantly below standard audio quality, with clarity dropping from 0.85 to 0.65 and dynamic range compressed from 90dB to 60dB, severely impacting the program's performance.

[0036] A performance venue used the audio processing system described in this application for live sound reinforcement. Before the performance, the system acquired a standard audio clip under ideal conditions, consisting of the lead vocals (-15dBFS, 100Hz-10kHz), accompanying music (-20dBFS, 20Hz-20kHz), and ambient reverberation (-40dBFS). The system then performed feature analysis on the standard audio and established an audio feature coordination index system. At the performance venue, the system first detected three main types of noise through real-time analysis: background noise in the audience area (500-2kHz, -50dBFS), operating noise from stage equipment (above 4kHz, -60dBFS), and low-frequency noise from the air conditioning system (100-200Hz, -45dBFS). The system then compared these noise characteristics with those of the standard audio, constructed a noise suppression index, and calculated spectral masking and energy suppression values ​​across 24 critical frequency bands. Based on these indicators, the system automatically adjusts the settings of the 31-band parametric equalizer: the Q value in the low-frequency band is increased to 8.65, with a -6dB attenuation setting; an adaptive notch filter with a Q value of 12.5 is used in the mid-frequency band; and a high-shelf filter with a slope of -12dB / octave is used in the high-frequency band. The dynamic processor parameters are also optimized based on the dynamic characteristics of standard audio: the compression ratio in the low-frequency band is increased to 3:1, and the threshold is lowered to -24dB; 4:1 compression is used in the mid-frequency band with a fast response time of 1ms; and an expander is used in the high-frequency band. Through these adaptive adjustments, the system successfully adjusts the characteristics of the live audio to a level close to that of standard audio, improving the signal-to-noise ratio by 15dB, maintaining a sound clarity index above 0.85, and maintaining a dynamic range of 85dB, effectively solving sound quality issues in complex acoustic environments.

[0037] For ease of understanding, the following describes the process of the method provided by this implementation in combination with the above scenario. Figure 1 , which is a flow chart of a method for processing digital audio signals of a mixing console in an embodiment of the present application.

[0038] S101: Acquire multiple audio input signals, and perform time-frequency domain conversion on the audio input signals to obtain a time-frequency feature map.

[0039] Among them, multi-channel audio input signals represent multiple independent audio signal channels from different audio sources, including but not limited to musical instrument signals, human voice signals, etc.; time-frequency domain conversion refers to converting the time domain audio signal into a time-frequency representation that contains both time and frequency information; the time-frequency feature map represents a two-dimensional image representation of the frequency distribution characteristics of the audio signal at different time points.

[0040] This step begins after the system boots up and detects the connection of an audio input device. Specifically, the system collects audio signals from multiple channels through the audio input interface and performs analog-to-digital conversion on each channel to generate a digital audio signal. The digital audio signal is then framed and time-frequency analysis methods, such as the short-time Fourier transform (STFT), are applied to each frame to determine the signal's distribution characteristics on the time-frequency plane, forming a time-frequency feature map.

[0041] In some embodiments, time-frequency domain signal processing can be achieved by: optionally, first preprocessing the input signal, including signal normalization, noise reduction, and framing, then windowing each frame of the signal and performing an FFT transform, and finally arranging the transform results in chronological order to form a time-frequency graph; optionally, using a wavelet transform to perform multi-resolution analysis on the signal, and constructing a time-frequency feature graph using wavelet coefficients of different scales to better characterize the local characteristics of the signal. It is understood that other time-frequency analysis methods can also be used to achieve time-frequency representation of audio signals, which is not limited here.

[0042] After this step, the following steps are also included:

[0043] Analyze the timbre feature distribution in the time-frequency feature graph, and identify the signal characteristics of different musical instruments and human voices based on the timbre feature distribution.

[0044] In this step, the time-frequency feature map refers to the time-frequency distribution of the audio signal obtained through short-time Fourier transform (SFT); the timbre feature distribution includes the spectral envelope, transient characteristics, and overtone structure; and the signal characteristics refer to the typical acoustic characteristic parameters of different sound sources. The SFT uses a 2048-point FFT with a 25ms frame length, a 12.5ms frame shift, and a Hanning window, analyzing the frequency range from 20Hz to 20kHz. The system performs a time-frequency analysis on the input mixed audio signal. First, a SFT is used to generate a time-frequency feature map, obtaining the temporal evolution characteristics of the spectrum. In the frequency domain, the system divides the frequency range from 20Hz to 20kHz into 24 critical frequency bands and calculates the energy distribution of each band. For vocal signals, the system identifies its characteristic formant distribution (F1≈500Hz, F2≈1500Hz, F3≈2500Hz) and typical pitch range (80-1000Hz). For piano signals, the system identifies its wide spectrum (27.5Hz-4186Hz) and typical overtone structure (integer multiples of the fundamental frequency). For guitar signals, the system identifies the concentrated energy in the mid-band (80-1200Hz) and the unique plucked string transient characteristics (attack time <5ms). In the time domain, the system analyzes the envelope characteristics of various signals: the gradual transition characteristics of human voice (attack time 20-50ms), the rapid decay characteristics of piano (decay time 500-1000ms), and the moderate decay characteristics of guitar (decay time 200-500ms). The system further quantifies the timbre characteristics of different sound sources by calculating characteristic parameters such as Mel-Frequency Cepstral Coefficients (MFCC) and spectral centroid.

[0045] Separate multiple audio input signals into independent tracks based on signal characteristics.

[0046] In this step, the multi-channel audio input signal refers to mixed multi-source audio signals; the independent audio tracks refer to separated single-source audio signals; and signal separation utilizes a deep learning-based source separation algorithm. The system performs audio signal separation. First, the input mixed audio undergoes preprocessing, including signal normalization and frequency domain transformation. The system then employs a convolutional neural network model with a U-Net architecture for signal separation. The input layer receives the time-frequency graph of the mixed signal (with a 25ms frame length and a 12.5ms frame shift). A five-layer encoder-decoder structure extracts features and reconstructs the spectrograms of each sound source. In the time-frequency masking layer, the system generates a soft masking matrix for each target sound source, with masking values ​​ranging from 0 to 1. The masked spectrum is optimized using a Wiener filter to reduce artifacts during the separation process. Finally, the time-domain signal is reconstructed using an inverse short-time Fourier transform (ISFT) to produce independent audio tracks for vocals, piano, and guitar. A phase reconstruction algorithm maintains phase consistency between the tracks, and residual compensation techniques are employed to improve separation quality.

[0047] Extract the fundamental frequency and overtone structure features of independent audio tracks and perform feature matching with the timbre feature fingerprint library.

[0048] In this step, the fundamental frequency refers to the lowest frequency component of the audio signal; the overtone structure refers to the energy distribution of integer multiples of the fundamental frequency; and the timbre fingerprint library contains standard feature templates for various musical instruments and vocals. The system performs feature extraction and matching on each separated audio track. For each track, the system first extracts the fundamental frequency using the autocorrelation method, computes the spectrum using a 2048-point FFT, and identifies the peak corresponding to the fundamental frequency. For vocal tracks, the system extracts the pitch contour (fundamental frequency range 80-1000Hz) and formant structure (F1, F2, and F3 positions and bandwidth). For piano tracks, the system analyzes the fundamental frequency (27.5-4186Hz) and overtone decay characteristics (energy ratios of the initial eight overtones) of all 88 keys. For guitar tracks, the system extracts the fundamental frequency (82-698Hz) and overtone structure of all six strings. The system matches the extracted feature vectors with templates in the timbre fingerprint library, using Euclidean distance to measure feature similarity. The feature template with the highest similarity is selected as the reference.

[0049] Calculate the audio feature coordination index of independent audio tracks and obtain the target parameter combination of each audio track.

[0050] In this step, the audio feature coordination indicators include spectral balance, dynamic consistency, and spatial correlation; the target parameter combination includes parameter settings for the equalizer, dynamic processor, and spatial processor. The system performs feature coordination analysis and parameter calculation. For each audio track, the system first calculates the spectral balance indicator: by performing variance analysis on the energy distribution of 24 critical frequency bands, the smoothness of the frequency response curve is calculated. The spectral balance of vocal tracks in the 200-8000Hz range must be greater than 0.85, the spectral balance of piano tracks in the full frequency range must be greater than 0.80, and the spectral balance of guitar tracks in the 80-5000Hz range must be greater than 0.75. The dynamic consistency indicator is obtained by calculating the rate of change of the short-term energy envelope, including attack time, release time, and dynamic range parameters. The spatial correlation indicator is obtained by analyzing the correlation coefficient and sound image localization characteristics of the stereo signal. Based on these metrics, the system calculates processing parameters for each track: the center frequency, Q value, and gain settings for the 31-band parametric equalizer; the compression ratio, threshold, attack time, and release time for the dynamics processor; and the panning position, width, and depth parameters for the spatial processor. All parameters are optimized based on the track's characteristic harmony metrics to ensure optimal sound quality after processing.

[0051] S102: Analyze the frequency distribution of each audio input signal based on the time-frequency feature graph, extract the fundamental frequency and overtone structure features of each audio input signal, and establish a timbre feature fingerprint library based on the fundamental frequency and overtone structure features.

[0052] Among them, the frequency distribution law represents the distribution characteristics of the audio signal spectrum energy in different frequency components; the fundamental frequency refers to the lowest basic frequency component in the audio signal; the overtone structure characteristics refer to the frequency and amplitude relationship of each harmonic component in the signal; the timbre feature fingerprint library is a database that stores different sound source feature templates.

[0053] This step is performed after obtaining the time-frequency feature map. Specifically, the system first performs peak detection on the time-frequency feature map to locate the main frequency components. It then determines the signal's fundamental frequency through methods such as autocorrelation analysis and identifies the individual harmonic components based on integer multiples of the fundamental frequency. The system analyzes the amplitude ratios and time-varying characteristics of each harmonic, extracting parameters that characterize the timbre. These parameters are organized into feature vectors and stored in a fingerprint library.

[0054] In some embodiments, feature extraction and fingerprint library construction can be achieved by the following methods: Optionally, using cepstrum analysis to extract the fundamental frequency, using harmonic-noise decomposition to extract the overtone structure, and combining energy envelope features to construct a timbre feature fingerprint; Optionally, using deep learning methods to directly learn timbre feature representations from time-frequency graphs, and automatically extracting features and establishing a fingerprint library using models such as convolutional neural networks. It is understood that other feature extraction methods can also be used to construct a timbre feature fingerprint library, which is not limited here.

[0055] S103: Obtain a preset standard audio signal, extract a multidimensional audio feature vector from the preset standard audio signal, match the multidimensional audio feature vector with a timbre feature fingerprint library, and obtain an audio feature parameter group.

[0056] Among them, the preset standard audio signal refers to a noise-free, high-quality audio signal recorded under an ideal environment; the multidimensional audio feature vector includes parameters such as spectral envelope, energy distribution, transient characteristics and spatial information; the timbre feature fingerprint library is a data set containing different types of audio feature templates; the audio feature parameter group refers to a complete parameter set used for subsequent audio processing.

[0057] This step is performed after the timbre feature fingerprint library is established. Specifically, the system first selects a standard audio signal with a style similar to the current input signal from a preset standard audio database as a reference sample. It then performs time-frequency analysis on the standard audio signal, extracting multidimensional feature parameters including spectral centroid, spectral flow, overtone ratio, timbre brightness, etc. to form a feature vector. The system then performs similarity matching on this feature vector with templates in the established timbre feature fingerprint library, selects the most matching feature template, and combines the characteristics of the current signal to generate a feature parameter group containing parameters such as timbre, dynamics, and space.

[0058] In some embodiments, the feature extraction and matching process can be implemented in a variety of ways: optionally, first perform multi-scale decomposition on the standard audio signal, extract the energy distribution characteristics of different frequency bands respectively, then calculate the statistical characteristics of the signal such as mean, variance, etc., and finally extract the dynamic characteristics of the signal such as envelope characteristics, and combine these features to form a feature vector for matching; optionally, use a machine learning method to automatically learn the representation of audio features by training a deep neural network, the network input is the time-frequency diagram of the audio signal, and the output is the corresponding feature vector, and then perform feature matching through similarity calculation. It is understandable that other feature extraction and matching methods can also be used to obtain the audio feature parameter group, which is not limited here.

[0059] For example, during the rehearsal phase before a concert, the system acquires a preset standard audio signal. This signal, captured in a professional recording studio (with background noise below -80dBFS), contains the lead vocals (-15dBFS, 100Hz-10kHz), instrumental accompaniment (-20dBFS, 20Hz-20kHz), and ambient reverberation (-40dBFS, RT60 of 1.2 seconds). The system extracts feature vectors from this standard audio: first, a 1024-point FFT is performed to obtain energy distribution curves for 31 frequency bands. The system then calculates the spectral centroid (3.2kHz) and spectral tilt (-2.5dB / octave). Envelope features, including attack time (5ms), release time (150ms), and dynamic range (90dB), are extracted. Spatial characteristics are analyzed to determine the sound image width (0.8) and depth (0.7). The system then matches these feature vectors against a library of timbre fingerprints containing over 1,000 preset audio feature templates. By calculating the Euclidean distance, the system finds the closest feature template and extracts the corresponding set of audio feature parameters: 31-band parametric equalizer settings (center frequency, Q value, and gain), dynamics processor parameters (compression ratio 2:1, threshold -18dB, attack time 5ms, release time 150ms), and spatial processing parameters (stereo correlation 0.85, sound image position ±30°). These parameters extracted from the standard audio serve as a reference for subsequent real-time audio processing. The system compares the feature vector of the real-time audio with these standard parameters to achieve precise audio optimization.

[0060] S104: Construct an audio feature coordination index system based on the audio feature parameter group.

[0061] Among them, the audio feature coordination index system represents a multi-dimensional evaluation system for evaluating the overall sound quality of audio signals; the overall coordination of sound refers to the balance index of the energy distribution in each frequency band of the audio signal; the voice level clarity represents the degree of distinguishability between different voice parts; and the spatial sound image balance refers to the rationality index of sound image positioning in the stereo signal.

[0062] This step is performed after obtaining the audio feature parameter set. Specifically, the system constructs a multi-dimensional evaluation index system based on the spectral, dynamic, and spatial parameters in the audio feature parameter set. This system establishes a quantitative standard for evaluating the overall quality of the audio signal by analyzing parameters such as frequency band energy distribution, voice loudness relationships, and channel signal characteristics. The system then combines these parameters according to preset weights to form a comprehensive index system that can be used to guide subsequent signal processing.

[0063] In some embodiments, the index system can be constructed in a variety of ways: optionally, first establish an evaluation model based on spectral energy distribution, calculate the energy ratio of adjacent frequency bands, then analyze the loudness hierarchy of the voice signals, and finally construct a spatial evaluation index based on the phase and amplitude characteristics of the channel signals; optionally, adopt a data-driven approach to analyze the characteristic distribution patterns of a large number of high-quality audio samples to establish an evaluation system based on statistical models, including energy distribution models, voice relationship models, and spatial feature models. It is understandable that other methods can also be used to construct an audio feature coordination index system, which is not limited here.

[0064] This step specifically includes:

[0065] Calculate the energy values ​​of different frequency bands in the audio feature parameter group, obtain the energy ratio between adjacent frequency bands, and calculate the overall sound coordination index based on the energy ratio.

[0066] In this step, the frequency band energy value refers to the root mean square energy value of the 31 frequency bands; the energy ratio indicates the energy comparison between adjacent frequency bands; and the overall sound harmony index indicates the degree of balance in the spectral energy distribution, ranging from 0 to 1. The system performs a spectral energy analysis process: first, the audio signal is analyzed in 31 frequency bands, divided into 1 / 3 octave intervals, covering the range of 20Hz-20kHz. A 2048-point FFT is used to calculate the energy value of each frequency band, and an exponential averaging method (averaging time constant 125ms) is used to obtain a stable energy reading. For the energy Ei of the i-th frequency band, its ratio to the energy of the adjacent frequency band is calculated: the upper ratio Ri_up = Ei / Ei+1, and the lower ratio Ri_down = Ei / Ei-1. Based on these ratios, the frequency band energy gradient Gi = (Ri_up + Ri_down) / 2 is calculated. When Gi falls between 0.8 and 1.25, the frequency band is considered energy-harmonious with its adjacent bands. The system calculates the coordination status of all frequency bands and calculates the ratio of coordinated frequency bands to the total number of frequency bands to obtain an overall coordination index. The system also performs variance analysis on the energy ratio sequence. A variance value less than 0.5 indicates a stable spectrum energy distribution. The final overall coordination index is calculated using a weighted average: coordination index = 0.6 × proportion of coordinated frequency bands + 0.4 × (1 - variance value).

[0067] The voice loudness value and frequency value in the audio feature parameter group are extracted, and the voice level clarity index corresponding to the overall sound coordination index is calculated.

[0068] In this step, the voice loudness value refers to the loudness level of each voice; the frequency value refers to the frequency center and bandwidth of the voice; and the voice layer clarity index indicates the degree of distinguishability of different voices in terms of loudness and frequency. The system performs a voice analysis process: using the loudness calculation model (ISO532B) to calculate the loudness level of each voice and simultaneously extract the frequency characteristics of the voice. For N voices, an N×N voice discrimination matrix D is constructed. The matrix element Dij represents the discrimination between the i-th voice and the j-th voice, calculated as: Dij = α × |Li-Lj| / Lmax + β × |Fi-Fj| / Fmax, where Li and Lj are loudness values, Fi and Fj are frequency center values, and α = 0.6 and β = 0.4 are weighting coefficients. The diagonal elements of the voice discrimination matrix are set to 1, indicating complete discrimination. The system then calculates the average of the off-diagonal elements of the matrix to obtain the overall voice discrimination. When the loudness difference between adjacent parts is greater than 6dB and the frequency interval is greater than the critical bandwidth, the parts are considered to have good layering. The system uses a weighted combination of part distinction and overall coordination to derive the part layer clarity index: clarity = 0.7 × part distinction + 0.3 × overall coordination.

[0069] According to the left and right channel signal values ​​in the audio feature parameter group, the phase difference and amplitude difference of the channel signals are calculated, and the sound image position distribution is determined in combination with the voice layer clarity index to obtain the spatial sound image balance index.

[0070] In this step, the phase difference refers to the time delay between the left and right channel signals; the amplitude difference refers to the energy difference between the channels; the image position distribution indicates the spatial positioning of the sound source in the stereo field; and the spatial image balance index reflects the uniformity of the image distribution. The system performs a spatial image analysis process: short-term cross-correlation analysis is used to calculate the phase difference between the left and right channels, with a time window of 20ms and a step of 10ms. For a frequency component f, the phase difference φ(f) is converted to an image azimuth angle θ(f) = arcsin(φ(f) × f / c), where c is the speed of sound. The image intensity is calculated by calculating the energy ratio between the left and right channels: I(f) = 10 × log(EL(f) / ER(f)). In the frequency domain, the system divides the frequency range from 20Hz to 20kHz into 24 critical frequency bands and calculates the image position vector P(k) = [θ(k), I(k)] for each band. The image distribution map is generated by statistically accumulating energy at different azimuth angles, ranging from -30° to +30°. The system calculates the standard deviation σ and skewness S of the image distribution to assess spatial uniformity. Combined with the voice layer clarity index C, the final spatial image balance is calculated as: E = w1 × (1-σ / 30) + w2 × (1-|S|) + w3 × C, where weighting coefficients w1 = 0.4, w2 = 0.3, and w3 = 0.3. When E is greater than 0.8, the image distribution has good spatial balance.

[0071] S105 , calculating audio feature coordination indices of multiple audio input signals in real time based on the audio feature coordination index system, comparing the audio feature coordination indices with coordination indices of preset standard audio signals, and obtaining signal adjustment values.

[0072] Among them, real-time calculation means continuous indicator calculation and update during the continuous input of audio signals; audio feature coordination index refers to the quantitative evaluation value calculated based on the index system, including the overall sound coordination index, the voice layer clarity index and the spatial sound image balance index; the coordination index of the preset standard audio signal represents the evaluation value of high-quality reference audio under the same index system; the signal adjustment value refers to the quantitative value of the parameter that needs to be adjusted obtained through index comparison, including the equalization adjustment value, dynamic adjustment value and spatial adjustment value.

[0073] This step is continuously executed after the audio feature coordination index system is established. Specifically, the system first segments the input multi-channel audio signals according to the preset time window, and calculates the index values ​​of the three dimensions of overall sound coordination, voice level clarity, and spatial sound image balance for each signal segment. During the calculation process, the system analyzes the energy distribution characteristics of the frequency band to calculate the coordination index, calculates the clarity index through the voice loudness and frequency characteristics, and calculates the balance index based on the phase and amplitude relationship of the channel signal. The system then extracts the index value of the corresponding time period from the preset standard audio signal, compares the calculated index value with the standard index value in multiple dimensions, and calculates the difference value of each dimension through the set threshold and weight calculation. Finally, the specific value of the signal adjustment is determined based on the matching result of the difference value and the timbre feature.

[0074] In some embodiments, the real-time index calculation and comparison process can be implemented in a variety of ways: optionally, the input signal is first adaptively segmented, the analysis window length is dynamically adjusted according to the changing characteristics of the signal, and then the weighted average method is used to calculate the index value of each dimension, followed by a comparison algorithm based on fuzzy logic for index comparison, and finally the optimal adjustment value is predicted by a neural network model; optionally, a parallel computing architecture is used to simultaneously process signals of multiple time windows, and a frequency domain decomposition method is used to extract features and calculate index values ​​for each window, and then a recursive optimization algorithm is used to compare indicators, and finally the calculation method of the adjustment value is dynamically optimized based on the historical adjustment effect. It is understandable that other methods can also be used to implement the real-time index calculation and comparison process, which are not limited here.

[0075] This step specifically includes:

[0076] According to the calculation rules of overall sound coordination, voice layer clarity and spatial sound image balance in the audio feature coordination index system, multiple audio input signals are segmented according to time intervals, and the coordination index of each signal segment is calculated.

[0077] Among them, the coordination index system includes three-dimensional calculation rules; the time interval segmentation adopts a 20ms analysis window and a 10ms step interval; the coordination index of each signal segment includes three values: overall coordination, hierarchical clarity and spatial balance. In this step, the system performs a segmented coordination analysis process. First, the input multi-channel audio signal is segmented, and each analysis window contains 1024 sampling points (sampling rate 48kHz). For each time window, the system performs three-dimensional index calculations: (1) Overall coordination calculation: divide the frequency band by a 31-band parametric equalizer, calculate the energy ratio of adjacent frequency bands (Ei / Ei+1), count the proportion of frequency bands with energy ratios falling within the range of 0.8-1.25, calculate the variance of spectral energy distribution, and obtain the overall coordination Ct; (2) Hierarchical clarity calculation: extract the loudness level Li and frequency center Fi of each voice part, and construct an N×N voice part distinction matrix, with the matrix element Dij=0.6×|Li -Lj| / Lmax+0.4×|Fi-Fj| / Fmax, calculate the average value of the off-diagonal elements to obtain the layer clarity Lt; (3) Spatial balance calculation: Analyze the phase difference φ(f) and energy ratio I(f) between the left and right channels, calculate the sound image position vector P(k)=[θ(k), I(k)] on 24 critical frequency bands, and calculate the spatial balance St=0.4×(1-σ / 30)+0.3×(1-|S|)+0.3×Lt through the standard deviation σ and skewness S of the sound image distribution. The system combines these three indicators into a coordination index vector [Ct, Lt, St] as the feature representation of this time window.

[0078] A coordination index of a corresponding time period is extracted from a preset standard audio signal, and numerically compared with the coordination indexes of multiple audio input signals to obtain a numerical comparison result.

[0079] In this step, the preset standard audio signal refers to ideal audio recorded in a noise-free environment; the corresponding time period refers to a time segment with the same musical content; and the numerical comparison includes calculating the differences in each dimension's indicators. The system performs the indicator comparison process. For each 20ms analysis window, the system locates the corresponding time position in the standard audio and extracts the standard audio's harmony indicator vector [Cs, Ls, Ss]. The system calculates the differences between the input signal and the standard signal in three dimensions: overall harmony difference ΔC = |Ct-Cs|, layer clarity difference ΔL = |Lt-Ls|, and spatial balance difference ΔS = |St-Ss|. For any time window i, the system calculates a weighted difference value Di = 0.4×ΔC + 0.35×ΔL + 0.25×ΔS. When Di exceeds a threshold of 0.2, the time window is marked as requiring adjustment. The system also records the specific difference direction and value in each dimension to provide a basis for subsequent parameter adjustments.

[0080] The signal adjustment value is determined based on the numerical comparison result and the matching characteristics of the timbre feature fingerprint library.

[0081] In this step, the numerical comparison results include the difference values ​​and directions for each dimension; the timbre feature fingerprint library contains standardized adjustment parameter templates; and the signal adjustment values ​​refer to the specific parameter settings of each processing module. The system performs the parameter adjustment calculation process. For each time window requiring adjustment, the system first searches the timbre feature fingerprint library for the most matching feature template. The matching process uses the Euclidean distance metric, and the feature template with the smallest distance is selected as the reference. Based on the difference analysis results, the system calculates specific adjustment parameters: (1) Frequency response adjustment: When ΔC indicates spectrum imbalance, the gain adjustment value of the 31-band equalizer is calculated based on the energy ratio: gi = -10 × log (Ei / Es,i), where Ei and Es,i are the frequency band energies of the current signal and the standard signal, respectively; (2) Dynamic range adjustment: When ΔL indicates insufficient layering, the compressor parameters are calculated: compression ratio r = 1 + 2 × ΔL, threshold th = -18-12 × ΔLdB; (3) Spatial processing adjustment: When ΔS indicates uneven spatial distribution, the sound image adjustment parameters are calculated: azimuth angle offset θadj = -30 × ΔS × sign (S) degrees, width adjustment wadj = 0.5 × ΔS. The system combines these adjustment parameters into complete processing instructions and applies them to audio signal processing.

[0082] S106 , calculating a target audio processing parameter combination based on the signal adjustment value and in combination with the timbre feature fingerprint library.

[0083] Among them, the signal adjustment value represents the specific numerical parameter that needs to be adjusted for the audio signal; the matching feature of the timbre feature fingerprint library refers to the parameter information contained in the feature template that is closest to the current audio signal; the target audio processing parameter combination represents the parameter set used to optimize the audio signal, including equalization parameters, dynamic parameters and spatial parameters; equalization parameters are used to adjust the frequency response characteristics; dynamic parameters are used to control the dynamic range of the signal; and spatial parameters are used to optimize the stereo image effect.

[0084] This step is performed after obtaining the signal adjustment value. Specifically, the system first analyzes the correlation between the signal adjustment value and the matching features in the timbre feature fingerprint library to establish a parameter mapping model. Then, based on the adjustment requirements of different frequency bands and the time-varying characteristics of the timbre features, the equalization parameters of each frequency band are calculated, including the center frequency, gain, and Q value. At the same time, based on the adjustment requirements of the dynamic range and the transient characteristics of the signal, the system calculates dynamic processing parameters such as compression ratio and threshold. In addition, the system also needs to calculate stereo processing parameters, including sound image position, width, and depth parameters, based on the adjustment requirements of the spatial sound image and the correlation of the channel signals.

[0085] In some embodiments, the parameter combination calculation process can be implemented in a variety of ways: optionally, first construct a parameter optimization model based on an expert system, establish a rule base based on historical processing experience, then determine the initial parameter values ​​through rule reasoning, then use a genetic algorithm to optimize the parameter combination, and finally ensure the stability of the parameters through cross-validation; optionally, use a deep reinforcement learning method to model the parameter adjustment process as a Markov decision process, continuously optimize the parameter selection strategy through interaction with the environment, and combine a multi-objective optimization algorithm to balance different processing goals, and finally obtain the optimal parameter combination. It is understandable that other methods can also be used to implement the calculation of the target audio processing parameter combination, which is not limited here.

[0086] The following is a more detailed description of the process of the method provided by this implementation. Figure 2 , is another flow chart of the method for processing digital audio signals of a mixing console in an embodiment of the present application.

[0087] S201. Calculate a target audio processing parameter combination based on the signal adjustment value and in combination with a timbre feature fingerprint library. The target audio processing parameter combination includes an equalization parameter, a dynamic parameter, and a spatial parameter.

[0088] Among them, the signal adjustment value represents the parameter value to be adjusted obtained after feature comparison; the timbre feature fingerprint library stores the pre-established timbre feature templates; the equalization parameters include the center frequency, Q value and gain value of each frequency band; the dynamic parameters include compression ratio, threshold, attack time and release time; the spatial parameters include sound image positioning, width and depth parameters.

[0089] The system first extracts the feature template that best matches the current signal from the timbre feature fingerprint library. The template contains a preset timbre type, spectrum envelope, and dynamic characteristics. Based on the extracted feature template, the system calculates the target frequency response curve for each frequency band. Specifically, it performs a 1 / 3 octave analysis of the spectrum and calculates the target gain values ​​for 30 frequency bands in the range of 31.5Hz-16kHz. For dynamic parameters, the system determines the compression threshold and compression ratio by calculating the peak level and root mean square level of the signal, and sets the compression curve in the range of -20dB to 0dB. For spatial parameters, the system analyzes the phase difference and level difference between the left and right channels and calculates the sound image positioning parameters within the range of ±60°. The system combines these parameters according to the preset priority to generate a target processing parameter combination containing 93 specific parameters.

[0090] S202: Analyze the timbre feature distribution in the time-frequency feature graph, and identify signal features of different musical instruments and human voices based on the timbre feature distribution.

[0091] Among them, the time-frequency feature diagram represents the spectral characteristics of the audio signal changing over time; the timbre feature distribution refers to the energy distribution pattern of different sound sources on the time-frequency plane; the signal characteristics include fundamental frequency, overtone structure and transient characteristics.

[0092] The system performs adaptive segmented analysis on the time-frequency feature graph, using a 32ms Hamming window with a step size of 16ms for each segment. The timbre feature parameters of each time window are extracted by calculating the spectral centroid, spectral flow, and sub-band energy ratio. The system uses Mel-frequency cepstral coefficients (MFCC) to extract 20 timbre feature coefficients and combines them with the fundamental pitch detection algorithm to identify the pitch of each part. By analyzing the overtone structure, the system identifies the characteristic frequency components of different instruments. For example, the piano has characteristic overtones of 2096Hz and 3144Hz at a fundamental frequency of 262Hz (C4). For human voice signals, the system identifies vowel features through resonance peak tracking and identifies consonant segments in combination with zero-crossing rate analysis. The system establishes a multidimensional discriminant model based on these features to achieve feature classification of different sound sources.

[0093] S203: Separate the multi-channel audio input signals into independent audio tracks according to signal characteristics.

[0094] Among them, the independent audio track represents the separated single sound source signal; signal separation refers to the process of extracting different sound source components from the mixed signal.

[0095] The system uses a deep learning-based sound source separation algorithm and a convolutional neural network with a U-Net structure to separate the signal. The input layer receives the time-frequency feature map and extracts hierarchical features through a 5-layer encoder-decoder structure. Each encoder layer contains two 3×3 convolutional layers and a maximum pooling layer, and the decoder uses transposed convolution for upsampling. The skipconnection structure preserves detailed features. The system applies Wiener filtering to the spectrum map output by the network to reconstruct the time domain signal. The processing uses a 2048-point FFT with 50% overlap, and the output is an independent audio track with a sampling rate of 48kHz. The system performs phase correction on the separated signal to eliminate the phase distortion introduced by the separation process and maintain the time alignment of the original signal. The final output independent audio track maintains the dynamic range and frequency characteristics of the original signal.

[0096] S204: Extract the fundamental frequency and overtone structure features of the independent audio track, and perform feature matching with the timbre feature fingerprint library.

[0097] Among them, the fundamental frequency represents the basic frequency component of the pitch; the overtone structure characteristics include the frequency and energy distribution of each harmonic; feature matching refers to the process of calculating the similarity between the extracted features and the preset template.

[0098] The system performs autocorrelation analysis on each individual audio track, using a Hamming window with a frame length of 1024 points and calculating the fundamental frequency using the autocorrelation function. For vocal signals, the system searches for the fundamental frequency within the 70-400Hz range; for instrumental signals, the system expands the search range to 20-2000Hz. Using a harmonic-noise decomposition algorithm, the system extracts the first 16 harmonic components and calculates the energy ratio of each harmonic relative to the fundamental frequency to form a harmonic feature vector. Simultaneously, the system calculates the frequency ratio of adjacent harmonics to identify changes in the harmonic structure. The system then calculates the cosine similarity between the extracted feature vectors and templates in the timbre feature fingerprint library, selecting those with a similarity exceeding 0.85 as matching results. When a significant change in the harmonic structure is detected (a change in the energy ratio of adjacent harmonics exceeding 6dB), the system updates the feature matching results.

[0099] S205: Calculate the audio feature coordination index of the independent audio tracks to obtain the target parameter combination of each audio track.

[0100] Among them, the audio feature coordination indicators include spectral balance, dynamic consistency and spatial correlation; the target parameter combination includes equalization, compression and spatial processing parameters.

[0101] The system calculates the coordination index for each independent audio track separately. In terms of spectral balance, the system divides the frequency range into five sub-bands: low frequency (20-200Hz), mid-low frequency (200-800Hz), mid-frequency (800-3kHz), mid-high frequency (3-8kHz) and high frequency (8-20kHz), and calculates the energy ratio of adjacent sub-bands and the change trend of the spectral centroid. Dynamic consistency is quantified by calculating the rate of change of the signal's peak factor (PeakFactor) and root mean square level (RMS). The system sets the target range of the peak factor to 8-12dB. Spatial correlation is evaluated by calculating the cross-correlation function and phase difference spectrum of the stereo signal to determine the optimal sound image position within the range of ±45°. Based on these indicators, the system generates a target parameter combination consisting of 31 frequency band equalization parameters, 4 sets of dynamic processing parameters and 3 spatial processing parameters.

[0102] S206: Extract frequency change features and energy change features from the time-frequency feature graph, and analyze the change patterns of the features.

[0103] Among them, the frequency change characteristics represent the changes in the signal frequency components over time; the energy change characteristics refer to the time-varying characteristics of the signal energy distribution; and the change law refers to the periodicity or trend of these characteristics.

[0104] The system continuously analyzes time-frequency feature maps using a 1024-point FFT and 75% frame overlap, resulting in a temporal resolution of 5.33ms. The system calculates the spectral centroid and spectral slope of each frame to generate a frequency variation curve. The system describes energy variation characteristics using a time series of subband energy ratios and constructs an energy envelope model by calculating the mean and standard deviation of 20 adjacent frames. To extract long-term variation patterns, the system performs wavelet decomposition on the feature sequence and multiresolution analysis using a 5-level Daubechies wavelet. The system fits feature variation trends using an autoregressive model and calculates the periodicity of frequency variations (such as rhythmic characteristics) and the envelope characteristics of energy variations (such as dynamic changes). For sudden changes, the system detects and marks them by setting a relative threshold of 0.5.

[0105] S207 : Predicting the fundamental frequency change trend and the overtone structure change trend in the next time period based on the change rule and the historical data of the audio feature parameter group.

[0106] Among them, the change pattern represents the frequency and energy change characteristics extracted from the time-frequency feature diagram; the historical data of the audio feature parameter group refers to the feature parameter records in the past 100ms; the fundamental frequency change trend refers to the change direction and amplitude of the fundamental frequency in the next 50ms; the overtone structure change trend includes the expected changes in the harmonic frequency ratio and energy ratio.

[0107] The system uses a long short-term memory (LSTM) network to perform predictive analysis on feature sequences. The input layer receives a feature sequence of 20 time frames, including the fundamental frequency trajectory and the energy ratios of 16 harmonic components. The LSTM network contains two hidden layers, each with 128 neurons, and uses a tanh activation function. The system normalizes the input features, standardizing the fundamental frequency to the range of 0-1 and compressing the harmonic energy ratios to the range of -60dB to 0dB through logarithmic transformation. The network training uses the root mean square error as the loss function and uses the Adam optimizer for parameter updates. The system obtains feature changes for the next 10 time frames through recursive prediction, and applies a Kalman filter to the prediction results for smoothing to obtain a continuous trend curve.

[0108] S208. Adjust the values ​​of various indicators in the audio feature coordination indicator system based on the fundamental frequency change trend and the overtone structure change trend.

[0109] Among them, the index values ​​include the quantitative values ​​of the overall coordination of sound, the clarity of the voice levels and the balance of spatial sound and image; index adjustment refers to the process of correcting the index calculation parameters according to the predicted trend.

[0110] The system establishes an indicator adjustment model based on a deep neural network. This model adopts a three-layer fully connected structure, with the input layer receiving the predicted trend feature vector. The first hidden layer contains 256 neurons and uses the ReLU activation function to process the fundamental frequency variation characteristics; the second hidden layer contains 128 neurons to process the overtone structure variation characteristics. The system divides the overall coordination index into five frequency band sub-indicators and adjusts the weight coefficient of each sub-indicator based on the predicted frequency changes. The voice level clarity index adapts to signal changes by dynamically adjusting the detection threshold (ranging from -40dB to -20dB). The spatial sound image balance index dynamically adjusts the phase correlation judgment threshold (ranging from 0.6 to 0.9) based on the predicted phase relationship changes. The system applies exponential smoothing to the adjusted index values ​​to ensure the continuity of the index changes.

[0111] S209 , substituting the adjusted values ​​of various indicators into the calculation process of the target audio processing parameter combination to obtain a pre-adjusted parameter group including equalization parameters, dynamic parameters, and space parameters.

[0112] The calculation process of the target audio processing parameter combination refers to the mapping method of converting the index value into specific processing parameters; the pre-adjusted parameter group includes 31 frequency band equalization parameters, 4 groups of dynamic processing parameters and 3 spatial processing parameters.

[0113] The system uses a gradient descent algorithm to optimize parameter mapping. The objective function combines evaluation metrics for equalization performance, dynamics, and spatial positioning. Equalization parameters are calculated using 31 1 / 3-octave filters with center frequencies ranging from 20Hz to 20kHz and a fixed Q value of 4.32. The system calculates the gain of each filter based on overall coordination, limited to ±12dB. Dynamic parameters include compression ratio (1:1 to 10:1), threshold (-40dB to 0dB), attack time (0.1ms to 50ms), and release time (10ms to 1000ms). Spatial parameters are calculated by calculating the sound image position (-60° to +60°), width (0 to 1), and depth (0 to 1). The system limits and smoothes these calculated parameters to ensure smooth parameter changes.

[0114] Taking a vocal signal as an example, the system performs parameter calculations. During the equalization phase, the system detects a harmonicity index of 0.75 (standard value: 0.85) in the 200-800Hz frequency band and a clarity index of 0.65 (standard value: 0.80) in the 2kHz-4kHz frequency band. Based on these values, the system sets the center frequency in the 200-800Hz band to 400Hz and a Q value of 4.32, calculating a gain of +3dB using the formula: Gain = 10×log(0.85 / 0.75). For the 2kHz-4kHz band, the system sets the center frequency to 3kHz and a Q value of 4.32, calculating a gain of +4.5dB using the formula: Gain = 10×log(0.80 / 0.65). During the dynamic parameter calculation phase, the system detects a dynamic range index of 0.60 (standard value: 0.75) and a crest factor of 15dB (target value: 12dB). Based on this calculation, the compression ratio is 2.5:1 (calculated using the formula: 1 + (15 - 12) / 2), the compression threshold is set to -18dB (based on the -3dB setting of the peak level), the attack time is set to 20ms, and the release time is set to 150ms. During the spatial parameter calculation phase, the system detected a stereo correlation index of 0.85 (standard value: 0.90) and a sound image localization index of 0.70 (standard value: 0.85). Based on this, the sound image position adjustment value is calculated as ±15° (based on the correlation difference), the sound image width is calculated as 0.8 (using the formula: width = 0.85 / 0.90), and the sound image depth is calculated as 0.7 (using the formula: depth = 0.70 / 0.85). The system combines these calculated parameters into a pre-adjusted parameter set for subsequent real-time audio processing. The parameter calculation process uses recursive averaging, where each newly calculated parameter value is blended with the previous one in a ratio of 0.7:0.3 to ensure smooth parameter changes. All parameters have change rate limits set. The maximum change rate of equalization gain is 6dB / s. Dynamic parameters change using logarithmic gradients, and spatial parameters change using linear interpolation.

[0115] S210 , acquiring an environmental noise signal, converting the environmental noise signal into a noise time-frequency feature graph, and extracting frequency distribution features and energy distribution features from the noise time-frequency feature graph.

[0116] Among them, the environmental noise signal represents the background noise collected by the microphone array; the noise time-frequency characteristic diagram refers to the time-frequency analysis result of the noise signal; the frequency distribution characteristics include the energy distribution curve of the noise in the frequency domain; the energy distribution characteristics represent the change characteristics of the noise intensity over time.

[0117] The system collects ambient noise using an 8-channel circular microphone array, with a sampling rate of 48kHz and a quantization accuracy of 24 bits. The collected noise signal undergoes a short-time Fourier transform (STFT) and generates a time-frequency feature map using a 2048-point FFT and a 50% frame overlap. The system then divides the time-frequency feature map into 24 subbands, equally spaced from the 20Hz-20kHz range, and calculates the energy spectral density of each subband. By calculating the energy difference between adjacent time frames, the system captures the time-varying characteristics of the noise and establishes an energy envelope model. For frequency distribution, the system calculates the mean and variance of each subband to form a noise spectral feature vector. The system quantifies the frequency concentration and spectral shape characteristics of the noise by calculating the spectral centroid and spectral flatness. Furthermore, the system calculates the temporal correlation of the energy of each subband to identify the periodic components of the noise.

[0118] S211 : Compare the frequency distribution characteristics and energy distribution characteristics of the noise time-frequency characteristic graph with the time-frequency characteristic graph of the audio input signal to determine the noise frequency band and noise energy.

[0119] The noise frequency band refers to the frequency range where noise energy is concentrated; the noise energy represents the intensity level of noise in each frequency band; and the comparison process refers to identifying noise components through feature similarity analysis.

[0120] The system uses a two-dimensional cross-correlation analysis method to compare the time-frequency feature maps. First, the two time-frequency feature maps are resampled according to the same time-frequency resolution to ensure one-to-one correspondence between the time-frequency points. The system calculates the correlation coefficient of each time-frequency point and constructs a correlation matrix. By setting a correlation threshold of 0.7, the system identifies time-frequency regions with high correlation, which are marked as noise-dominated regions. For each frequency band, the system calculates the signal-to-noise ratio (SNR) and marks the frequency band with an SNR lower than 6dB as a noise band. The system estimates the noise energy level of each frequency band using the minimum mean square error criterion and establishes a noise power spectrum model. At the same time, the system distinguishes between steady-state noise and impulsive noise components by analyzing the fluctuation characteristics of the energy envelope.

[0121] S212: Construct a noise suppression index according to the noise frequency band and noise energy, and add the noise suppression index to the audio feature coordination index system.

[0122] The noise suppression index includes a spectrum masking index and an energy suppression index. The spectrum masking index indicates the degree of spectrum interference of noise on valid signals. The energy suppression index indicates the noise energy level that needs to be reduced.

[0123] The system constructs noise suppression indicators based on the principle of auditory masking. First, a masking threshold curve is established for 24 critical frequency bands, and the critical frequency band width is divided according to the Bark scale. The system calculates the signal-to-noise ratio and audibility of each frequency band to generate a spectrum masking indicator. When the signal energy exceeds the masking threshold by 12dB, the frequency band is marked as a signal band that needs to be protected. The system designs a suppression curve based on the noise energy level, using an attenuation slope of -12dB / octave for steady-state noise and an attenuation slope of -18dB / octave for impact noise. The system adds the constructed noise suppression indicator to the existing indicator system and combines it with other indicators through a weight coefficient of 0.3. The new indicator participates in the subsequent parameter optimization process to guide the system in noise suppression processing.

[0124] Taking a studio recording of a human voice as an example, the system constructs noise suppression metrics. First, the system detects air conditioning noise at a -45dBFS energy level in the 100-300Hz frequency band, ambient resonance at a -50dBFS energy level at 500Hz, and high-frequency equipment noise at a -60dBFS energy level above 8kHz. Based on these noise characteristics, the system constructs masking threshold curves across 24 critical frequency bands. In the 100-300Hz frequency band, due to a -35dBFS signal energy and a 10dB signal-to-noise ratio, the system sets a spectral masking index of 0.8. In the 500Hz frequency band, with a -30dBFS signal energy and a 20dB signal-to-noise ratio, the system sets a spectral masking index of 0.9. In the 8kHz frequency band, with a -40dBFS signal energy and a 20dB signal-to-noise ratio, the system sets a spectral masking index of 0.85. For the energy suppression metric, the system sets an attenuation curve of -12dB / octave in the low-frequency band, -9dB / octave in the mid-frequency band, and -15dB / octave in the high-frequency band. These metrics are added to the existing audio feature coordination metric system through a weighted average (weight coefficient of 0.3), with a weight of 0.4 for the low-frequency band, 0.4 for the mid-frequency band, and 0.2 for the high-frequency band. The system implements a dynamic update mechanism for these new metrics, updating the values ​​every 50ms and using exponential smoothing (smoothing coefficient of 0.8) to ensure continuity of changes. The noise suppression metric works together with the existing spectral balance, dynamic consistency, and spatial correlation metrics to form a more comprehensive audio quality evaluation system.

[0125] S213 : Based on the value of the noise suppression index, adjust the equalization parameter and the dynamic parameter in the target audio processing parameter combination.

[0126] Among them, the values ​​of the noise suppression index include the spectrum masking values ​​and energy suppression values ​​of 24 critical frequency bands; the equalization parameters refer to the center frequency, bandwidth and gain values ​​of the 31-band parametric equalizer; the dynamic parameters include the threshold, ratio, attack time and release time of the compressor; the adjustment process represents the calculation method for modifying the audio processing parameters according to the noise characteristics.

[0127] The system's parameter adjustment process is as follows: First, the 31-band parametric equalizer is adjusted. The system expands the noise suppression performance of 24 critical bands to 31 bands using cubic spline interpolation. For each band, the system calculates a compensation gain based on the spectral mask value, with the compensation range limited to -12dB to +6dB. When the signal-to-noise ratio (SNR) in a band falls below 6dB, the system increases the Q value in that band (from 4.32 to 8.65) to improve frequency selectivity. Regarding dynamic parameters, the system adjusts the compressor parameters based on the energy suppression value. In bands with high noise energy (SNR < 0dB), the compression ratio is increased by 50% from the original value, up to a maximum of 10:1. The compressor threshold is adjusted based on the noise energy level, calculated as: New Threshold = Original Threshold - (Noise Level + 6dB). For impulsive noise, the system shortens the compressor attack time to 0.1ms and increases the release time to 500ms. For steady-state noise, the attack time is set to 1ms and the release time to 100ms. The system smoothes the adjusted parameters through adaptive filters to avoid audio distortion caused by sudden parameter changes. Finally, the system applies the adjusted parameters to the audio processing chain to achieve precise noise suppression.

[0128] Taking an audio clip recorded in a noisy environment as an example, the system performs parameter adjustment. The system detects air conditioning noise in the low-frequency band (100-300Hz), with a spectral masking value of 0.8 and an energy suppression value of -12dB; environmental resonance in the mid-frequency band (500Hz), with a spectral masking value of 0.9 and an energy suppression value of -9dB; and equipment noise in the high-frequency band (above 8kHz), with a spectral masking value of 0.85 and an energy suppression value of -15dB. In response to these noise characteristics, the system first adjusted the settings of the 31-band parametric equalizer: in the 100-300Hz frequency band, the Q value was increased from 4.32 to 8.65 to improve frequency selectivity, and the gain values ​​of the five equalizers in this frequency band were calculated through cubic spline interpolation, ranging from -8dB to -4dB; in the 500Hz frequency band, a narrowband notch filter was set with a Q value of 12.5 and a gain of -6dB; in the frequency band above 8kHz, a high-shelf filter was used with a slope set to -12dB / octave. For dynamic parameters, the system increases the compression ratio from 2:1 to 3:1 in the low-frequency noise range, lowers the threshold to -24dB (from -6dB of the original signal level), shortens the attack time to 0.1ms to quickly respond to sudden noise fluctuations, and extends the release time to 500ms to smooth the effect. In the mid-frequency resonance range, the compression ratio is set to 4:1, the threshold is set to -20dB, the attack time is 1ms, and the release time is 200ms. In the high-frequency noise range, the expander is applied with a ratio of 1:2, the threshold is set to -45dB, the attack time is 5ms, and the release time is 100ms. The system uses recursive averaging for all parameter adjustments, blending the new parameter value with the original parameter value in a ratio of 0.6:0.4. Parameter change rate is also limited: the maximum change rate of equalization gain is 8dB / s, the compression parameter changes are logarithmic, and the time constant adjustment uses linear interpolation to ensure smooth and continuous processing.

[0129] The following describes the digital audio signal processing system for a mixing console in an embodiment of the present invention from the perspective of hardware processing. Figure 3 , is a schematic diagram of the physical device structure of the mixing console digital audio signal processing system in an embodiment of the present application.

[0130] It should be noted that Figure 3 The structure of the digital audio signal processing system for a mixing console shown is only an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0131] like Figure 3As shown, the mixing console digital audio signal processing system includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 302 or programs loaded from a storage unit 308 into a random access memory (RAM) 303, such as the methods described in the above embodiments. RAM 303 also stores various programs and data required for system operation. CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to bus 304.

[0132] The following components are connected to the I / O interface 305: an input section 306 including an audio input device, push button switches, and the like; an output section 307 including a liquid crystal display (LCD), an audio output device, indicator lights, and the like; a storage section 308 including a hard disk and the like; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. Removable media 311, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 310 as needed, so that computer programs read from the removable media can be installed in the storage section 308 as needed.

[0133] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for executing the methods illustrated in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication section 309 and / or installed from removable media 311. When executed by the central processing unit (CPU) 301, the computer program performs the various functions defined in the present invention.

[0134] It should be noted that specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0135] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. Each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings.

[0136] Specifically, the mixing console digital audio signal processing system of this embodiment includes a processor and a memory. The memory stores a computer program. When the computer program is executed by the processor, the mixing console digital audio signal processing method provided in the above embodiment is implemented.

[0137] As another aspect, the present invention further provides a computer-readable storage medium, which may be included in the mixing console digital audio signal processing system described in the above embodiments, or may exist independently and not be incorporated into the mixing console digital audio signal processing system. The storage medium carries one or more computer programs, which, when executed by a processor of the mixing console digital audio signal processing system, enable the mixing console digital audio signal processing system to implement the mixing console digital audio signal processing method provided in the above embodiments.

[0138] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

[0139] As used in the above embodiments, the term “when” may be interpreted to mean “if” or “after” or “in response to determining that” or “in response to detecting that”, depending on the context. Similarly, the phrases “upon determining that” or “if (stated condition or event) is detected” may be interpreted to mean “if determining that” or “in response to determining that” or “upon detecting (stated condition or event)” or “in response to detecting (stated condition or event)”, depending on the context.

[0140] Those skilled in the art will appreciate that all or part of the process steps in the above-described method embodiments can be implemented by a computer program instructing the relevant hardware. The program can be stored in a computer-readable storage medium, and when executed, the program can include the process steps in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A method for processing digital audio signals of a mixing console, characterized in that: Applied to a digital audio signal processing system of a mixing console, the method comprises: Acquire multiple audio input signals, and perform time-frequency domain conversion on the audio input signals to obtain a time-frequency feature map; Analyzing the frequency distribution of each of the audio input signals based on the time-frequency feature graph, extracting the fundamental frequency and overtone structure features of each of the audio input signals, and establishing a timbre feature fingerprint library based on the fundamental frequency and overtone structure features; Acquiring a preset standard audio signal, extracting a multidimensional audio feature vector from the preset standard audio signal, and matching the multidimensional audio feature vector with the timbre feature fingerprint library to obtain an audio feature parameter group; Calculating energy values ​​of different frequency bands in the audio feature parameter group, obtaining energy ratios between adjacent frequency bands, and calculating an overall sound coordination index based on the energy ratios; Extracting the voice part loudness value and frequency value in the audio feature parameter group, and calculating the voice part level clarity index corresponding to the overall sound coordination index; Calculating the phase difference and amplitude difference of the channel signals based on the left and right channel signal values ​​in the audio feature parameter group, determining the sound image position distribution in combination with the voice level clarity index, and obtaining a spatial sound image balance index, wherein the audio feature coordination index system includes overall sound coordination, voice level clarity, and spatial sound image balance; Calculating the audio feature coordination index of the multiple audio input signals in real time based on the audio feature coordination index system, dividing the multiple audio input signals into segments according to time intervals, and calculating the coordination index of each segment according to the overall sound coordination calculation rules, the part layer clarity calculation rules, and the spatial sound image balance calculation rules in the audio feature coordination index system; Extracting a coordination index of a corresponding time period from the preset standard audio signal, and performing a numerical comparison with the coordination indexes of the multiple audio input signals to obtain a numerical comparison result; Determining a signal adjustment value based on the value comparison result and a matching feature of the timbre feature fingerprint library to obtain a signal adjustment value; Based on the signal adjustment value, a target audio processing parameter combination is calculated in combination with the timbre feature fingerprint library, where the target audio processing parameter combination includes an equalization parameter, a dynamic parameter, and a space parameter.

2. The method according to claim 1, characterized in that After the step of adjusting the value based on the signal and calculating the target audio processing parameter combination in combination with the timbre feature fingerprint library, the method further includes: Extracting characteristic change patterns of the preset standard audio signal and establishing a characteristic change record table; Analyzing the adjustment direction of the target audio processing parameter combination according to the feature change record table to generate a parameter adjustment step; The target audio processing parameter combination is adjusted according to the parameter adjustment step size.

3. The method according to claim 1, characterized in that After the step of obtaining multiple audio input signals and performing time-frequency domain conversion on the audio input signals to obtain a time-frequency feature map, the method further includes: Analyzing the timbre feature distribution in the time-frequency feature graph, and identifying signal features according to the timbre feature distribution; the signal features include musical instrument sounds or human voices; Separating the multiple audio input signals into independent audio tracks according to the signal characteristics; Extracting fundamental frequency and overtone structure features of the independent audio track and performing feature matching with the timbre feature fingerprint library; The audio feature coordination index of the independent audio tracks is calculated to obtain a target audio processing parameter combination for each audio track.

4. The method according to claim 1, wherein After the step of calculating a target audio processing parameter combination based on the signal adjustment value and in combination with the matching features of the timbre feature fingerprint library, the method further includes: Extracting frequency change features and energy change features from the time-frequency feature graph, and determining change patterns of the frequency change features and the energy change features; Predicting a fundamental frequency change trend and a harmonic structure change trend in a next time period based on the change rule and historical data of the audio feature parameter group; Adjusting the values ​​of various indicators in the audio feature coordination index system based on the fundamental frequency change trend and the overtone structure change trend; Substituting the adjusted values ​​of the various indicators into the calculation process of the target audio processing parameter combination, a pre-adjustment parameter group including equalization parameters, dynamic parameters and space parameters is obtained.

5. The method according to claim 4, characterized in that After the step of adjusting the value based on the signal and calculating the target audio processing parameter combination in combination with the timbre feature fingerprint library, the method further includes: Acquire an ambient noise signal, convert the ambient noise signal into a noise time-frequency feature graph, and extract frequency distribution features and energy distribution features from the noise time-frequency feature graph; Comparing the frequency distribution characteristics and energy distribution characteristics of the noise time-frequency characteristic graph with the time-frequency characteristic graph of the audio input signal to determine the noise frequency band and noise energy; constructing a noise suppression index according to the noise frequency band and noise energy, and adding the noise suppression index to the audio feature coordination index system; Based on the value of the noise suppression index, the equalization parameter and the dynamic parameter in the target audio processing parameter combination are adjusted.

6. A digital audio signal processing system for a mixing console, characterized in that: The mixing console digital audio signal processing system includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the mixing console digital audio signal processing system to execute the method described in any one of claims 1-5.

7. A computer-readable storage medium comprising instructions, characterized in that: When the instruction is executed on a mixing console digital audio signal processing system, the mixing console digital audio signal processing system is caused to execute the method according to any one of claims 1 to 5.

8. A computer program product, characterized in that When the computer program product is run on a mixing console digital audio signal processing system, the mixing console digital audio signal processing system is enabled to perform the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Audio processing method and device, electronic equipment and storage medium

    CN114422897A

  • Balanced adjustment method, system and device for sound console volume controller and medium

    CN119603605A