Sound processing method, system, medium, and device based on audio spectral features

CN122551817APending Publication Date: 2026-08-11GUANGZHOU HAVIT COMP TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-22
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0007]本发明的目的是提供一种基于音频频谱特征的声音处理方法、系统、介质及设备,用于解决现有技术中混合内容主导成分识别粗糙、音效切换生硬以及参数调整缺乏协同性的问题

Benefits of technology

1.本发明通过提取子带频谱质心、频谱通量与互通道电平差三类频域特征,从频谱重心、变化烈度和空间扩散度三个独立维度精细刻画音频帧的声学形态属性;利用两级联分类结构由粗到精进行高效率、高精度的成分判别,有效解决了在混合音频中区分语音、瞬态音效与持续乐音这一关键技术难点。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551817A_ABST
    Figure CN122551817A_ABST
Patent Text Reader

Abstract

This invention discloses a sound processing method, system, medium, and device based on audio spectral features. The method includes: acquiring multi-channel audio signals to be processed; performing time-frequency transformation on the signal of each channel to extract multi-dimensional frequency domain feature parameters; inputting the multi-dimensional frequency domain feature parameters into a cascaded classifier model to obtain membership scores for the current audio frame belonging to the speech-dominated type, transient sound effect-dominated type, and continuous musical sound-dominated type; triggering parameter state transitions in the sound effect processing pipeline based on the smoothed result of the membership scores on the time axis, and coordinating the adjustment of the virtual sound field expansion width, the gain curve of the multi-segment dynamic equalizer, and the release time of the peak limiter. This invention achieves fine-grained discrimination of dominant sound types in complex mixed content and smooth adaptive switching of sound effect strategies by performing multi-dimensional spectral morphology analysis on audio signals, solving the problems of coarse scene judgment granularity and abrupt parameter switching in existing technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio signal processing technology, and more specifically, to a sound processing method, system, medium, and device based on audio spectral characteristics. Background Technology

[0002] In multimedia content consumption scenarios, audio signals are typically composed of a mixture of various acoustic components with distinct properties. For example, movie audio simultaneously includes dialogue, ambient sound effects, and background music; in video games, voice prompts, skill sound effects, and background music are mixed and output in real time. Auditory perception research indicates that the human ear has differentiated expectations for the "ideal listening experience" of different dominant sound types: speech requires clarity and intelligibility as a priority, transient sound effects need impact and spatial positioning, and sustained musical sounds seek natural balance and dynamic integrity.

[0003] However, existing audio processing equipment typically uses fixed processing parameters or simple energy threshold judgments, making it difficult to accurately identify and adapt to the dominant components that change in real time in mixed audio. The existing technologies have the following technical solutions and drawbacks: 1. Manual sound effect mode switching solution: Users select preset sound effects such as "movie mode" or "music mode" in the player or device menu; the disadvantage is that the operation interrupts the continuity of the content experience, and ordinary users have difficulty in judging and switching to the appropriate mode in time, resulting in the sound effect parameters being disconnected from the content most of the time.

[0004] 2. An automatic scene detection scheme based on energy statistics simply divides "dynamic" and "static" scenes by calculating macroscopic statistics such as the overall RMS energy and the proportion of low-frequency energy of the audio signal. Its disadvantage is that the energy dimension is too singular, and it cannot distinguish between equally high-energy sounds (transient single-point sound sources) and symphonic ensembles (continuous wide-area sound sources), nor can it distinguish between equally low-to-medium energy whispered dialogue and suspenseful ambient sounds.

[0005] 3. A two-stage processing scheme based on sound source separation uses a deep learning network to separate the input signal into independent tracks such as "human voice" and "background sound" in real time and then process them separately. Its disadvantage is that the amount of computation is extremely large, making it difficult to meet the real-time requirements on mobile or embedded devices. In addition, the spectral damage and artificial artifacts introduced during the separation process will reduce the output sound quality.

[0006] 4. Existing adaptive solutions generally use hard-decision output type labels, lacking a smooth transition mechanism between states, which causes abrupt switching of sound effect parameters at type boundaries, resulting in a perceptible auditory disconnect and destroying the immersive experience. Summary of the Invention

[0007] The purpose of this invention is to provide a sound processing method, system, medium, and device based on audio spectrum characteristics to solve the problems of rough identification of the dominant component of mixed content, abrupt sound effect switching, and lack of coordination in parameter adjustment in the prior art.

[0008] The first aspect of this invention provides a sound processing method based on audio spectral features, comprising the following steps: Acquire multi-channel audio signals, divide the spectrum of each channel audio frame into sub-bands, and extract the spectral centroid and spectral flux of each sub-band; The cross-channel level difference between corresponding sub-bands between channels is calculated, and a multi-dimensional frequency domain feature vector is constructed based on the spectral centroid, spectral flux, and cross-channel level difference. The multi-dimensional frequency domain feature vector is input into the cascaded decision model. The first-level coarse classifier performs flow division based on the high-frequency subband spectral flux and the cross-channel level difference. The second-level fine classifier outputs the soft decision membership scores of the current frame for the speech-dominated class, transient sound effect-dominated class, and continuous musical sound-dominated class, respectively. The membership scores are smoothed by time-axis filtering. When the smoothed membership score of any category exceeds the preset transition threshold and continues to reach the preset confirmation time, the state transition is triggered, and the current global sound state is switched. Based on the switched processing state, the virtual sound field width parameter, the multi-segment dynamic equalizer gain curve parameter, and the peak limiter release time parameter in the sound effect processing pipeline are adjusted synchronously to complete adaptive sound effect processing.

[0009] In this solution, the acquisition of multi-channel audio signals, the division of the spectrum of each channel audio frame into sub-bands, and the extraction of the spectral centroid and spectral flux of each sub-band specifically include: The spectrum of the audio frame is divided into low-frequency sub-band, mid-low-frequency sub-band, mid-high-frequency sub-band, and high-frequency sub-band on the frequency axis according to the octave band method. The frequency range of the low-frequency sub-band is 20Hz to 80Hz, the frequency range of the mid-low-frequency sub-band is 80Hz to 500Hz, the frequency range of the mid-high-frequency sub-band is 500Hz to 4000Hz, and the frequency range of the high-frequency sub-band is 4000Hz to 20000Hz. Iterate through all frequency points within each sub-band and calculate the spectral centroid of the sub-band according to a preset formula; Calculate the amplitude difference between the current frame's spectral amplitude and the previous frame's spectral amplitude at each corresponding frequency point, and use the sum of the squares of the amplitude differences at all frequency points as the spectral flux value of the current sub-band. The spectral flux value is used to characterize the degree of drastic change in the spectral structure within the sub-band in the time domain.

[0010] In this scheme, the calculation of the cross-channel level difference between corresponding sub-bands between audio channels is based on a multi-dimensional frequency domain feature vector constructed from the spectral centroid, spectral flux, and cross-channel level difference, specifically including: In a multi-channel audio signal, the left and right channels, which are spatially symmetrical in the channel configuration, are selected to form at least one channel pair. Based on the channel pair, calculate the total energy of the left channel sub-band and the total energy of the right channel sub-band for each sub-band of the left channel; The cross-channel level difference at the corresponding decibel scale of the current sub-band is calculated based on the total energy of the left channel sub-band and the total energy of the right channel sub-band, wherein the cross-channel level difference is used to reflect the spatial bias of the current sound source on the corresponding sub-band; The spectral centroid values, spectral flux values, and inter-channel level differences of the four sub-bands are concatenated in a preset order to form a twelve-dimensional frequency domain feature vector, which serves as the multi-dimensional frequency domain feature vector of the current audio frame.

[0011] In this scheme, the training process of the cascaded decision model includes the following steps: Construct a training sample set, which includes multiple labeled pure speech segments, pure transient sound effect segments, and pure continuous musical sound segments. The annotation information of each segment includes the category label of the dominant sound type of the corresponding segment. For each audio segment in the training sample set, the multidimensional frequency domain feature vector is extracted frame by frame, and each frame is assigned a category label corresponding to its segment. Using the high-frequency subband spectral flux of each frame and the cross-channel level difference of the four subbands as input features, and the concentrated sound image class label or the diffuse sound image class label as output target, a binary classification decision tree model of the first-level coarse classifier is trained. Frames labeled as pure speech segments and pure transient sound effect segments are labeled as concentrated sound image class, and frames labeled as pure continuous musical sound segments are labeled as diffuse sound image class. The mid-to-low frequency subband spectral centroid, mid-to-low frequency subband spectral flux, and high frequency subband spectral flux of all training frames labeled as concentrated audio-visual classes are used as input features, and the speech-dominant class label and transient sound effect-dominant class label are used as output targets to train the speech / transient sound effect discrimination support vector machine branch model in the second-level fine classifier. The mid-to-high frequency subband spectral centroid and mid-to-high frequency subband spectral flux of all training frames labeled as diffuse sound image class are used as input features. The continuous musical tone dominant class label is used as the positive output target. The continuous musical tone discrimination support vector machine branch model in the second-level fine classifier is trained, and a sigmoid function mapping layer is added to the output of the branch model to map the decision value of the support vector machine to the membership score in the interval of 0 to 1.

[0012] In this solution, the membership score is smoothed over time. When the smoothed membership score of any category exceeds a preset transition threshold and continues for a preset confirmation duration, a state transition is triggered to switch the current global sound state. Specifically, this includes: Maintain a smooth membership score variable for each of the speech-dominated, transient sound effect-dominated, and continuous musical sound-dominated classes, with the initial value of each smooth membership score variable being zero. For each new input audio frame, after obtaining the three types of soft-decision membership scores of the current frame, a first-order infinite impulse response low-pass filter is used to recursively smooth the three types of membership scores respectively. The frame count limit is calculated based on the preset confirmation duration. When the smooth membership score exceeds the preset migration threshold, the corresponding frame counter is started, and the number of consecutive frames reaches the frame limit. Then, it is confirmed that the state migration condition is met, and the current global sound state variable is updated to the processing state identifier.

[0013] In this solution, based on the switched processing state, the virtual sound field width parameter, the multi-segment dynamic equalizer gain curve parameter, and the peak limiter release time parameter in the sound effect processing pipeline are synchronously adjusted to complete adaptive sound effect processing. Specifically, this includes: The audio processing pipeline includes a virtual sound field synthesis module, a multi-band dynamic equalizer module, and a peak limiter module connected in series or in parallel. Maintain a sound effect parameter matrix, which stores preset values ​​of virtual sound field width parameters corresponding to voice optimization state, transient enhancement state and music reproduction state, preset values ​​of gain curve parameters of each frequency band of multi-band dynamic equalizer, and preset values ​​of peak limiter release time parameters. When a switch in the current global sound state variable is detected, all preset values ​​of sound effect parameters corresponding to the post-switching processing state are read from the sound effect parameter matrix; Within the preset transition time window, the sound field width control parameters of the virtual sound field synthesis module, the gain curve control parameters of each frequency band of the multi-band dynamic equalizer module, and the release time control parameters of the peak limiter are smoothly transitioned from their current values ​​to the newly read preset values, thus completing the synchronous update of the sound effect processing parameters.

[0014] A second aspect of the present invention also provides a sound processing system based on audio spectrum features, including a memory and a processor. The memory includes a sound processing method program based on audio spectrum features, which, when executed by the processor, performs the following steps: The input multi-channel audio signal is acquired, and the spectrum of the audio frame of each channel is divided into multiple sub-bands. The spectral centroid and spectral flux of each sub-band are extracted. Calculate the cross-channel level difference between each channel on the corresponding sub-band, and construct a multi-dimensional frequency domain feature vector based on the spectral centroid, the spectral flux, and the cross-channel level difference; The multi-dimensional frequency domain feature vector is input into the trained cascaded decision model. The first-level coarse classifier calculates the spatial diffusion score and the high-frequency transient attitude score based on the spectral flux of the high-frequency sub-band and the cross-channel level difference. The second-level fine classifier calculates the soft decision membership score of the current frame for the speech-dominated class, transient sound effect-dominated class and continuous musical sound-dominated class respectively under the corresponding path. The membership scores are smoothed by time-axis filtering. When the smoothed membership score of any category exceeds the preset migration threshold and the overshoot state continues for a preset confirmation time, it is determined that a sound state migration has occurred. At this time, the current global sound state is switched to the processing state corresponding to the current category. Based on the switched processing state, the virtual sound field width parameter, the gain curve parameters of each frequency band of the multi-band dynamic equalizer, and the release time parameter of the peak limiter in the sound effect processing pipeline are adjusted synchronously to complete adaptive sound effect processing that matches the current sound state.

[0015] A third aspect of the present invention provides a computer-readable storage medium comprising a machine program for a sound processing method based on audio spectrum features, wherein when executed by a processor, the sound processing method program based on audio spectrum features implements the steps of the sound processing method based on audio spectrum features as described in any of the preceding claims.

[0016] A fourth aspect of the present invention provides a computer program product comprising computer program code, wherein when the computer program code is run on a computer, the computer implements the steps of a sound processing method based on audio spectral features as described in any of the preceding claims.

[0017] A fifth aspect of the present invention provides an electronic device comprising: a processor and a memory; wherein the memory is configured to store a computer program, and the processor is configured to execute the computer program stored in the memory to cause the electronic device to perform the steps of a sound processing method based on audio spectral features as described in any of the preceding claims.

[0018] The present invention discloses a sound processing method, system, medium, and device based on audio spectrum characteristics, which has the following beneficial effects: 1. This invention extracts three types of frequency domain features—subband spectral centroid, spectral flux, and cross-channel level difference—to finely characterize the acoustic morphological properties of audio frames from three independent dimensions: spectral centroid, intensity of change, and spatial diffusion. It utilizes a two-tiered classification structure to perform high-efficiency and high-precision component discrimination from coarse to fine, effectively solving the key technical challenge of distinguishing speech, transient sound effects, and continuous musical tones in mixed audio.

[0019] 2. This invention replaces hard classification labels with soft decision membership scores, and with a smooth state transition mechanism with hysteresis time axis, the state switch is only triggered after a certain type of dominant feature has been stable for a period of time. This eliminates the abrupt auditory discontinuity caused by the "jump" of sound effect parameters in traditional solutions, and achieves a seamless and smooth transition between different dominant sound segments.

[0020] 3. This invention provides a complete three-dimensional collaborative parameter matrix of "virtual sound field width - multi-segment dynamic equalization - peak limiter release time", which provides customized optimal parameter combinations for the three dominant types of speech, transient and musical sounds, so that each component can achieve its own ideal listening experience in terms of spatial sense, timbre balance and dynamic elasticity.

[0021] 4. This invention is based on lightweight signal processing and shallow machine learning models. It does not rely on deep neural network sound source separation, has low computational load, and can run in real time on mainstream Bluetooth audio SoCs or mobile DSPs, making it highly feasible for industrial applications. Attached Figure Description

[0022] Figure 1 The diagram illustrates the steps of a sound processing method based on audio spectral features according to the present invention. Figure 2 A schematic flowchart of a sound processing method based on audio spectral features according to the present invention is shown; Figure 3 This diagram illustrates the switching of sound processing states in a sound processing method based on audio spectrum features according to the present invention. Figure 4 A block diagram of a sound processing system based on audio spectral features according to the present invention is shown. Detailed Implementation

[0023] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments of the present invention and the features thereof can be combined with each other.

[0024] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.

[0025] Specifically, Figure 1 The diagram illustrates the steps of a sound processing method based on audio spectral features according to the present invention.

[0026] like Figure 1 As shown, this invention discloses a sound processing method based on audio spectral features, comprising the following steps: S102, acquire the input multi-channel audio signal, divide the spectrum of the audio frame of each channel into multiple sub-bands, and extract the spectral centroid and spectral flux of each sub-band; S104, calculate the cross-channel level difference between each channel on the corresponding sub-band, and construct a multi-dimensional frequency domain feature vector based on the spectral centroid, the spectral flux and the cross-channel level difference; S106, input the multi-dimensional frequency domain feature vector into the trained cascaded decision model to obtain the spatial diffusion score, high-frequency instantaneous attitude score, and membership score. S108, perform time-axis smoothing filtering on the membership score. When the smoothed membership score of any category exceeds the preset migration threshold and the overshoot state continues for a preset confirmation time, determine that a sound state migration has occurred and switch the sound processing state. S110, based on the switched processing state, synchronously adjust the virtual sound field width parameter, the gain curve parameters of each frequency band of the multi-band dynamic equalizer, and the release time parameter of the peak limiter in the sound effect processing pipeline to complete adaptive sound effect processing that matches the current sound state.

[0027] It should be noted that, in this embodiment, a sound processing method based on audio spectrum features is provided, which can run on electronic devices with audio processing capabilities, such as smartphones, tablets, personal computers, Bluetooth audio SoCs (System on Chips) or embedded DSPs (Digital Signal Processors). Through refined frequency domain feature extraction and a coarse-to-fine cascaded classification strategy, it achieves accurate discrimination of the three dominant types of mixed audio: speech, transient sound effects, and continuous musical sounds. Furthermore, by combining soft-decision membership scores with a smooth state transition mechanism with hysteresis characteristics, and a set of three-dimensional collaborative sound effect parameter matrices, it solves the problems of coarse scene judgment granularity, abrupt switching of processing parameters, and lack of synergy in the prior art.

[0028] Specifically, in this embodiment, as Figure 2As shown in the flowchart, the process begins with the acquisition of multi-channel audio signals. After framing and time-frequency transformation, the spectrum is divided into four sub-bands by octave bands, and the spectral centroid and spectral flux of each sub-band are extracted. Next, the inter-channel level difference between corresponding sub-bands is calculated, and the three types of features are concatenated into a twelve-dimensional frequency domain feature vector. Then, the feature vector is input into a two-stage cascaded classifier—first, a coarse classifier sorts the signals based on spatial and transient information, and then a dedicated fine classifier outputs three soft-decision membership scores for the current frame: speech, transient sound effects, or continuous musical sounds. Next, the scores are recursively smoothed and filtered. When a certain smoothing score consistently exceeds a threshold and remains for a preset duration, a global sound state transition is triggered. Finally, the virtual sound field width, multi-segment dynamic equalizer gain curve, and peak limiter release time are synchronously adjusted according to the new state to achieve adaptive sound effect processing with parameter coordination.

[0029] According to an embodiment of the present invention, the step of acquiring the input multi-channel audio signal, dividing the spectrum of the audio frame of each channel into multiple sub-bands, and extracting the spectral centroid and spectral flux of each sub-band specifically includes: The spectrum of the audio frame is divided into low-frequency sub-band, mid-low-frequency sub-band, mid-high-frequency sub-band, and high-frequency sub-band on the frequency axis according to the octave band method. The frequency range of the low-frequency sub-band is 20Hz to 80Hz, the frequency range of the mid-low-frequency sub-band is 80Hz to 500Hz, the frequency range of the mid-high-frequency sub-band is 500Hz to 4000Hz, and the frequency range of the high-frequency sub-band is 4000Hz to 20000Hz. Iterate through all frequency points within each sub-band and calculate the spectral centroid of the sub-band according to a preset formula, as follows: ; in, For the spectral centroid, For the child belt The center frequency of each frequency point For the first spectral amplitude values ​​at each frequency point; Calculate the amplitude difference between the current frame's spectral amplitude and the previous frame's spectral amplitude at each corresponding frequency point, and use the sum of the squares of the amplitude differences at all frequency points as the spectral flux value of the current sub-band. The spectral flux value is used to characterize the degree of drastic change in the spectral structure within the sub-band in the time domain.

[0030] It should be noted that, in this embodiment, firstly, a real-time multi-channel audio signal, such as stereo (left and right dual channels) or surround sound signal, is obtained from an audio source. The continuous audio stream is then processed by frame division with a fixed frame length and frame shift. For example, each frame is 20 milliseconds long and the frame shift is 10 milliseconds. For each frame of audio signal, a time-frequency transformation, such as a short-time fourier transform (STFT), is performed on each channel to obtain the spectrum of that channel.

[0031] Furthermore, in this embodiment, to finely characterize the acoustic characteristics of different frequency bands, the present invention divides the spectrum of the audio frame into four sub-bands on the frequency axis using an octave division method: Low-frequency sub-band: frequency range of 20Hz to 80Hz, mainly capturing ultra-low bass energy; Mid-low frequency sub-band: frequency range of 80Hz to 500Hz, corresponding to the fundamental frequency and main harmonic energy region of speech, as well as the basic range of most musical instruments; Mid-high frequency sub-band: frequency range of 500Hz to 4000Hz, corresponding to the important region of consonant clarity in speech and overtones of musical instruments; High-frequency sub-band: frequency range of 4000Hz to 20000Hz, mainly reflecting the airiness, detail, and environmental atmosphere of transient sound effects. For each sub-band of each channel, the following two types of features are extracted: the spectral centroid, used to characterize the spectral energy distribution within the sub-band. The calculation method for SC is as follows: traverse all frequency points within the sub-band, and calculate the weighted average of the frequencies using the square of the spectral amplitude of each frequency point as the weight. The higher the spectral centroid, the more concentrated the energy of the sub-band is in the high-frequency part, and vice versa. Spectral Flux (SF), used to characterize the degree of drastic change in the spectral structure within the sub-band in the time domain, is calculated by calculating the amplitude difference between the current frame's spectral amplitude and the previous frame's spectral amplitude at each corresponding frequency point, and using the sum of the squares of the amplitude differences at all frequency points as the spectral flux value of the current sub-band. The higher the value, the more drastic the spectral change, which often occurs at the beginning of transient sound effects.

[0032] According to an embodiment of the present invention, the calculation of the cross-channel level difference between each channel on the corresponding sub-band, based on the spectral centroid, the spectral flux, and the cross-channel level difference to construct a multi-dimensional frequency domain feature vector, specifically includes: In a multi-channel audio signal, the left and right channels, which are spatially symmetrical in the channel configuration, are selected to form at least one channel pair. Based on the channel pair, calculate the total energy of the left channel sub-band and the total energy of the right channel sub-band for each sub-band of the left channel; The cross-channel level difference at the corresponding decibel scale of the current sub-band is calculated based on the total energy of the left channel sub-band and the total energy of the right channel sub-band, wherein the cross-channel level difference is used to reflect the spatial bias of the current sound source on the corresponding sub-band; The spectral centroid values, spectral flux values, and inter-channel level differences of the four sub-bands are concatenated in a preset order to form a twelve-dimensional frequency domain feature vector, which serves as the multi-dimensional frequency domain feature vector of the current audio frame.

[0033] It should be noted that, in this embodiment, in order to capture the spatial attributes of the sound signal, this embodiment introduces the Inter-channel Level Difference (ILD) feature. The specific process is as follows: (1) Select a channel pair. In a multi-channel audio signal, select the left and right channels that are spatially symmetrical in the channel configuration to form a channel pair for analysis. For a multi-channel system, multiple channel pairs can be selected; (2) Calculate the total energy of the sub-band. Calculate the total energy of the left and right channels in each sub-band, that is, sum the square values ​​of the spectral amplitudes of all frequency points in the sub-band to obtain the total energy of the left channel sub-band and the total energy of the right channel sub-band; (3) Calculate the inter-channel level difference: Based on the total energy of the left and right channel sub-bands, calculate the inter-channel level difference at the corresponding decibel scale of the current sub-band. The formula is as follows: ; in, The difference in channel levels. The total energy of the left vocal tract. This represents the total energy of the right channel sub-band. Specifically, in this embodiment, the magnitude and sign of this value reflect the spatial bias of the current sound source in the corresponding sub-band. For example, A value close to 0dB indicates that the sound source is centered, while a larger positive value indicates that the sound source is more biased to the left.

[0034] Further, in this embodiment, (4) splicing feature vectors, specifically splicing the spectral centroid values, spectral flux values, and interchannel level differences of the four sub-bands in a preset order (e.g., sub-band 1 centroid, ..., sub-band 4 centroid, sub-band 1 flux, ..., sub-band 4 flux, sub-band 1 ILD, ..., sub-band 4 ILD) to form a twelve-dimensional frequency domain feature vector, which serves as the multi-dimensional frequency domain feature vector of the current audio frame. This vector integrates information from three dimensions: spectral centroid, intensity of change, and spatial diffusion.

[0035] According to an embodiment of the present invention, the training process of the cascaded decision model includes the following steps, specifically: Construct a training sample set, which includes multiple labeled pure speech segments, pure transient sound effect segments, and pure continuous musical sound segments. The annotation information of each segment includes the category label of the dominant sound type of the corresponding segment. For each audio segment in the training sample set, the multidimensional frequency domain feature vector is extracted frame by frame, and each frame is assigned a category label corresponding to its segment. Using the high-frequency subband spectral flux of each frame and the cross-channel level difference of the four subbands as input features, and the concentrated sound image class label or the diffuse sound image class label as output target, a binary classification decision tree model of the first-level coarse classifier is trained. Frames labeled as pure speech segments and pure transient sound effect segments are labeled as concentrated sound image class, and frames labeled as pure continuous musical sound segments are labeled as diffuse sound image class. The mid-to-low frequency subband spectral centroid, mid-to-low frequency subband spectral flux, and high frequency subband spectral flux of all training frames labeled as concentrated audio-visual classes are used as input features, and the speech-dominant class label and transient sound effect-dominant class label are used as output targets to train the speech / transient sound effect discrimination support vector machine branch model in the second-level fine classifier. The mid-to-high frequency subband spectral centroid and mid-to-high frequency subband spectral flux of all training frames labeled as diffuse sound image class are used as input features. The continuous musical tone dominant class label is used as the positive output target. The continuous musical tone discrimination support vector machine branch model in the second-level fine classifier is trained, and a sigmoid function mapping layer is added to the output of the branch model to map the decision value of the support vector machine to the membership score in the interval of 0 to 1.

[0036] It should be noted that, in this embodiment, the training process of the cascaded decision model mentioned in this embodiment is also within the protection scope of this invention. Its specific process has been clearly defined in the above embodiments, and will not be repeated here.

[0037] According to an embodiment of the present invention, the step of performing time-axis smoothing filtering on the membership scores, when the smoothed membership score of any category exceeds a preset migration threshold and the overshoot state continues for a preset confirmation duration, determines that a sound state migration has occurred. At this time, the current global sound state is switched to the processing state corresponding to the current category, specifically including: Maintain a smooth membership score variable for each of the speech-dominated, transient sound effect-dominated, and continuous musical sound-dominated classes, with the initial value of each smooth membership score variable being zero. For each new input audio frame, after obtaining the three soft-decision membership scores of the current frame, a first-order infinite impulse response low-pass filter is used to recursively smooth the three membership scores respectively. The recursive formula is as follows: ; in, For category indexing, The current frame number. For category The original membership score in the current frame. For category The smooth membership score in the current frame. For category The smooth membership score in the (n-1)th frame, The preset smoothing coefficient; The frame count limit is calculated based on a preset confirmation duration, using the following formula: ,in, For frame rate limits, To preset the confirmation time, For audio frame rate; When the smooth membership score exceeds the preset migration threshold, the corresponding frame counter is started, and the number of consecutive frames reaches the frame limit. Then, it is confirmed that the state migration condition is met, and the current global sound state variable is updated to the processing state identifier.

[0038] It should be noted that, in this embodiment, as Figure 3 As shown, this is a schematic diagram of switching sound processing states. To avoid frequent switching of sound effect parameters due to slight jitter in the frame-by-frame discrimination results, this invention introduces a smooth state transition mechanism with hysteresis characteristics. Here, a smooth membership score variable is maintained for the speech-dominant class (i=1), transient sound effect-dominant class (i=2), and continuous musical tone-dominant class (i=3), and their initial values ​​are all zero.

[0039] Specifically, in this embodiment, recursive smoothing is performed. For each new input audio frame, after obtaining the three original membership scores of the output, a first-order infinite impulse response low-pass filter is used for recursive smoothing to eliminate high-frequency jitter. The recursive formula is as follows: ,in, For category indexing, The current frame number. For category The original membership score in the current frame. For category The smooth membership score in the current frame. For category The smooth membership score in the (n-1)th frame, The preset smoothing coefficient (e.g., 0.9) is used. The closer the value is to 1, the stronger the smoothing effect, the more stable the state, but also the less responsive it is.

[0040] Furthermore, in this embodiment, the dual-condition state transition judgment specifically involves setting a preset transition threshold (e.g., 0.7). When the smooth membership score of a certain category exceeds this threshold for the first time, a frame counter dedicated to that category is activated. To avoid instantaneous noise interference, a preset confirmation duration (e.g., 300 milliseconds) is set. The required consecutive frame count limit is calculated based on the audio frame rate (e.g., 100 frames / second). Only when the smooth membership score remains above the threshold and the consecutive frame count reaches the frame count limit is the "sound state transition" event finally determined to have occurred. At this time, the variable representing the current global sound state is updated to the processing state identifier corresponding to category i. This mechanism ensures that the switching of sound effect strategies follows a substantial and continuous change in the content-driven type, rather than a brief fluctuation, which is key to achieving a seamless listening experience.

[0041] According to an embodiment of the present invention, based on the switched processing state, the virtual sound field width parameter, the gain curve parameters of each frequency band of the multi-band dynamic equalizer, and the release time parameter of the peak limiter in the sound effect processing pipeline are synchronously adjusted to complete adaptive sound effect processing that matches the current sound state, specifically including: The audio processing pipeline includes a virtual sound field synthesis module, a multi-band dynamic equalizer module, and a peak limiter module connected in series or in parallel. Maintain a sound effect parameter matrix, which stores preset values ​​of virtual sound field width parameters corresponding to voice optimization state, transient enhancement state and music reproduction state, preset values ​​of gain curve parameters of each frequency band of multi-band dynamic equalizer, and preset values ​​of peak limiter release time parameters. When a switch in the current global sound state variable is detected, all preset values ​​of sound effect parameters corresponding to the post-switching processing state are read from the sound effect parameter matrix; Within the preset transition time window, the sound field width control parameters of the virtual sound field synthesis module, the gain curve control parameters of each frequency band of the multi-band dynamic equalizer module, and the release time control parameters of the peak limiter are smoothly transitioned from their current values ​​to the newly read preset values, thus completing the synchronous update of the sound effect processing parameters.

[0042] It should be noted that in this embodiment, when the global sound state is switched, the parameters of the sound effect processing pipeline will be updated immediately. It is assumed that the sound effect processing pipeline is connected in series with the virtual sound field synthesis module, the multi-band dynamic equalizer module and the peak limiter module. A sound effect parameter matrix is ​​maintained, as shown in Table 1, which is displayed as the sound effect parameter matrix table.

[0043] Table 1. Sound Effect Parameter Matrix

[0044] Furthermore, in this embodiment, when a change in the global sound state variable is detected, all preset values ​​of sound effect parameters corresponding to the "musical sound restoration state" are immediately read from the aforementioned parameter matrix. To avoid abrupt parameter changes, within a very short preset transition time window (e.g., 20-50 milliseconds), linear interpolation or a first-order smoothing algorithm is used to smoothly transition the control parameters of the three modules from their current values ​​to the newly read preset values. Specifically, this includes the sound field width control parameter of the virtual sound field synthesis module, which smoothly shrinks from 120 degrees to 65 degrees; the high-frequency gain curve of the multi-band dynamic equalizer module, which smoothly drops from +4dB to 0dB; and the release time parameter of the peak limiter, which smoothly adjusts from a fixed 180ms to a dynamic value.

[0045] Furthermore, in this embodiment, the special processing of the musical tone state means that the release time of the peak limiter is not a fixed value in the "musical tone restoration state," but a dynamically calculated adaptive release value. Accordingly, the calculation process is as follows: extract the spectral amplitude sequence of the low-frequency sub-band in the audio in real time and calculate its amplitude envelope; perform fast Fourier transform analysis on the envelope signal in the range of 0.5Hz to 4Hz to search for the peak frequency as the estimated value of the beat frequency of the current music; multiply the reciprocal of this frequency (i.e., the beat interval duration) by a preset scaling factor (e.g., 0.4), and use the result as the release time of the current period, thereby better maintaining the dynamic rhythm of the music.

[0046] At this point, the audio processing pipeline has been switched to a parameter combination that perfectly matches the newly determined dominant sound type, achieving real-time, adaptive optimization of the audio signal and ultimately outputting highly immersive and expressive audio.

[0047] It is worth mentioning that the process of inputting the multi-dimensional frequency domain feature vector into the trained cascaded decision model involves the first-level coarse classifier calculating the spatial diffusion score and high-frequency transient attitude score based on the spectral flux of the high-frequency subband and the cross-channel level difference, and the second-level fine classifier calculating the soft-decision membership score of the current frame for the speech-dominated, transient sound effect-dominated, and continuous musical sound-dominated classes under the corresponding paths. Specifically, this includes: Extract the high-frequency subband spectral flux value and the cross-channel level difference value of the four subbands from the multi-dimensional frequency domain feature vector of the current audio frame, and concatenate them into a coarse classification input vector; The coarse classification input vector is input into the decision tree model of the trained first-level coarse classifier. The decision tree model outputs the result path of whether the current frame belongs to the concentrated sound image class or the diffuse sound image class according to the decision rules of the branch nodes. If the first-level coarse classifier determines that the current frame belongs to the concentrated sound image class, then the mid-low frequency sub-band spectral centroid, mid-low frequency sub-band spectral flux, and high frequency sub-band spectral flux are extracted from the multi-dimensional frequency domain feature vector and input to the speech / transient sound effect discrimination support vector machine branch model. The branch model outputs the first membership score of the current frame belonging to the speech-dominated class and the second membership score of the current frame belonging to the transient sound effect-dominated class, while setting the membership score of the current frame belonging to the continuous musical sound-dominated class to zero. If the first-level coarse classifier determines that the current frame belongs to the diffuse sound image class, then the mid-to-high frequency sub-band spectral centroid and mid-to-high frequency sub-band spectral flux are extracted from the multi-dimensional frequency domain feature vector and input to the continuous musical sound discrimination support vector machine branch model. The branch model outputs the third membership score of the current frame belonging to the continuous musical sound dominant class through sigmoid mapping, while the membership scores of the current frame belonging to the speech dominant class and the transient sound effect dominant class are both set to zero.

[0048] It should be noted that, in this embodiment, the first stage is a coarse classifier, which extracts the high-frequency subband spectral flux value and the cross-channel level difference value of the four subbands from the twelve-dimensional feature vector of the current frame to form a coarse classification input vector. This vector is then input into the trained first-stage coarse classifier—a binary classification decision tree model. Based on the decision rules learned during training (e.g., the absolute value of ILD is generally small and the high-frequency flux is high, tending to be concentrated sound image class), the decision tree model outputs the path classification of the current frame: Path A: concentrated sound image class, representing sound sources that are relatively concentrated in space, such as speech or transient sound effects from point sources; Path B: diffuse sound image class, representing sound sources that are widely distributed in space, such as continuous musical sounds or ambient sounds containing a lot of reverberation information.

[0049] Furthermore, in this embodiment, the second level is a fine-grained classifier for precise judgment. If the target path is A (concentrated audio-visual class), the mid-to-low frequency sub-band spectral centroid, mid-to-low frequency sub-band spectral flux, and high frequency sub-band spectral flux are extracted from the multi-dimensional frequency domain feature vector and used as input features. These features are then fed into the speech / transient sound effect discrimination support vector machine branch model in the second-level fine-grained classifier. This branch model will simultaneously output the first membership score (range 0-1) of the current frame belonging to the "speech-dominated class" and the second membership score of the current frame belonging to the "transient sound effect-dominated class", and forcibly set the membership score of the "continuous musical tone-dominated class" to 0. For example, speech typically has a stable centroid in the low to mid frequencies, while transient sound effects can produce high and sharp flux peaks in both low to mid frequencies and high frequencies. If the input is path B (diffuse sound image class), the centroid and flux of the mid to high frequency sub-band spectrum are extracted from the multi-dimensional frequency domain feature vector and used as input features. These are fed into the continuous musical tone discrimination support vector machine branch model in the second-level fine classifier. Correspondingly, this branch model outputs a raw decision value, which is then converted into a soft decision probability value in the range of 0 to 1 by a Sigmoid function mapping layer attached to its output. This is the third membership score of the current frame belonging to the "continuous musical tone dominant class". At the same time, the membership scores of both speech and transient sound effects are forcibly set to 0.

[0050] Specifically, in this embodiment, the present invention uses this cascaded structure to first divide the traffic according to spatial attributes, and then uses specially designed features to perform fine classification under different paths, which greatly improves the accuracy of discrimination and computational efficiency.

[0051] It is worth mentioning that the preset parameter values ​​stored in the sound effect parameter matrix specifically include: When the processing state is voice optimization state, the preset value of the virtual sound field width parameter is the first preset width value corresponding to the 30-degree narrow sound field of stereo. The preset value of the gain curve parameter of the multi-band dynamic equalizer in the low-to-mid frequency band is the dynamic equalization curve that applies automatic attenuation to the frequency band exceeding the preset energy threshold. The preset value of the gain curve parameter in the mid-to-high frequency band is the shelving gain curve that applies a 2dB to 3dB broadband boost. The preset value of the peak limiter release time parameter is the first preset fast release value in the range of 5ms to 10ms. When the processing state is transient enhancement state, the preset value of the virtual sound field width parameter is the second preset width value corresponding to the 120-degree full sound field of the surround sound, the preset value of the gain curve parameter of the multi-segment dynamic equalizer in the high frequency band is the shelf-type gain curve with a 3dB to 5dB high frequency boost, and the preset value of the peak limiter release time parameter is the second preset slow release value in the range of 150ms to 200ms. When the processing state is music reproduction state, the preset value of the virtual sound field width parameter is the third preset width value corresponding to the stereo 60-degree to 70-degree natural sound field, the preset value of the gain curve parameter of each frequency band of the multi-band dynamic equalizer is a flat reference gain curve of 0dB, and the preset value of the peak limiter release time parameter is an adaptive release value dynamically calculated according to the real-time music beat.

[0052] It should be noted that, in this embodiment, the core of configuring the preset values ​​for the sound effect parameter matrix lies in carefully matching a set of three-dimensional collaborative parameters, including "virtual sound field width, multi-band dynamic equalization, and peak limiter release time," for different dominant sound types to achieve a unique and ideal listening experience. Specifically, in voice optimization mode, a 30-degree narrow sound field is used to enhance center positioning, eliminating muddiness through dynamic attenuation of mid-low frequencies, slightly boosting mid-high frequencies to increase clarity, and setting a fast release time of 5-10ms to protect transient details and intelligibility of speech; in transient enhancement mode, the sound field is expanded to 120 degrees to create a sense of immersion and impact, significantly enhancing sound effect details and airiness through a 3-5dB boost in high frequencies, and setting a slow release time of 150-200ms to fully preserve the tail sound of transient impact; and in music reproduction mode, a 65-degree natural sound field is returned, and a 0dB flat equalization curve is used to achieve high-fidelity, uncolored music playback, with its release time more intelligently adjusted dynamically according to the real-time music beat.

[0053] It is worth mentioning that the dynamic calculation of the adaptive release value under the musical tone restoration state includes the following steps: In the music sound restoration state, the spectrum amplitude sequence of the low-frequency sub-band in the audio signal is extracted in real time, and the amplitude envelope signal is calculated based on the spectrum amplitude sequence. Spectral analysis is performed on the amplitude envelope signal. Within the preset music beat frequency search range of 0.5Hz to 4Hz, the frequency value corresponding to the peak value of the spectral amplitude is searched, and the searched peak frequency is used as the estimated value of the current music beat frequency. The reciprocal of the estimated music beat frequency is used as the beat interval duration. The beat interval duration is multiplied by a preset scaling factor, and the resulting product is used as the adaptive release value. The preset scaling factor ranges from 0.3 to 0.5.

[0054] It should be noted that, in this embodiment, the dynamic calculation of the release time under the musical tone restoration state is achieved through a process of "beat detection - interval calculation - proportional mapping". First, under the musical tone state, the spectral amplitude sequence is extracted from the low-frequency sub-band of the audio in real time and its amplitude envelope signal is calculated. Then, the envelope signal is subjected to spectral analysis within the typical musical beat frequency range of 0.5Hz to 4Hz (corresponding to 30 to 240 beats per minute), and the strongest frequency component corresponding to the amplitude peak is automatically searched as the estimated value of the beat frequency of the current music. Finally, the reciprocal of the beat frequency is calculated to obtain the beat interval duration, and then multiplied by a proportional coefficient of 0.3 to 0.5. This product is the adaptive release value. In this embodiment, the dynamic range controller can follow the rhythm of the music and release in time before the next strong beat, thereby maintaining the dynamic rhythm of the music and avoiding the loss of musicality caused by excessive compression.

[0055] Figure 4 A block diagram of a sound processing system based on audio spectral features according to the present invention is shown.

[0056] like Figure 4 As shown, this invention discloses a sound processing system based on audio spectrum features, including a memory and a processor. The memory includes a sound processing method program based on audio spectrum features. When the sound processing method program based on audio spectrum features is executed by the processor, it performs the following steps: The input multi-channel audio signal is acquired, and the spectrum of the audio frame of each channel is divided into multiple sub-bands. The spectral centroid and spectral flux of each sub-band are extracted. Calculate the cross-channel level difference between each channel on the corresponding sub-band, and construct a multi-dimensional frequency domain feature vector based on the spectral centroid, the spectral flux, and the cross-channel level difference; The multi-dimensional frequency domain feature vector is input into the trained cascaded decision model. The first-level coarse classifier calculates the spatial diffusion score and the high-frequency transient attitude score based on the spectral flux of the high-frequency sub-band and the cross-channel level difference. The second-level fine classifier calculates the soft decision membership score of the current frame for the speech-dominated class, transient sound effect-dominated class and continuous musical sound-dominated class respectively under the corresponding path. The membership scores are smoothed by time-axis filtering. When the smoothed membership score of any category exceeds the preset migration threshold and the overshoot state continues for a preset confirmation time, it is determined that a sound state migration has occurred. At this time, the current global sound state is switched to the processing state corresponding to the current category. Based on the switched processing state, the virtual sound field width parameter, the gain curve parameters of each frequency band of the multi-band dynamic equalizer, and the release time parameter of the peak limiter in the sound effect processing pipeline are adjusted synchronously to complete adaptive sound effect processing that matches the current sound state.

[0057] It should be noted that when the audio spectrum feature-based sound processing system disclosed in this invention is applied, the specific process corresponds to the audio spectrum feature-based sound processing method described in the above embodiments. Since the specific implementation details of the system application are consistent with the content of the above-mentioned audio spectrum feature-based sound processing method, no further details will be provided in this embodiment.

[0058] A third aspect of the present invention provides a computer-readable storage medium comprising a sound processing method program based on audio spectrum features, wherein when executed by a processor, the sound processing method program based on audio spectrum features implements the steps of the sound processing method based on audio spectrum features as described in any of the preceding claims.

[0059] A fourth aspect of the present invention provides a computer program product comprising: computer program code, which, when run on a computer, causes the computer to perform any of the methods described in the embodiments of the sound processing method based on audio spectrum features.

[0060] A fifth aspect of the present invention provides an electronic device comprising: a processor and a memory; wherein the memory is configured to store a computer program, and the processor is configured to execute the computer program stored in the memory to cause the electronic device to perform the steps of a sound processing method based on audio spectral features as described in any of the preceding claims.

[0061] The terms “component,” “module,” “system,” etc., used in this specification are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. As illustrated, applications running on computing devices and computing devices can both be components. One or more components may reside in a process and / or an execution thread, and components may be located on a single computer and / or distributed among two or more computers. Furthermore, these components can be executed from various computer-readable media on which various data structures are stored. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).

[0062] Those skilled in the art will recognize that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0063] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0064] In the embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0065] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0066] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0067] In the above embodiments, the functions of each functional unit can be implemented entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. A computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, all or part of the flow or function according to the embodiments of the present invention is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs (DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).

[0068] This invention discloses a sound processing method, system, medium, and device based on audio spectrum characteristics. By performing multi-dimensional spectrum morphology analysis on audio signals, it achieves fine-grained discrimination of dominant sound types in complex mixed content and smooth adaptive switching of sound effect strategies, solving the problems of coarse scene judgment granularity and abrupt switching of processing parameters in the prior art.

[0069] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.

Claims

1. A sound processing method based on audio spectral features, characterized by, Includes the following steps: Acquire multi-channel audio signals, divide the spectrum of each channel audio frame into sub-bands, and extract the spectral centroid and spectral flux of each sub-band; The cross-channel level difference between corresponding sub-bands between channels is calculated, and a multi-dimensional frequency domain feature vector is constructed based on the spectral centroid, spectral flux, and cross-channel level difference. The multi-dimensional frequency domain feature vector is input into the cascaded decision model. The first-level coarse classifier performs flow division based on the high-frequency subband spectral flux and the cross-channel level difference. The second-level fine classifier outputs the soft decision membership scores of the current frame for the speech-dominated class, transient sound effect-dominated class, and continuous musical sound-dominated class, respectively. The membership scores are smoothed by time-axis filtering. When the smoothed membership score of any category exceeds the preset transition threshold and continues to reach the preset confirmation time, the state transition is triggered, and the current global sound state is switched. Based on the switched processing state, the virtual sound field width parameter, the multi-segment dynamic equalizer gain curve parameter, and the peak limiter release time parameter in the sound effect processing pipeline are adjusted synchronously to complete adaptive sound effect processing.

2. The sound processing method based on audio spectral features according to claim 1, characterized in that, The process of acquiring multi-channel audio signals, dividing the spectrum of each channel audio frame into sub-bands, and extracting the spectral centroid and spectral flux of each sub-band specifically includes: The spectrum of the audio frame is divided into low-frequency sub-band, mid-low-frequency sub-band, mid-high-frequency sub-band, and high-frequency sub-band on the frequency axis according to the octave band method. The frequency range of the low-frequency sub-band is 20Hz to 80Hz, the frequency range of the mid-low-frequency sub-band is 80Hz to 500Hz, the frequency range of the mid-high-frequency sub-band is 500Hz to 4000Hz, and the frequency range of the high-frequency sub-band is 4000Hz to 20000Hz. Iterate through all frequency points within each sub-band and calculate the spectral centroid of the sub-band according to a preset formula; Calculate the amplitude difference between the current frame's spectral amplitude and the previous frame's spectral amplitude at each corresponding frequency point, and use the sum of the squares of the amplitude differences at all frequency points as the spectral flux value of the current sub-band. The spectral flux value is used to characterize the degree of drastic change in the spectral structure within the sub-band in the time domain.

3. The sound processing method based on audio spectral features according to claim 1, characterized in that, The calculation of the cross-channel level difference between corresponding sub-bands between audio channels is based on a multi-dimensional frequency domain feature vector constructed from the spectral centroid, spectral flux, and cross-channel level difference, specifically including: In a multi-channel audio signal, the left and right channels, which are spatially symmetrical in the channel configuration, are selected to form at least one channel pair. Based on the channel pair, calculate the total energy of the left channel sub-band and the total energy of the right channel sub-band for each sub-band of the left channel; The cross-channel level difference at the corresponding decibel scale of the current sub-band is calculated based on the total energy of the left channel sub-band and the total energy of the right channel sub-band, wherein the cross-channel level difference is used to reflect the spatial bias of the current sound source on the corresponding sub-band; The spectral centroid values, spectral flux values, and inter-channel level differences of the four sub-bands are concatenated in a preset order to form a twelve-dimensional frequency domain feature vector, which serves as the multi-dimensional frequency domain feature vector of the current audio frame.

4. The sound processing method based on audio spectral features according to claim 1, characterized in that, The training process of the cascaded decision model includes the following steps: Construct a training sample set, which includes multiple labeled pure speech segments, pure transient sound effect segments, and pure continuous musical sound segments. The annotation information of each segment includes the category label of the dominant sound type of the corresponding segment. For each audio segment in the training sample set, the multidimensional frequency domain feature vector is extracted frame by frame, and each frame is assigned a category label corresponding to its segment. Using the high-frequency subband spectral flux of each frame and the cross-channel level difference of the four subbands as input features, and the concentrated sound image class label or the diffuse sound image class label as output target, a binary classification decision tree model of the first-level coarse classifier is trained. Frames labeled as pure speech segments and pure transient sound effect segments are labeled as concentrated sound image class, and frames labeled as pure continuous musical sound segments are labeled as diffuse sound image class. The mid-to-low frequency subband spectral centroid, mid-to-low frequency subband spectral flux, and high frequency subband spectral flux of all training frames labeled as concentrated audio-visual classes are used as input features, and the speech-dominant class label and transient sound effect-dominant class label are used as output targets to train the speech / transient sound effect discrimination support vector machine branch model in the second-level fine classifier. The mid-to-high frequency subband spectral centroid and mid-to-high frequency subband spectral flux of all training frames labeled as diffuse sound image class are used as input features. The continuous musical tone dominant class label is used as the positive output target. The continuous musical tone discrimination support vector machine branch model in the second-level fine classifier is trained, and a sigmoid function mapping layer is added to the output of the branch model to map the decision value of the support vector machine to the membership score in the interval of 0 to 1.

5. The sound processing method based on audio spectral features according to claim 1, characterized in that, The membership score is smoothed over time. When the smoothed membership score of any category exceeds a preset transition threshold and continues for a preset confirmation duration, a state transition is triggered, switching the current global sound state. This specifically includes: Maintain a smooth membership score variable for each of the speech-dominated, transient sound effect-dominated, and continuous musical sound-dominated classes, with the initial value of each smooth membership score variable being zero. For each new input audio frame, after obtaining the three types of soft-decision membership scores of the current frame, a first-order infinite impulse response low-pass filter is used to recursively smooth the three types of membership scores respectively. The frame count limit is calculated based on the preset confirmation duration. When the smooth membership score exceeds the preset migration threshold, the corresponding frame counter is started, and the number of consecutive frames reaches the frame limit. Then, it is confirmed that the state migration condition is met, and the current global sound state variable is updated to the processing state identifier.

6. The sound processing method based on audio spectral features according to claim 1, characterized in that, Based on the switched processing state, the virtual sound field width parameter, multi-segment dynamic equalizer gain curve parameter, and peak limiter release time parameter in the sound effect processing pipeline are adjusted synchronously to complete adaptive sound effect processing, specifically including: The audio processing pipeline includes a virtual sound field synthesis module, a multi-band dynamic equalizer module, and a peak limiter module. Maintain the sound effect parameter matrix, which stores the preset values ​​of virtual sound field width parameter, multi-band dynamic equalizer gain curve parameter, and peak limiter release time parameter corresponding to the voice optimization state, transient enhancement state, and music reproduction state, respectively. When a change in the current global sound state variable is detected, the preset values ​​of the sound effect parameters corresponding to the post-change processing state are read from the sound effect parameter matrix; Within the preset transition time window, the control parameters of each module are smoothly transitioned from the current value to the newly read preset value, thus completing the synchronous update of the sound effect processing parameters.

7. A sound processing system based on audio spectral features, characterized by, The system includes a memory and a processor. The memory contains a sound processing method program based on audio spectrum features. When the sound processing method program based on audio spectrum features is executed by the processor, it performs the following steps: Acquire multi-channel audio signals, divide the spectrum of each channel audio frame into sub-bands, and extract the spectral centroid and spectral flux of each sub-band; The cross-channel level difference between corresponding sub-bands between channels is calculated, and a multi-dimensional frequency domain feature vector is constructed based on the spectral centroid, spectral flux, and cross-channel level difference. The multi-dimensional frequency domain feature vector is input into the cascaded decision model. The first-level coarse classifier performs flow division based on the high-frequency subband spectral flux and the cross-channel level difference. The second-level fine classifier outputs the soft decision membership scores of the current frame for the speech-dominated class, transient sound effect-dominated class, and continuous musical sound-dominated class, respectively. The membership scores are smoothed by time-axis filtering. When the smoothed membership score of any category exceeds the preset transition threshold and continues to reach the preset confirmation time, the state transition is triggered, and the current global sound state is switched. Based on the switched processing state, the virtual sound field width parameter, the multi-segment dynamic equalizer gain curve parameter, and the peak limiter release time parameter in the sound effect processing pipeline are adjusted synchronously to complete adaptive sound effect processing.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a sound processing method program based on audio spectrum features, which, when executed by a processor, implements the steps of a sound processing method based on audio spectrum features as described in any one of claims 1 to 6.

9. A computer program product, characterised in that, The computer program product includes computer program code, which, when run on a computer, causes the computer to implement the steps of a sound processing method based on audio spectral features as described in any one of claims 1 to 6.

10. An electronic device, comprising: The electronic device includes a processor and a memory; wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to cause the electronic device to perform the steps of a sound processing method based on audio spectrum features as described in any one of claims 1 to 6.