Adaptive scene sound effect adjustment method and device, and storage medium

By extracting acoustic fingerprint vectors through pre-trained convolutional neural networks and constructing noise reduction gain and sound field optimization matrices, the problem that existing sound effect adjustment schemes cannot adapt to different entertainment scenarios is solved, achieving efficient noise reduction and improved sound quality adaptability.

CN121284478BActive Publication Date: 2026-03-17CHENGDU XIAOCHANG TECH CO LTD

Patent Information

Application Number
CN202511845803.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-09
Publication Date
2026-03-17
Estimated Expiration
2045-12-09

AI Technical Summary

Technical Problem

Existing sound effect adjustment solutions fail to effectively address the time-frequency characteristics of noise in different entertainment scenarios, resulting in a disconnect between noise reduction effects and scenario requirements, and an inability to meet users' diverse needs for immersion, clarity, and dynamic performance.

Method used

A pre-trained convolutional neural network is used to extract acoustic fingerprint vectors. Through acoustic feature deconstruction and scene recognition, a time-frequency mask is constructed to generate a noise reduction gain matrix. The sound effect is then adjusted in conjunction with a sound field optimization matrix to specifically eliminate noise and optimize sound field parameters.

Benefits of technology

It achieves dynamic sound effect adjustment according to different scenarios, improves noise reduction accuracy and multi-scenario sound quality adaptability, and ensures the directional noise suppression effect and immersive sound quality in audio playback scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121284478B_ABST
    Figure CN121284478B_ABST
Patent Text Reader

Abstract

This invention relates to the field of audio processing technology, specifically to an adaptive scene sound effect adjustment method, device, and storage medium. The method includes: acquiring a mixed audio stream and the original playback audio; extracting an acoustic fingerprint vector using a pre-trained convolutional neural network; deconstructing the acoustic fingerprint vector to obtain multiple acoustic feature components; based on these components, identifying and determining the audio playback scene according to preset scene determination rules; constructing a time-frequency mask based on the audio playback scene; generating a noise reduction gain matrix based on the mask; simultaneously enhancing the sound field of the audio playback scene to obtain a sound field optimization matrix; and performing noise reduction and sound effect adjustment on the original playback audio based on the noise reduction gain matrix and the sound field optimization matrix, outputting the optimized playback audio. This invention can dynamically adapt sound effect adjustment according to different scenes, improve noise reduction accuracy, and enhance multi-scene sound quality adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio processing technology, and more specifically to an adaptive scene sound effect adjustment method, apparatus, and storage medium. Background Technology

[0002] With the rapid development of entertainment devices, integrated audio devices, represented by cinema karaoke systems, have gradually become the core of entertainment activities. They need to simultaneously meet the needs of voice optimization during movie watching, gaming, and karaoke, surround sound immersion during movie watching, dynamic sound effect reproduction during gaming, and high-fidelity requirements for daily music playback. This places higher demands on the adaptive capabilities of sound effect adjustment.

[0003] Currently, some multi-scenario adaptive sound effect adjustment solutions exist for smart speakers. For example, initial sound field parameters are generated through room spatial structure (such as length, width, and height dimensions, and wall materials) and sound field modeling. Noise reduction is achieved through multi-band gain adjustment, and sound effect mode switching is realized by combining audio content type recognition, thereby optimizing the consistency of audio spatial perception. Although this sound effect adjustment solution achieves sound effect adjustment based on environment and content to a certain extent, it does not consider the dynamic characteristics and type differences of noise in different entertainment scenarios. It only performs indiscriminate noise reduction through uniform multi-band gain and does not design targeted suppression strategies for the time-frequency characteristics of scene-specific noise (such as the temporal burst nature of transient noise and the frequency distribution of environmental noise). This results in a disconnect between noise reduction effect and scene requirements, either over-suppressing effective sound or failing to specifically eliminate target noise. Meanwhile, the existing sound effect adjustment solution switches sound effect modes based on the type of audio content, without relating it to the core sound quality requirements of entertainment scenarios. It does not extract the scene-specific acoustic characteristics, nor does it configure sound field parameters according to scene requirements. As a result, the sound effect adjustment only matches the content type and cannot adapt to the different needs of users for immersion, clarity, and dynamic performance in different entertainment scenarios. Ultimately, this leads to the problem of insufficient adaptability of the same mode for multiple scenarios. Summary of the Invention

[0004] In view of the shortcomings of the existing technology, the present invention provides an adaptive scene sound effect adjustment method, device and storage medium to improve the noise reduction accuracy of sound effect adjustment and the sound quality adaptability of multiple scenes.

[0005] To achieve the above objectives, the present invention employs the following technical solution:

[0006] This invention provides an adaptive scene sound effect adjustment method, including:

[0007] Acquire mixed audio streams and original playback audio, and use a pre-trained convolutional neural network to extract acoustic fingerprint vectors;

[0008] The acoustic fingerprint vector is deconstructed to obtain multiple acoustic feature components, and the multiple acoustic feature components are identified and judged according to the preset scene judgment rules to determine the audio playback scene.

[0009] A time-frequency mask is constructed based on the audio playback scenario, a noise reduction gain matrix is ​​generated based on the time-frequency mask, and the scene sound field is enhanced for the audio playback scenario to obtain a sound field optimization matrix.

[0010] The original playback audio is denoised and its sound effects are adjusted based on the noise reduction gain matrix and the sound field optimization matrix, and the optimized playback audio is output.

[0011] Preferably, the step of extracting the acoustic fingerprint vector using a pre-trained convolutional neural network involves the following steps:

[0012] Acquire the mixed audio stream captured by the microphone array and read the currently playing raw audio from the audio system;

[0013] A short-time Fourier transform is performed synchronously on the mixed audio stream and the original playback audio to obtain the spectrum matrix of the mixed audio stream and the spectrum matrix of the original playback audio.

[0014] The spectral matrix of the mixed audio stream, the spectral matrix of the original playback audio, the mixed audio stream, and the original playback audio are input into a pre-trained convolutional neural network to extract features, resulting in noise features, content features, and spatial features.

[0015] The noise features, content features, and spatial features are normalized, and the normalized noise features, content features, and spatial features are concatenated to form an acoustic fingerprint vector.

[0016] Preferably, the pre-trained convolutional neural network includes a noise processing branch, a content extraction branch, and a spatial localization branch; the step of inputting the spectrum matrix of the mixed audio stream, the spectrum matrix of the original playback audio, the mixed audio stream, and the original playback audio into the pre-trained convolutional neural network for feature extraction involves the following steps:

[0017] The noise processing branch of a pre-trained convolutional neural network is used to extract features from the spectral matrix of the mixed audio stream and the spectral matrix of the original playback audio, respectively, to obtain basic texture features. Then, the basic texture features are abstracted and fused to obtain noise features.

[0018] The original playback audio is input into the content extraction branch of the pre-trained convolutional neural network. The original playback audio is pre-emphasized and windowed in frames to extract the 13th order Mel-frequency cepstral coefficients. First-order difference MFCC and second-order difference MFCC are performed based on the 13th order Mel-frequency cepstral coefficients. The 13th order Mel-frequency cepstral coefficients are concatenated with the first-order difference MFCC and the second-order difference MFCC to obtain the content features.

[0019] The mixed audio stream acquired by the microphone array is input into the spatial localization branch of the pre-trained convolutional neural network. Based on the phase difference of the audio signals in the mixed audio stream, a multi-signal classification algorithm is used to estimate the direction of arrival of the sound field and obtain spatial features.

[0020] Preferably, the acoustic fingerprint vector is deconstructed to obtain multiple acoustic feature components, and the multiple acoustic feature components are then used to identify and determine the audio playback scene according to a preset scene determination rule. The following steps are performed:

[0021] Feature analysis is performed on the acoustic fingerprint vector to obtain noise sub-vector, content sub-vector and spatial sub-vector respectively;

[0022] The acoustic feature components are obtained by calculating the feature components of the noise subvector, content subvector, and spatial subvector respectively. The acoustic feature components include transient intensity, spectral entropy, dynamic range, spectral centroid, human voice intelligibility, surround sound intensity, and low-frequency energy proportion.

[0023] According to the preset scene determination rules, it is determined whether the multiple acoustic feature components meet the preset scene determination threshold conditions, and the preset scene that meets the determination threshold conditions is taken as the audio playback scene.

[0024] Preferably, the steps include: constructing a time-frequency mask based on the audio playback scenario, generating a noise reduction gain matrix based on the time-frequency mask, configuring specific sound field parameters based on the characteristics of the audio playback scenario and combining content- and space-related acoustic feature components, integrating the parameters into a matrix and calibrating through correlation constraints to obtain a sound field optimization matrix; and performing the following steps:

[0025] Based on the audio playback scenario, a time-frequency mask is constructed by selecting the corresponding masking method, and a noise reduction gain matrix is ​​generated based on the time-frequency mask; among which, the masking methods include soft mask, hard mask and hybrid mask;

[0026] Acoustic feature components related to content and space are selected, a sound field template is configured in conjunction with the audio playback scenario, and the sound field template is converted into a sound field parameter matrix. The sound field parameter matrix is ​​then collaboratively calibrated using preset content-space correlation constraints to obtain a sound field optimization matrix. The sound field optimization matrix includes frequency band gain parameters, spatial positioning parameters, and dynamic range parameters.

[0027] Preferably, the steps of selecting acoustic feature components related to content and space, configuring a sound field template in conjunction with the audio playback scenario, and performing the following steps are performed:

[0028] The spectral centroid, dynamic range, and low-frequency energy ratio were selected as content-related acoustic feature components, while the surround sound intensity was selected as spatially related acoustic feature components.

[0029] Configure the sound field template for the audio playback scene based on the content- and space-related acoustic feature components:

[0030] Movie viewing scenario: If the surround sound intensity is >0.7, multi-channel gain allocation is executed; at the same time, the subwoofer gain is adjusted according to the proportion of low-frequency energy.

[0031] Game scenario: If the surround sound intensity is >0.6, dynamically adjust the channel gain according to the direction of the sound source, and adjust the transient response parameters in combination with the dynamic range;

[0032] Karaoke scenario: If the surround sound intensity is <0.5, reduce the gain of the rear channels, adjust the EQ of the 3-5kHz frequency band by 1.5 dB according to the centroid of the spectrum, and configure the compression ratio according to the dynamic range.

[0033] Preferably, the original playback audio is denoised according to the denoising gain matrix, and the original playback audio is adjusted for sound effects according to the sound field optimization matrix to output the optimized playback audio. The following steps are performed:

[0034] A short-time Fourier transform is performed on the original playback audio to obtain the time-frequency matrix. The time-frequency matrix is ​​then multiplied element-wise with the noise reduction gain matrix. The noise reduction process is completed by directional attenuation of the noise-dominant time-frequency points through the gain coefficient.

[0035] The frequency band gain parameters in the sound field optimization matrix are called to optimize the frequency band energy of the noise-reduced time-spectrum matrix. The spatial distribution of the multi-channel signal is adjusted in combination with the spatial positioning parameters, and the dynamic range parameters are loaded simultaneously to optimize the audio transient response to enhance the scene sound effects.

[0036] The time-spectrum matrix after sound effect adjustment is converted into a time-domain audio signal by performing an inverse short-time Fourier transform. The real-time signal-to-noise ratio (SNR) and scene sound effect matching degree of the time-domain audio signal are then detected. If the real-time SNR is ≥15dB and the scene sound effect matching degree is ≥0.8, the optimized playback audio is directly output. Otherwise, the noise reduction gain and sound field parameters are finely adjusted based on the real-time SNR and scene sound effect matching degree detection results, and the detection is repeated until an audio signal that meets the sound effect requirements is output.

[0037] Preferably, it further includes: sampling the optimized playback audio in real time and calculating the signal-to-noise ratio and scene matching degree;

[0038] When the signal-to-noise ratio is <15dB or the scene matching degree is <0.8, secondary parameter optimization is triggered: the noise reduction gain matrix is ​​iteratively updated; the sound field optimization matrix calls the scene preset parameter package for compensation.

[0039] The optimization process continues until the signal-to-noise ratio and scene matching meet the standards or the maximum number of iterations is reached.

[0040] This invention also provides an adaptive scene sound effect adjustment device, comprising:

[0041] The feature extraction module is used to acquire the mixed audio stream and the original playback audio, and uses a pre-trained convolutional neural network to extract acoustic fingerprint vectors;

[0042] The scene recognition module is used to deconstruct the acoustic fingerprint vector into multiple acoustic feature components, and to perform scene recognition and judgment on the multiple acoustic feature components according to the preset scene judgment rules to determine the audio playback scene.

[0043] The parameter optimization calculation module is used to construct a time-frequency mask based on the audio playback scene, generate a noise reduction gain matrix based on the time-frequency mask, and simultaneously enhance the scene sound field of the audio playback scene to obtain a sound field optimization matrix.

[0044] The sound effect adjustment module is used to perform noise reduction and sound effect adjustment on the original playback audio based on the noise reduction gain matrix and the sound field optimization matrix, and output the optimized playback audio.

[0045] The present invention also provides a computer storage medium for storing program data, which, when executed by a computer, is used to implement the above-described adaptive scene sound effect adjustment method.

[0046] In summary, the beneficial effects of the present invention are as follows:

[0047] This invention extracts an acoustic fingerprint vector that integrates noise, content, and spatial features through a pre-trained convolutional neural network. After feature deconstruction to determine the audio playback scene, a noise reduction gain matrix is ​​constructed specifically, a sound field template is configured and converted into a sound field optimization matrix, and the noise reduction gain matrix and the sound field optimization matrix are used to process the original playback audio in a coordinated manner. While reducing noise in the original playback audio, the sound effects of the original playback audio are adjusted and optimized through the sound field optimization matrix. This achieves the goal of dynamically adapting sound effect adjustment strategies according to different scenes, improving noise reduction accuracy and multi-scene sound quality adaptability, and ensuring the noise-directed suppression effect and immersive sound quality of the audio in the audio playback scene. Attached Figure Description

[0048] Figure 1 This is a schematic diagram of the adaptive scene sound effect adjustment method of the present invention;

[0049] Figure 2 This is a flowchart of the acoustic fingerprint vector extraction process of the present invention;

[0050] Figure 3 This is a flowchart of the scene recognition process of the present invention;

[0051] Figure 4 This is a flowchart of the parameter optimization calculation process of the present invention;

[0052] Figure 5 This is a structural diagram of the adaptive scene sound effect adjustment device module of the present invention. Detailed Implementation

[0053] The present invention will be further described in detail below with reference to the accompanying drawings.

[0054] To make the objectives, solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments and accompanying drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and are not intended to limit the present invention.

[0055] The following is in conjunction with the appendix of this invention. Figures 1-5 The embodiments of the present invention will be described in detail below.

[0056] like Figure 1 As shown in the figure, this embodiment of the invention provides an adaptive scene sound effect adjustment method, the specific steps of which are as follows:

[0057] S1: Acquire the mixed audio stream and the original playback audio, and use a pre-trained convolutional neural network to extract the acoustic fingerprint vector.

[0058] Because mixed audio streams contain non-stationary noise such as transient pulses, broadband spectra, and harmonic drift, and pre-trained convolutional neural networks have strong local perception and weight sharing characteristics, they can efficiently capture subtle changes and complex structures in non-stationary noise signals and extract non-stationary noise features. Therefore, this embodiment of the invention uses a pre-trained convolutional neural network to extract an acoustic fingerprint vector containing noise features, content features, and spatial features. The acoustic fingerprint vector is a high-dimensional feature vector used to uniformly represent the three-dimensional attributes of audio noise, content, and space, and can serve as the core data basis for subsequent acoustic feature deconstruction, audio playback scene recognition, and the generation of noise reduction gain matrix and sound field optimization matrix.

[0059] S2: Perform acoustic feature deconstruction on the acoustic fingerprint vector to obtain multiple acoustic feature components. Based on the multiple acoustic feature components, perform scene recognition and judgment on the multiple acoustic feature components according to the preset scene judgment rules to determine the audio playback scene.

[0060] Because the feature components of the acoustic fingerprint vector generated in S1 are intertwined, directly performing scene recognition based on this vector may not accurately extract the contribution value of a single feature to scene determination. By deconstructing acoustic features, key features such as noise intensity, audio content category, and spatial reflection coefficient can be refined and separated. Then, based on multiple refined acoustic feature components, scene recognition and determination are performed according to preset scene determination rules, thereby more accurately determining the audio playback scene.

[0061] S3: Construct a time-frequency mask based on the audio playback scene, generate a noise reduction gain matrix based on the time-frequency mask, and simultaneously enhance the scene sound field of the audio playback scene to obtain a sound field optimization matrix.

[0062] Among them, by constructing a time-frequency mask to generate a noise reduction gain matrix, environmental noise can be accurately identified and suppressed, improving audio purity; while the sound field enhancement operation can optimize the sound field spatial sense in a targeted manner according to the acoustic characteristics of different playback scenarios (such as cinemas, cars, indoors, etc.), making the audio more in line with the needs of the scenario, and ultimately achieving the dual effects of noise reduction and sound field optimization, providing a high-quality audio foundation for subsequent sound effects processing.

[0063] S4: Based on the noise reduction gain matrix and sound field optimization matrix, the original playback audio is denoised and the sound effects are adjusted to output the optimized playback audio.

[0064] In S4, the noise reduction gain matrix and sound field parameter matrix generated in the early stage are used to perform targeted noise reduction processing on the original playback audio, eliminating environmental noise and interference. At the same time, the sound field effect of the audio is precisely adjusted according to the sound field parameter matrix to optimize the spatial sense, stereo effect and other sound characteristics of the audio. Finally, high-quality optimized playback audio that is adapted to the needs of specific scenarios is output, thereby bringing users a better listening experience.

[0065] In this embodiment of the invention, in step S1, a mixed audio stream and the original playback audio are acquired, and an acoustic fingerprint vector is extracted using a pre-trained convolutional neural network, referring to... Figure 2 As shown, the specific implementation steps are as follows:

[0066] S11: Acquire the mixed audio stream captured by the microphone array and read the currently playing raw audio from the audio system. The raw audio includes various pre-recorded audio content such as music streams and movie soundtracks, serving as the reference signal source for adaptive sound effect adjustment.

[0067] The microphone array collects mixed signals in the environment in real time to form a mixed audio stream. In addition to the reflected sound of the target audio, it also contains environmental noise such as air conditioner operation and human voices. Its signal characteristics change dynamically with physical space and time.

[0068] S12: Perform a short-time Fourier transform (STFT) on the mixed audio stream and the original playback audio simultaneously to obtain the spectrum matrix of the mixed audio stream and the spectrum matrix of the original playback audio.

[0069] Specifically, when using Short-Time Fourier Transform (STFT) to synchronize the mixed audio stream and the original playback audio, the parameters are set as follows: a Hanning window is selected as the window function, and the time-domain signals corresponding to the mixed audio stream and the original playback audio are converted into a two-dimensional spectral matrix Mspec through a sliding window. The vertical axis of this matrix represents the frequency dimension (0-44.1kHz), and the horizontal axis represents the time-domain dimension (in frames). Each element value represents the energy intensity of the corresponding time-frequency unit. By performing joint time-frequency domain analysis using STFT, the visualization and structured representation of the mixed audio stream and the original playback audio in both time and frequency domains are achieved.

[0070] S13: The spectral matrix of the mixed audio stream, the spectral matrix of the original playback audio, the mixed audio stream, and the original playback audio are input into a pre-trained convolutional neural network for feature extraction to obtain noise features, content features, and spatial features. Among them, noise features are extracted through the residual structure of the pre-trained convolutional neural network, content features are calculated using Mel-frequency cepstral coefficients (MFCC), and spatial features are obtained through sound field direction of arrival (DOA) estimation.

[0071] S14: Normalize the noise features, content features, and spatial features, and then concatenate the normalized noise features, content features, and spatial features to form an acoustic fingerprint vector, which serves as the core basis for subsequent sound effect adjustments.

[0072] In this embodiment of the invention, a mixed audio stream and the original playback audio are first acquired. The former includes ambient noise and target audio reflections, while the latter serves as a reference signal source. Then, the two are converted into a spectrum matrix using Short Time Fourier Transform (STFT) to achieve time-frequency dual-domain visualization and structured expression. This allows for a direct display of the frequency component distribution of the signal at different time points, helping to quickly locate noise frequency bands and effective signal regions. Next, noise, content, and spatial features are extracted using a pre-trained convolutional neural network. Compared to traditional single-dimensional feature extraction, this provides a more comprehensive description of the audio scene. Finally, the features are normalized and concatenated to form an acoustic fingerprint vector, which can be dynamically adjusted according to environmental changes. Compared to fixed-parameter sound effect adjustment techniques, this can adapt to complex and varied scenarios such as air conditioner operation noise and noisy human voices.

[0073] In embodiments of this invention, the pre-trained convolutional neural network for extracting acoustic fingerprint vectors includes a noise processing branch, a content extraction branch, and a spatial localization branch. The noise processing branch is specifically composed of basic residual blocks from existing pre-trained convolutional neural networks, achieving layer-by-layer abstraction of the input data through alternating stacking of 3×3 convolutional layers and batch normalization layers. The content extraction branch is used to calculate the Mel-frequency cepstral coefficients (MFCC) of the original played audio. The spatial localization branch obtains spatial features through direction of arrival (DOA) estimation using the Multiple Signal Classification (MUSIC) algorithm. Specifically, the noise processing branch uses existing ResNet-18 / ResNet-34 models, and the content extraction branch uses the existing Wav2Vec 2.0 model. Both the ResNet-18 / ResNet-34 and Wav2Vec 2.0 models are obtained from the torchvision.models model library in PyTorch. The spatial localization branch uses the existing SincNet model, obtained from an open-source repository on GitHub. In addition, other existing convolutional neural network models can be retrieved according to processing needs. The specific selection should be made according to the actual situation, which will not be elaborated here.

[0074] The pre-training process of a convolutional neural network is as follows:

[0075] First, a multi-scenario, multi-dimensional training dataset is constructed. The noise processing branch training data consists of paired samples of the spectrogram matrices of the mixed audio streams from movie-watching, gaming, and karaoke scenarios, and the corresponding spectrogram matrices of the original playback audio. The content extraction branch training data consists of MFCC feature data of the original audio from the three scenarios after pre-emphasis and frame-segmentation windowing processing. The spatial positioning branch training data consists of audio signals (including single-source and multi-source scenarios) collected by a microphone array from different scenarios, along with their corresponding Direction of Arrival (DOA) labels. The DOA labels are calculated offline using a multi-signal classification algorithm, with an azimuth accuracy of ±1°.

[0076] Next, we define hierarchical training labels. The overall scene labels can be labeled into three categories: movie watching, game playing, and karaoke singing. The sub-labels of each branch are as follows: the noise processing branch labels the noise type and noise energy percentage; the content extraction branch labels the audio content category (such as dialogue, sound effects, and human voice) and spectral feature parameters (such as spectral centroid and dynamic range); and the spatial positioning branch labels the azimuth and elevation angles.

[0077] Subsequently, each branch is trained in stages: the noise processing branch is based on paired spectrum matrix samples, using the loss function of minimizing the noise residual between the mixed spectrum and the original spectrum, and iteratively trained through a residual structure of alternating stacks of 3×3 convolutional kernels and batch normalization layers until the noise residual error is <5%; the content extraction branch is based on MFCC feature data, aiming to accurately classify audio content categories, and uses the cross-entropy loss function to train a cascaded feature extraction network of 13th-order MFCC and first-order and second-order difference MFCC; the spatial localization branch is based on microphone array signals and DOA annotation values, using the loss function of minimizing DOA estimation error to train a phase difference feature extraction and angle regression network.

[0078] Finally, multi-branch joint fine-tuning is performed. The features output by the three branches are concatenated and input into the fully connected layer. With the goal of maximizing the scene classification accuracy, the network is adapted to the association mapping of "noise-content-space" features in the three scenes. Finally, a pre-trained convolutional neural network is obtained, ensuring that its scene classification accuracy is ≥92%, noise feature extraction error is ≤3%, and DOA estimation error is ≤±3°.

[0079] Based on this pre-trained convolutional neural network structure, in step S12, the spectral matrix of the mixed audio stream, the spectral matrix of the original playback audio, the mixed audio stream, and the original playback audio are input into the pre-trained convolutional neural network for feature extraction to obtain noise features, content features, and spatial features. The specific implementation steps are as follows:

[0080] S121: The noise processing branch of the pre-trained convolutional neural network is used to extract features from the spectrum matrix of the mixed audio stream and the spectrum matrix of the original playback audio, respectively, to obtain basic texture features. This lays the foundation for the extraction of more abstract and representative deep features by the subsequent deeper network. The basic texture features are then abstracted and fused to obtain noise features. In environments with complex signal-to-noise ratios, this improves the recognition accuracy of common interferences such as Gaussian white noise and mechanical vibration noise.

[0081] S122: The original playback audio is input into the content extraction branch of a pre-trained convolutional neural network. The original playback audio undergoes pre-emphasis, frame-by-frame windowing, and 13th-order Mel-spectral coefficients (MSCs). Based on these MSCs, first-order and second-order difference MFCCs are performed. The MSCs are then concatenated with these MFCCs to obtain content features. MSCs simulate human auditory characteristics and can effectively characterize the timbre of audio. Based on the MSCs, the first-order difference MFCC reflects the rate of change of features over time, and the second-order difference MFCC captures the trend of this rate of change, thus fully preserving core content information such as the fundamental frequency of speech and the melody of music.

[0082] S123: Input the mixed audio stream acquired by the microphone array into the spatial localization branch of the pre-trained convolutional neural network. Based on the phase difference of the audio signals in the mixed audio stream, use a multi-signal classification algorithm to estimate the direction of arrival of the sound field and obtain spatial features.

[0083] Among them, the multi-signal classification algorithm constructs an array covariance matrix, uses eigenvalue decomposition to separate the signal subspace and noise subspace, scans the spectral peaks on a three-dimensional spatial grid of azimuth and elevation angles, and outputs a 6-dimensional spatial vector containing azimuth, elevation angle and their confidence. The algorithm can control the localization error of a single sound source within ±3° under an environment with a reverberation time RT60=300ms.

[0084] In this embodiment of the invention, step S2 involves acoustic feature deconstruction of the acoustic fingerprint vector to obtain multiple acoustic feature components. Based on these multiple acoustic feature components, scene recognition and determination are performed according to preset scene determination rules to determine the audio playback scene. Figure 3 As shown, the specific implementation steps are as follows:

[0085] S21: Perform feature analysis on the acoustic fingerprint vector to obtain the noise sub-vector, content sub-vector, and spatial sub-vector, providing a data foundation for subsequent component calculations.

[0086] S22: Calculate the feature components of the noise sub-vector, content sub-vector, and spatial sub-vector respectively to obtain the acoustic feature components.

[0087] The acoustic feature components include transient intensity, spectral entropy, dynamic range, spectral centroid, vocal intelligibility, surround sound intensity, and low-frequency energy proportion. Transient intensity captures the explosive power at the beginning of a sound, such as the characteristic identification of sudden, strong impact sounds like drumbeats or skill-triggered sound effects. Spectral entropy measures the uniformity of the spectral distribution. A low entropy value indicates that energy is concentrated in a few frequency bands, resulting in a pure sound (such as a monotone signal); a high entropy value indicates that energy is dispersed, resulting in a complex sound (such as white noise). Dynamic range refers to the level difference between the strongest and weakest parts of an audio signal. Spectral centroid indicates the location of the center of gravity of spectral energy. Vocal intelligibility assesses the intelligibility of a speech signal. Surround sound intensity quantifies the spatial distribution and intensity relationship of sound energy in each channel of a surround sound system, reflecting the spatial immersion and positioning accuracy of the sound. Low-frequency energy proportion statistically represents the proportion of low-frequency energy in the total energy of the audio signal, reflecting the weight and impact of the sound, such as the intensity characteristics of a deep bass effect. In the process of calculating acoustic feature components, for the dimensional data types of noise sub-vectors, content sub-vectors, and spatial sub-vectors, the corresponding feature indicators are calculated dimension by dimension of the sub-vector: for the time-domain energy and frequency-domain distribution data of each dimension of the noise sub-vector, the transient intensity and spectral entropy are calculated dimension by dimension; for the Mel-frequency cepstral coefficients (MFCC) and difference data of each dimension of the content sub-vector, the dynamic range, spectral centroid, vocal intelligibility, and low-frequency energy proportion are calculated dimension by dimension; for the sound field arrival direction and channel energy difference data of each dimension of the spatial sub-vector, the surround sound intensity is calculated dimension by dimension. Finally, acoustic feature components containing transient intensity, spectral entropy, dynamic range, spectral centroid, vocal intelligibility, surround sound intensity, and low-frequency energy proportion are obtained.

[0088] S23: According to the preset scene determination rules, determine whether multiple acoustic feature components meet the preset scene determination threshold conditions, and use the preset scenes that meet the determination threshold conditions as audio playback scenes. Among them, preset scenes include movie-watching scenes, game scenes, and karaoke scenes. In addition, audio playback scenes can be added according to the needs of other playback scenes. The specific settings and definitions can be made according to the actual situation, which will not be elaborated here.

[0089] Specifically, in one embodiment of the present invention, the determination threshold condition for the preset scenario in sub-step S23 is as follows:

[0090] Viewing scenario: Voice clarity > 0.8 indicates that the dialogue clarity meets the high-quality standard; surround sound intensity > 0.7 indicates that the multi-channel spatial sense is significant; spectral entropy < 1.5 is used to eliminate background noise interference.

[0091] Game scenario: Transient intensity > 0.7, indicating frequent occurrence of pulse signals such as explosions and collisions; Low frequency energy ratio > 60%, indicating that the sound effects are dominated by low frequency components such as engines and explosions; Dynamic verification: Dynamic range > 30dB, indicating that the sound effect intensity changes drastically.

[0092] Karaoke scenario: Dynamic range > 40dB, indicating significant fluctuations in singing volume. Spectral centroid > 2000Hz, indicating that high-frequency overtones dominate the vocals. Secondary verification: Vocal intelligibility > 0.6, used to ensure accuracy in non-instrumental solo scenarios.

[0093] In step S2 of this embodiment of the invention, accurate audio playback scene determination is achieved by deconstructing and recognizing the acoustic fingerprint vector. Compared with existing technologies, this step no longer relies on single-dimensional audio feature analysis, but decomposes the acoustic fingerprint vector into noise, content, and spatial sub-vectors. By calculating feature components dimension by dimension and combining them with preset judgment rules, it can more comprehensively and accurately identify different scenes such as watching movies, playing games, and singing karaoke. This effectively solves the problem of scene misjudgment in complex audio environments using traditional methods, significantly improves the reliability and adaptability of scene recognition, and provides a more accurate basis for subsequent targeted sound effect adjustments.

[0094] In this embodiment of the invention, step S3 involves constructing a time-frequency mask based on the audio playback scenario, generating a noise reduction gain matrix based on the time-frequency mask, and simultaneously configuring specific sound field parameters based on the characteristics of the audio playback scenario and the acoustic feature components related to content and space. These parameters are then integrated into a matrix and calibrated through correlation constraints to obtain a sound field optimization matrix. Figure 4 As shown, the specific implementation steps are as follows:

[0095] S31: Select the corresponding masking method based on the audio playback scenario to construct a time-frequency mask, and generate a noise reduction gain matrix based on the time-frequency mask. The matrix element values ​​range from [0,1]. Among them, the masking methods include soft mask, hard mask and hybrid mask.

[0096] Specifically, in one embodiment of the present invention, the sub-step S31 of selecting the corresponding masking method based on the audio playback scenario to construct the time-frequency mask specifically involves:

[0097] For movie-watching scenarios, a soft mask (mask=0.3~0.8) is used to preserve environmental sound details, creating an immersive movie-watching experience for users and enhancing the sense of immersion and atmosphere.

[0098] Using hard masking (mask=0~0.2) for game scenes suppresses transient interference in the scene, which helps players capture key sound effects in the game more clearly, such as footsteps and skill trigger sound effects, improving the accuracy and reaction speed of game operations and enhancing the game experience.

[0099] For karaoke scenarios, a hybrid masking approach is used. An adaptive threshold mask (where the threshold dynamically adjusts with the signal-to-noise ratio) is applied to the vocal frequency band, while a statistical mask is used for other frequency bands. This adapts to complex and ever-changing sound environments, ensuring that sound effects are always optimal. By developing differentiated masking strategies, the sound effect optimization needs of each scenario are precisely matched. The scenarios mentioned above are all preset scenarios; other scenarios can be set according to requirements to meet specific sound effect adjustment needs.

[0100] S32: Select acoustic feature components related to content and space, configure a sound field template in conjunction with the audio playback scene, convert the sound field template into a sound field parameter matrix, and use the pre-set content and space correlation constraints to perform collaborative calibration of the sound field parameter matrix to ensure the synergy of frequency band gain, spatial positioning and dynamic response, and finally generate the optimal sound field parameter matrix adapted to the current scene.

[0101] The sound field optimization matrix includes frequency band gain parameters, spatial positioning parameters, and dynamic range parameters.

[0102] This invention generates a noise reduction gain matrix by constructing a time-frequency mask and obtains a sound field optimization matrix by configuring and calibrating sound field parameters. This allows for precise sound effect adjustment based on different audio playback scenarios. Compared to existing technologies, its advantages are: firstly, it employs a differentiated masking strategy, selecting soft masks, hard masks, and hybrid masks for different scenarios such as movie watching, gaming, and karaoke, precisely meeting the sound effect optimization needs of each scenario. For example, it preserves environmental sound details during movie watching, suppresses interference and highlights key sound effects in games, and adapts to complex sound environments in karaoke. Secondly, by selecting acoustic feature components, configuring and calibrating the sound field template according to the scene, it optimizes parameters such as frequency band gain, spatial positioning, and dynamic range, providing users with a more immersive, clear, and adaptable sound experience. This effectively solves the problem that existing technologies have strong universality in sound effect adjustment but lack scene specificity.

[0103] In one embodiment, sub-step S32 selects acoustic feature components related to content and space, and configures a sound field template in conjunction with the audio playback scenario. The specific implementation steps are as follows:

[0104] S321: Select the spectral centroid, dynamic range, and low-frequency energy ratio as content-related acoustic feature components, and select the surround sound intensity as spatially related acoustic feature components.

[0105] S322: Configure the sound field template for the audio playback scene based on the content- and space-related acoustic feature components.

[0106] Movie viewing scenario: If the surround sound intensity is >0.7, multi-channel gain allocation will be implemented, such as front channel ×1.2, rear surround channel ×1.5, and center vocal channel ×1.3; at the same time, the subwoofer gain will be adjusted according to the proportion of low frequency energy, for example, when the proportion of low frequency energy is <20%, it will be automatically increased by 6dB to ensure that the surround immersion and low frequency impact are adapted to the movie's sound effects requirements.

[0107] Game scenario: If the surround sound intensity is >0.6, dynamically adjust the channel gain according to the direction of the sound source, and adjust the transient response parameters in combination with the dynamic range.

[0108] For example, by increasing the directional channel of the skill trigger sound effect by 2dB and decreasing the inverse channel by 1dB, and combining this with the dynamic range of >30dB, the transient response parameters are optimized. Specifically, the Attack time is set to 3ms and the Release time to 40ms, thereby enhancing the spatial positioning and dynamic performance of the pulse sound effect.

[0109] Karaoke scenario: If the surround sound intensity is <0.5, reduce the gain of the rear channels to weaken the surround sound effect and highlight the center positioning of the vocals. Then, adjust the EQ of the 3-5kHz frequency band by 1.5 dB based on the centroid of the frequency spectrum, and configure the compression ratio according to the dynamic range. For example, when the dynamic range is >40dB (at which point the volume fluctuation is large), configure the compression ratio to 2:1 to avoid overloading the vocals.

[0110] In this embodiment, by selecting acoustic feature components related to content and space, such as the spectral centroid and dynamic range, and accurately configuring sound field templates for different audio playback scenarios such as watching movies, playing games, and singing karaoke based on their component values, the sound effects of the original playback audio are adjusted by matching the sound field optimization matrix corresponding to the sound field template. This can achieve effects such as surround immersion, low-frequency impact, spatial positioning, and prominent vocals, meeting diverse sound effect adjustment needs.

[0111] In this embodiment of the invention, step S4 involves denoising and adjusting the sound effects of the original playback audio based on the noise reduction gain matrix and the sound field optimization matrix, and outputting the optimized playback audio. The specific implementation steps are as follows:

[0112] S41: Perform a short-time Fourier transform on the original playback audio to obtain the time-frequency matrix. Multiply the time-frequency matrix element by element with the noise reduction gain matrix. Use the gain coefficient to perform directional attenuation on the noise-dominant time-frequency points to complete the noise reduction process. While preserving the effective audio signal (such as ambient sound in a movie-watching scene or human voice in a karaoke scene), the noise component is directionally weakened to avoid the sound quality distortion caused by traditional global noise reduction.

[0113] S42: Call the frequency band gain parameters in the sound field optimization matrix to optimize the frequency band energy of the noise-reduced time-spectrum matrix, combine the spatial positioning parameters to adjust the spatial distribution of the multi-channel signal, and simultaneously load the dynamic range parameters to optimize the audio transient response to enhance the scene sound effects, so that the sound effect adjustment is deeply adapted to the sound quality requirements of the current playback scene (movie watching / game / karaoke) to achieve scene-specific sound quality enhancement.

[0114] S43: Perform an inverse short-time Fourier transform on the time-spectrum matrix after sound effect adjustment to convert it into a time-domain audio signal, and perform real-time signal-to-noise ratio and scene sound effect matching degree detection on the time-domain audio signal. If the real-time signal-to-noise ratio is ≥15dB and the scene sound effect matching degree is ≥0.8, the optimized playback audio is directly output. Otherwise, the noise reduction gain and sound field parameters are finely adjusted based on the real-time signal-to-noise ratio and scene sound effect matching degree detection results, and the detection is repeated until an audio signal that meets the sound effect requirements is output, ensuring that the output audio meets the preset standard.

[0115] This invention systematically processes the original playback audio using a noise reduction gain matrix and a sound field optimization matrix, achieving precise noise reduction and scene-appropriate sound effect adjustment. It effectively suppresses noise interference and significantly improves the scene-specific sound quality of the audio. Compared with traditional global noise reduction methods, it can effectively suppress environmental noise (such as background noise in movie-watching scenes and electrical noise in karaoke scenes) while completely preserving the core audio signal. It avoids problems such as blurred vocals and distorted instruments caused by excessive noise reduction, achieving a balance between noise reduction and sound quality.

[0116] In this embodiment of the invention, the adaptive scene sound effect adjustment method further includes S5:

[0117] The optimized playback audio is sampled in real time to calculate the signal-to-noise ratio (SNR) and scene matching degree. When the SNR is <15dB or the scene matching degree is <0.8, secondary parameter optimization is triggered: the noise reduction gain matrix is ​​iteratively updated; the sound field optimization matrix calls the scene preset parameter package for compensation; the optimization process continues until the SNR and scene matching degree meet the standards or the maximum number of iterations is reached, ensuring high-quality output audio and improving the accuracy and stability of adaptive scene sound effect adjustment.

[0118] To implement the above-mentioned adaptive scene sound effect adjustment method, embodiments of the present invention also provide an adaptive scene sound effect adjustment device, such as... Figure 5 As shown, the device includes:

[0119] The feature extraction module is used to acquire the mixed audio stream and the original playback audio, and uses a pre-trained convolutional neural network to extract acoustic fingerprint vectors;

[0120] The scene recognition module is used to deconstruct the acoustic fingerprint vector into multiple acoustic feature components, and to perform scene recognition and judgment on the multiple acoustic feature components according to the preset scene judgment rules to determine the audio playback scene.

[0121] The parameter optimization calculation module is used to construct a time-frequency mask based on the audio playback scene, generate a noise reduction gain matrix based on the time-frequency mask, and simultaneously enhance the scene sound field of the audio playback scene to obtain a sound field optimization matrix.

[0122] The sound effect adjustment module is used to perform noise reduction and sound effect adjustment on the original playback audio based on the noise reduction gain matrix and the sound field optimization matrix, and output the optimized playback audio.

[0123] In this embodiment of the invention, adaptive sound effect adjustment is achieved through the coordinated operation of four major modules:

[0124] First, the feature extraction module collects the mixed audio stream acquired by the microphone and the original playback audio output by the sound system. It then uses a pre-trained convolutional neural network (including residual structure) to extract non-steady-state noise features, Mel-frequency cepstral coefficient (MFCC) content features, and direction of arrival (DOA) spatial features, and fuses them to generate a multi-dimensional acoustic fingerprint vector.

[0125] Next, the scene recognition module receives the acoustic fingerprint vector, deconstructs seven feature components such as transient intensity and surround sound intensity, and determines the current audio playback scene (movie / game / karaoke) based on preset rules (e.g., for movie-watching scenes, human voice clarity > 0.8 and surround sound intensity > 0.7). Subsequently, the parameter optimization calculation module constructs a dedicated time-frequency mask for the determined scene (e.g., for game scenes, a hard mask is used to suppress transient interference), generates a noise reduction gain matrix, and combines scene features to configure a sound field template (e.g., for karaoke scenes, EQ adjustment is applied to the 3-5kHz frequency band), integrating them into a sound field optimization matrix.

[0126] Finally, the sound effect adjustment module applies the noise reduction gain matrix and the sound field optimization matrix to the original playback audio, and completes directional noise reduction and scene-based sound effect enhancement through time-frequency domain calculations, outputting optimized playback audio. If the signal-to-noise ratio is detected to be <15dB or the scene matching degree is <0.8 in real time, it will also trigger secondary parameter optimization to ensure that the sound effect is stable and meets the standards.

[0127] The adaptive scene sound effect adjustment device module of this invention has a clear division of labor and works together efficiently. It relies on multi-dimensional acoustic fingerprint vectors and multi-feature scene judgment rules to achieve accurate scene adaptation. It can generate scene-specific noise reduction and sound field parameters to avoid one-size-fits-all sound effect adjustment. It ensures stable output audio quality through a real-time feedback optimization mechanism. At the same time, it is compatible with multiple audio types and terminal devices, can adapt to complex noise environments, and takes into account both a high-quality user experience and a wide range of applications.

[0128] This invention also provides a computer storage medium for storing program data, which, when executed by a computer, is used to implement the adaptive scene sound effect adjustment method described above.

[0129] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An adaptive scene sound effect adjustment method, characterized in that, The method comprises the following steps: Collecting mixed audio stream and original playing audio, and extracting acoustic fingerprint vector by using pre-trained convolutional neural network; the acoustic fingerprint vector is formed by normalizing and splicing noise features, content features and space features, and is used for uniformly representing three-dimensional attributes of audio noise, content and space; Acoustic feature deconstruction is performed on the acoustic fingerprint vector to obtain a plurality of acoustic feature components, and scene recognition and determination are performed on the plurality of acoustic feature components according to a preset scene determination rule to determine an audio playing scene; The acoustic feature deconstruction on the acoustic fingerprint vector comprises: feature analysis on the acoustic fingerprint vector to obtain noise sub-vectors, content sub-vectors and space sub-vectors respectively; and feature component calculation is performed on the noise sub-vectors, content sub-vectors and space sub-vectors respectively to obtain acoustic feature components; A time-frequency mask is constructed according to the audio playing scene, a noise reduction gain matrix is generated based on the time-frequency mask, and based on the characteristics of the audio playing scene, exclusive sound field parameters are configured by combining the content and space related acoustic feature components, the parameters are integrated into a matrix and calibrated through correlation constraint to obtain a sound field optimization matrix; The original playing audio is subjected to noise reduction according to the noise reduction gain matrix, and sound effect adjustment is performed on the original playing audio according to the sound field optimization matrix, and the optimized playing audio is output.

2. The adaptive scene audio adjustment method of claim 1, wherein, The following steps are performed to extract the acoustic fingerprint vector by using the pre-trained convolutional neural network: The mixed audio stream collected by the microphone array is obtained, and the original playing audio currently played by the sound system is read; Short-time Fourier transform is performed on the mixed audio stream and the original playing audio synchronously to obtain a frequency spectrum matrix of the mixed audio stream and a frequency spectrum matrix of the original playing audio; The frequency spectrum matrix of the mixed audio stream and the frequency spectrum matrix of the original playing audio, the mixed audio stream and the original playing audio are input into the pre-trained convolutional neural network for feature extraction to obtain noise features, content features and space features; The noise features, content features and space features are normalized, and the normalized noise features, content features and space features are spliced to form an acoustic fingerprint vector.

3. The adaptive scene audio adjustment method of claim 2, wherein, The pre-trained convolutional neural network comprises a noise processing branch, a content extraction branch and a space positioning branch; the following steps are performed to input the frequency spectrum matrix of the mixed audio stream and the frequency spectrum matrix of the original playing audio, the mixed audio stream and the original playing audio into the pre-trained convolutional neural network for feature extraction: The noise processing branch of the pre-trained convolutional neural network is used to perform feature extraction processing on the frequency spectrum matrix of the mixed audio stream and the frequency spectrum matrix of the original playing audio respectively to obtain basic texture features, and the basic texture features are abstractly fused to obtain noise features; The original playing audio is input into the content extraction branch of the pre-trained convolutional neural network, pre-emphasis and frame windowing processing are performed on the original playing audio, 13-order mel-frequency cepstral coefficients are extracted, first-order difference MFCC and second-order difference MFCC are performed based on the 13-order mel-frequency cepstral coefficients, and the 13-order mel-frequency cepstral coefficients, the first-order difference MFCC and the second-order difference MFCC are cascaded to obtain content features; The mixed audio stream collected by the microphone array is input into a spatial positioning branch of a pre-trained convolutional neural network, and based on the phase difference of the audio signals in the mixed audio stream, a multiple signal classification algorithm is used to estimate the direction of arrival of the sound field to obtain spatial features.

4. The adaptive scene audio adjustment method of claim 1, wherein, The acoustic fingerprint vector is deconstructed based on acoustic features to obtain a plurality of acoustic feature components, and scene recognition and determination are performed on the plurality of acoustic feature components according to a preset scene determination rule to determine an audio playing scene, and the following steps are performed: The acoustic fingerprint vector is deconstructed based on acoustic features to obtain a plurality of acoustic feature components, and scene recognition and determination are performed on the plurality of acoustic feature components according to a preset scene determination rule to determine an audio playing scene, and the following steps are performed: The acoustic fingerprint vector is deconstructed based on acoustic features to obtain a plurality of acoustic feature components, and scene recognition and determination are performed on the plurality of acoustic feature components according to a preset scene determination rule to determine an audio playing scene, and the following steps are performed: The acoustic feature components include transient intensity, spectral entropy, dynamic range, spectral centroid, vocal clarity, surround sound intensity, and low-frequency energy proportion. According to the preset scene determination rule, it is judged whether the plurality of acoustic feature components meet the preset scene determination threshold condition, and the preset scene meeting the determination threshold condition is taken as the audio playing scene.

5. The adaptive scene audio adjustment method of claim 1, wherein, According to the audio playing scene, a time-frequency mask is constructed, a noise reduction gain matrix is generated based on the time-frequency mask, and based on the characteristics of the audio playing scene, the content and space related acoustic feature components are combined to configure exclusive sound field parameters, the parameters are integrated into a matrix and calibrated through relevance constraints to obtain a sound field optimization matrix, and the following steps are performed: Based on the audio playing scene, a corresponding mask mode is selected to construct a time-frequency mask, and a noise reduction gain matrix is generated based on the time-frequency mask; wherein the mask mode includes soft mask, hard mask and hybrid mask. Selecting acoustic feature components related to content and space, configuring a sound field template in combination with the audio playing scene, and converting the sound field template into a sound field parameter matrix, and using the preset relevance constraints of content and space to cooperatively calibrate the sound field parameter matrix to obtain a sound field optimization matrix; wherein the sound field optimization matrix includes frequency band gain parameters, spatial positioning parameters and dynamic range parameters.

6. The adaptive scene audio adjustment method of claim 5, wherein, Selecting acoustic feature components related to content and space, configuring a sound field template in combination with the audio playing scene, and the following steps are performed: Selecting spectral centroid, dynamic range, and low-frequency energy proportion as content-related acoustic feature components, and selecting surround sound intensity as space-related acoustic feature components; According to the content and space related acoustic feature components, the sound field template of the audio playing scene is configured: Movie scene: if the surround sound intensity is greater than 0.7, perform multi-channel gain distribution; at the same time, adjust the subwoofer gain according to the low-frequency energy proportion; Game scene: if the surround sound intensity is greater than 0.6, dynamically adjust the channel gain according to the sound source direction, and adjust the transient response parameter in combination with the dynamic range; K song scene: if the surround sound intensity is less than 0.5, reduce the rear channel gain, adjust the 3-5 kHz frequency band by 1.5 dB EQ according to the spectral centroid, and configure the compression ratio according to the dynamic range.

7. The adaptive scene audio adjustment method of claim 1, wherein, According to the noise reduction gain matrix, the original playing audio is subjected to noise reduction, and according to the sound field optimization matrix, the original playing audio is subjected to sound effect adjustment, and an optimized playing audio is output, and the following steps are performed: The short-time Fourier transform is performed on the original playing audio to obtain a time-frequency spectrum matrix, the time-frequency spectrum matrix is multiplied by a noise reduction gain matrix element by element, and the time-frequency points dominated by noise are attenuated in direction through the gain coefficient to complete the noise reduction processing; The frequency band energy of the time-frequency spectrum matrix after noise reduction is optimized by calling the frequency band gain parameters in the sound field optimization matrix, the spatial distribution of the multi-channel signal is adjusted in combination with the spatial positioning parameters, and the dynamic range parameters are simultaneously loaded to optimize the audio transient response to strengthen the scene sound effect; The time-frequency spectrum matrix after sound effect adjustment processing is converted into a time-domain audio signal by performing inverse short-time Fourier transform, and real-time signal-to-noise ratio and scene sound effect matching degree detection is performed on the time-domain audio signal, if the real-time signal-to-noise ratio is greater than or equal to 15 dB and the scene sound effect matching degree is greater than or equal to 0.8, the optimized playing audio is directly output, otherwise, the noise reduction gain and the sound field parameters are fine-tuned based on the real-time signal-to-noise ratio and the scene sound effect matching degree detection results, and the detection is repeated until the audio signal meeting the sound effect requirements is output.

8. The adaptive scene audio adjustment method of claim 1, wherein, Further comprising: real-time sampling of the optimized playing audio, calculation of the signal-to-noise ratio and the scene matching degree; When the signal-to-noise ratio is less than 15 dB or the scene matching degree is less than 0.8, parameter secondary optimization is triggered: the noise reduction gain matrix is iteratively updated; the sound field optimization matrix calls the scene preset parameter package for compensation; The optimization process continues until the signal-to-noise ratio and the scene matching degree meet the standards or the maximum number of iterations is reached.

9. An adaptive scene sound effect adjustment apparatus, characterized by, Comprise: A feature extraction module is configured to collect a mixed audio stream and an original playing audio, and extract an acoustic fingerprint vector by using a pre-trained convolutional neural network; the acoustic fingerprint vector is formed by normalizing and splicing three features of noise features, content features and spatial features, and is used as a feature vector for uniformly representing three-dimensional attributes of audio noise, content and space; A scene recognition module is configured to deconstruct acoustic features of the acoustic fingerprint vector to obtain a plurality of acoustic feature components, and perform scene recognition and determination on the plurality of acoustic feature components according to a preset scene determination rule to determine an audio playing scene; The acoustic feature deconstruction of the acoustic fingerprint vector comprises: feature analysis of the acoustic fingerprint vector to obtain noise sub-vectors, content sub-vectors and spatial sub-vectors respectively; and feature component calculation of the noise sub-vectors, the content sub-vectors and the spatial sub-vectors respectively to obtain acoustic feature components; A parameter optimization calculation module is configured to construct a time-frequency mask according to the audio playing scene, generate a noise reduction gain matrix based on the time-frequency mask, and perform scene sound field strengthening on the audio playing scene to obtain a sound field optimization matrix; An sound effect adjustment module is configured to perform noise reduction and sound effect adjustment on the original playing audio according to the noise reduction gain matrix and the sound field optimization matrix, and output the optimized playing audio.

10. A computer storage medium, characterized in that, The computer storage medium is configured to store program data, which, when executed by a computer, implements the adaptive scene sound effect adjustment method of any one of claims 1 to 8.

Citation Information

Patent Citations

  • Audio rendering method, storage medium and electronic device

    CN119729331A

  • Sound box sound effect intelligent adjustment method and system based on data analysis, and storage medium

    CN120751310A

Cited By

  • Bluetooth sound equipment intelligent sound effect adjusting system based on adaptive noise reduction

    CN122093725A