Sound box sound effect optimization method and system based on deep learning
By constructing a reference-free audio quality characterization learning model and integrating the human auditory perception model, combined with physiological signal feedback, personalized sound optimization is achieved, which solves the limitations of existing audio quality assessment methods and improves audio playback quality and user experience.
Patent Information
- Application Number
- CN202510678289.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-10-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing audio quality assessment methods rely on reference audio that is difficult to obtain, have limited generalization capabilities, and cannot be personalized to adapt to user auditory characteristics. Traditional sound optimization strategies cannot meet the differentiated needs of different users and do not take into account the perceptual characteristics of the human auditory system.
Build a reference-free audio quality characterization learning model that integrates the human auditory perception model and physiological signal feedback. Use deep learning technology to extract quality-sensitive features, perform personalized sound optimization, and adjust sound parameters in real time to adapt to the user's individual auditory characteristics.
It achieves accurate audio quality assessment without relying on reference audio, personalized sound optimization, improves audio playback quality and user listening experience, meets the differentiated needs of different users, and improves the refinement and consistency of sound optimization.
Smart Images

Figure CN120751309A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio signal processing, and more specifically, to a method and system for optimizing sound effects based on deep learning. Background Art
[0002] As people's demand for audio content consumption continues to increase, sound quality optimization has become one of the key technologies for various audio devices. However, current audio systems generally have the following problems in sound quality optimization:
[0003] Traditional audio quality assessment mainly relies on full-reference methods, which require comparing and analyzing the playback audio with high-quality reference audio. For example, commonly used objective evaluation metrics such as PESQ (Perceptual Evaluation of Speech Quality) and PEAQ (Perceptual Evaluation of Audio Quality) rely on original reference audio. However, in actual audio playback environments, it is often difficult to obtain reference audio corresponding to the current playback content, which severely limits the practicality of traditional methods. Existing no-reference audio quality assessment methods are mostly based on hand-designed features or shallow machine learning models, such as 3QUEST and ANIQUE+ based on statistical features. These methods are usually designed for specific types of distortion. When faced with complex and diverse audio content and quality issues, their generalization capabilities are limited, making it difficult to accurately identify and evaluate different types of quality issues. Traditional sound optimization methods mostly use preset equalizer curves or fixed sound processing parameters, such as Dolby audio and DTS audio. These optimization methods use fixed rules and parameter adjustments, without considering the specific quality issues of different audio content, nor the perceptual characteristics of the human auditory system, resulting in a gap between the optimization effect and the user's subjective feelings. Different users have individual differences in their perception of audio content, which stems from factors such as the physiological differences in the individual auditory system, the degree of hearing damage, and personal preferences. Existing technologies often use a "one-size-fits-all" optimization strategy, which is unable to provide personalized sound optimization based on the auditory characteristics and preferences of individual users, making it difficult to meet the differentiated needs of different users. Deep learning technology has made significant breakthroughs in computer vision, natural language processing, and other fields, and has begun to be applied to the field of audio signal processing. However, these methods often require a large amount of labeled data for training and do not fully consider the human auditory perception mechanism. In practical applications, they still face problems of insufficient generalization ability and poor subjective consistency.
[0004] Therefore, there is an urgent need for a technical solution that does not rely on reference audio, can accurately evaluate audio quality and perform personalized sound optimization to improve the quality of audio playback and user listening experience. Summary of the Invention
[0005] The present invention provides a sound effect optimization method and system based on deep learning, which solves the technical problems of reference-free sound quality evaluation, personalized adaptation and auditory perception model fusion in related technologies.
[0006] The present invention provides a method for optimizing sound effects based on deep learning, comprising:
[0007] Build a reference-free audio quality characterization learning model to extract quality-sensitive feature representations from audio signals;
[0008] By integrating the human auditory perception model, the audio quality characterization model is optimized using quality-sensitive feature representation;
[0009] Based on the optimization results of the audio quality characterization model, personalized fine-tuning is performed using physiological signal feedback to adapt to the user's individual hearing characteristics;
[0010] The system uses a personalized, fine-tuned audio quality characterization model to evaluate the quality of the audio playback in real time, and dynamically adjusts sound effect parameters based on the evaluation results for adaptive sound effect calibration.
[0011] By combining the real-time evaluation of the quality results of the played audio and the dynamic adjustment of the sound effect parameters based on the evaluation results, multi-dimensional quality problem diagnosis and fine-tuning are carried out, and targeted optimization is carried out for quality problems in different dimensions.
[0012] Furthermore, the step of constructing a no-reference audio quality characterization learning model includes:
[0013] The input original audio signal is framed and windowed, and converted into time-frequency representation;
[0014] Construct a self-supervised learning task to apply multiple known types of distortion to the input audio and train the model to predict the type and severity of distortion from the distorted audio;
[0015] A convolutional neural network is used as the encoder backbone network to extract hierarchical features through multi-layer convolution and pooling operations;
[0016] A multi-task learning framework is constructed to jointly optimize distortion classification, distortion degree regression, and feature reconstruction tasks.
[0017] Furthermore, the self-supervised learning task also includes a context prediction task and a contrastive learning task, wherein:
[0018] The context prediction task segments the audio into adjacent segments and trains the model to predict the temporal relationship between the segments;
[0019] The contrastive learning task considers different time-frequency transformations of the same audio as positive sample pairs, and transformations of different audio as negative sample pairs, and trains the model to learn invariant features related to perceptual quality.
[0020] Furthermore, the step of integrating the human auditory perception model includes:
[0021] Build a computational model of human hearing that includes auditory filter banks, loudness perception nonlinearities, and time-frequency masking effects;
[0022] In self-supervised learning tasks, the distortion introduction method is dynamically adjusted according to the characteristics of human hearing;
[0023] A loss function based on auditory perception is introduced to replace the conventional reconstruction loss.
[0024] Furthermore, the human auditory computational model also includes binaural hearing characteristics and auditory fatigue models, wherein:
[0025] Binaural hearing characteristics are implemented through a binaural cross-correlation function model to evaluate the spatial perception quality of stereo audio;
[0026] The auditory fatigue model simulates the dynamic changes of auditory sensitivity after long-term listening through a time-varying gain control function.
[0027] Furthermore, the step of using physiological signal feedback to perform personalized fine-tuning includes:
[0028] Collecting electroencephalogram (EEG) signals and galvanic skin response (GSR) signals provided by the physiological signal acquisition device worn by the user;
[0029] Train the physiological signal feedback model to learn the mapping from physiological signals to audio perception states;
[0030] Based on the physiological signal feedback model, the audio quality characterization model is personalized and fine-tuned;
[0031] Continuously collect physiological signal feedback data from users during daily use and periodically update personalized model parameters.
[0032] Furthermore, the physiological signals also include heart rate variability, pupil diameter changes and eye movement trajectories;
[0033] The physiological signal feedback model adopts a multimodal fusion architecture, including a dedicated encoder for each signal type, and performs feature fusion through an attention mechanism.
[0034] Furthermore, the step of evaluating the quality of the played audio in real time and dynamically adjusting the sound effect parameters based on the evaluation results to perform adaptive sound effect calibration includes:
[0035] Split the audio stream into frames, perform feature extraction on each frame, and output quality scores and probability distribution of quality problem types;
[0036] Based on the quality assessment results, identify existing quality issues and determine the parameters that need to be adjusted through mapping relationships;
[0037] Dynamically adjust sound processing parameters, with the amount of parameter adjustment proportional to the confidence level of the problem;
[0038] Introducing smooth transition processing to avoid auditory discomfort caused by sudden parameter changes;
[0039] Perform real-time processing on the input audio stream and output optimized audio.
[0040] Furthermore, the steps of multi-dimensional quality problem diagnosis and fine-tuning include:
[0041] Build a dedicated quality decoding network to decode fine-grained quality information from quality feature representations;
[0042] Conduct refined quality problem diagnosis based on multi-dimensional quality characterization;
[0043] Determine optimization priorities based on the severity of quality issues and user physiological feedback;
[0044] Perform refined parameter optimization for high-priority quality issues;
[0045] Establish a feedback loop mechanism for optimization effects and dynamically adjust optimization strategies.
[0046] The present invention provides a deep learning-based sound effect optimization system for executing the above-mentioned deep learning-based sound effect optimization method, comprising:
[0047] A no-reference audio quality representation learning module for extracting quality-sensitive feature representations from audio signals;
[0048] The human auditory perception fusion module is used to optimize the audio quality characterization module so that its quality assessment results are closer to human subjective feelings;
[0049] Physiological signal feedback and personalized fine-tuning module to adapt to the user's individual hearing characteristics;
[0050] Real-time quality assessment and adaptive calibration module, used to evaluate the quality of played audio and dynamically adjust sound effect parameters;
[0051] The multi-dimensional quality problem diagnosis and fine-tuning module is used to optimize quality problems in different dimensions in a targeted manner.
[0052] The beneficial effects of the present invention are: through the self-supervised learning method, a quality assessment model that does not rely on reference audio is constructed, which can extract quality-sensitive feature representations from any audio signal. The system can identify common types of audio distortion, solving the limitation of traditional audio quality assessment relying on high-quality reference audio; by integrating the key perceptual characteristics of the human auditory system model, the quality assessment results are more in line with human subjective feelings, filling the gap between the traditional method based on signal level analysis and human subjective perception; using physiological signals as implicit feedback, the system can adapt to the individual auditory characteristics of different users, solve the problem of different users' perception of the same audio content, and realize personalized sound effect optimization; based on real-time quality assessment results, the system can dynamically adjust the sound effect parameters according to the specific quality problems identified, and improve the subjective listening quality of the damaged audio through adaptive calibration; through multi-dimensional quality problem diagnosis, the system can make fine adjustments to quality problems in different dimensions, avoid a "one-size-fits-all" optimization method, improve the overall listening experience, and form a closed loop with user physiological feedback to ensure the consistency of optimization and user perception. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 This is a flow chart of a method for optimizing sound effects based on deep learning according to the present invention;
[0054] Figure 2 is a flow chart of step 1 in the present invention;
[0055] Figure 3 It is a flow chart of step 2 in the present invention;
[0056] Figure 4 It is a flow chart of step 3 in the present invention;
[0057] Figure 5 is a flow chart of step 4 in the present invention;
[0058] Figure 6 It is a flow chart of step 5 in the present invention. DETAILED DESCRIPTION
[0059] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.
[0060] At least one embodiment of the present invention discloses a method for optimizing sound effects based on deep learning, such as Figures 1 to 6Shown, including:
[0061] Step 1: Build a reference-free audio quality characterization learning model to extract quality-sensitive feature representations from audio signals.
[0062] According to one embodiment of the present application, this step uses a self-supervised learning method to build a no-reference audio quality characterization model. This model can extract quality-sensitive feature representations from audio signals without relying on labeled data and can learn from a large amount of unlabeled audio. The specific implementation is as follows:
[0063] Step 1.1, audio signal preprocessing;
[0064] The input original audio signal is framed and windowed, and converted into a time-frequency representation (such as a spectrogram obtained by short-time Fourier transform (STFT)) as input for subsequent processing.
[0065] Step 1.2, construct a self-supervised learning task;
[0066] Based on the characteristics of audio quality assessment, the following self-supervised learning tasks are constructed:
[0067] Distortion prediction task: Apply multiple known types of distortion (such as compression distortion, noise interference, and clipping distortion) to the input audio, and train the model to predict the distortion type and severity from the distorted audio;
[0068] Feature reconstruction task: Randomly mask part of the time-frequency region of the input audio spectrum and train the model to reconstruct the masked region;
[0069] Optionally, in some implementations, the following auxiliary self-supervision tasks can also be constructed:
[0070] Context prediction task: Segment the audio into adjacent segments, train the model to predict the temporal relationship between the segments, and enhance the model's understanding of the temporal coherence of the audio.
[0071] Contrastive learning task: Different time-frequency transformations of the same audio are regarded as positive sample pairs, and transformations of different audio are regarded as negative sample pairs. The model is trained to learn invariant features related to perceptual quality.
[0072] Step 1.3, build the encoder network;
[0073] Convolutional Neural Networks (CNN) is used as the encoder backbone network, and the input is the spectrogram representation of the audio, and hierarchical features are extracted through multi-layer convolution and pooling operations. According to an embodiment of the present application, the CNN encoder includes 5 convolution blocks, each of which consists of a convolution layer, a batch normalization layer, a ReLU activation function, and a maximum pooling layer. The first two convolution blocks use 3×3 convolution kernels to extract local time-frequency features, and the last three convolution blocks use 5×5 and 7×7 convolution kernels to gradually expand the receptive field and capture a wider range of time-frequency patterns. The final feature map is converted into a vector of fixed dimension through a global average pooling layer. It should be understood that the mathematical expression of encoder E is:
[0074] z quality =E(X input );
[0075] Where E represents the encoder neural network, X input represents the spectrogram representation of the input audio, z quality A quality-sensitive feature representation vector representing the output.
[0076] According to another embodiment of the present application, the encoder network can also adopt a hybrid structure design, combining a convolutional network and a self-attention mechanism. In this embodiment, local features are first extracted through three convolutional blocks, and then two transformer encoder layers are introduced to process long-distance dependencies. Each transformer layer contains a multi-head self-attention module and a feedforward neural network. This design is particularly suitable for scenarios where it is necessary to capture long-term dependencies in audio, such as analyzing reverberation or spatial effect quality issues. The self-attention part can be expressed as:
[0077] H=MultiHeadAttention(Q,K,V);
[0078] Where H is the self-attention output feature, MultiHeadAttention(·) is the multi-head self-attention function, Q is the query matrix, K is the key matrix, and V is the value matrix. The query matrix, key matrix, and value matrix are all obtained by transforming the convolutional feature map.
[0079] For example, when processing high dynamic range audio content (such as symphonies), the hybrid structure can focus on both transient details and long-term dynamic changes, providing a more comprehensive quality assessment, and showing an accuracy improvement of about 15% compared to the pure convolutional structure in such scenarios.
[0080] Step 1.4, build the task head network;
[0081] For the self-supervised task constructed in step 1.2, build the corresponding task head network:
[0082] Distortion classification head: It consists of a 3-layer fully connected network with a structure of [512, 256, N classes,1 ], where N classes,1 is the number of predefined distortion types, ReLU activation function and Dropout (0.3) regularization are used between each layer, and the input is z quality , the output is the probability distribution of distortion type;
[0083] Distortion regression head: consists of a 2-layer fully connected network with a structure of [256, 1], using ReLU and linear activation functions, and the input is z quality , the output is an estimate of the degree of distortion;
[0084] Feature reconstruction head: It consists of 4 layers of transposed convolutional networks, which gradually upsample the feature vector to the original spectrogram size. Each layer uses a transposed convolution operation with a step size of 2, combined with batch normalization and ReLU activation function. The input is z quality , the output is the reconstructed spectrogram.
[0085] Step 1.5, model training and optimization;
[0086] The multi-task learning framework is used to jointly optimize the above tasks, and the loss function is:
[0087]
[0088] in is the total loss function of step 1, L class is the cross entropy loss for distortion classification, L reg is the mean square error loss of distortion degree regression, L recon is the reconstruction loss for feature reconstruction, α 1,1 , α 2,1 , α 3,1 Represent the weight coefficients of distortion classification loss, regression loss and reconstruction loss respectively.
[0089] Through the above steps, a trained no-reference audio quality representation model is obtained, in which the encoder E can extract quality-sensitive feature representation z from any audio signal. quality , which is used in subsequent steps. In addition, the model does not rely on labeled data and can be self-supervised learned from a large amount of unlabeled audio, with strong generalization ability.
[0090] Step 2: By integrating the human auditory perception model, the audio quality representation model is optimized using quality-sensitive feature representation;
[0091] According to another embodiment of the present application, this step incorporates the key perceptual characteristics of the human auditory system (HAS) into the audio quality characterization model constructed in step 1, so that the quality assessment results of the model are closer to human subjective perception. The specific implementation is as follows:
[0092] Step 2.1, construct a human auditory model;
[0093] Implement a computational model that incorporates the following key auditory perception features:
[0094] Auditory filter bank: Based on critical bandwidth theory, a set of filters is constructed to simulate the frequency selectivity of the human ear. According to the embodiment of this application, a Gammatone filter bank is used, which contains 24 filters covering the auditory range of 20Hz-20kHz, and the filter center frequencies are distributed according to the ERB (equivalent rectangular bandwidth) scale;
[0095] Loudness perception nonlinearity: realizes the nonlinear mapping relationship between loudness and sound pressure level. According to the embodiment of the present application, Stevens power law is used for modeling, which is expressed as:
[0096] L=k1·(P / P0) 0.6 ;
[0097] Where L is the perceived loudness, P is the sound pressure, P0 is the reference sound pressure, k1 is the proportionality coefficient, and 0.6 is the exponential parameter in Stevens' power law, which represents the nonlinear relationship between loudness and sound pressure;
[0098] Time-frequency masking effect: Achieve mutual masking effects of sounds in the time and frequency dimensions. According to the embodiments of the present application, frequency masking is achieved by calculating the masking threshold function of each frequency point, and time masking is achieved by forward and backward masking window functions. The masking threshold calculation formula is:
[0099]
[0100] Where T mask (f) is the masking threshold at frequency f, T abs (f) is the absolute hearing threshold at frequency f, represents the sum of all masking tones i1, is the intensity of the i1th masker, is the frequency f and the center frequency of the masking sound The masked spread function between .
[0101] Optionally, in another embodiment, the human auditory model may be extended to include the following additional features:
[0102] Binaural hearing: This feature simulates the spatial localization and separation capabilities of the human binaural auditory system. This feature is achieved by establishing an Interaural Cross Correlation (ICC) model to evaluate the spatial perception quality of stereo audio.
[0103] Auditory fatigue model: simulates the dynamic changes in auditory sensitivity after long-term listening. This model uses a time-varying gain control function:
[0104]
[0105] Where G(f, t1) is the gain of frequency f at time t1, G0(f) is the initial gain, α fatigue,1 is the fatigue coefficient, I(f, τ1) is the sound intensity of frequency f at time τ1, is an exponential decay function, τ r,1 is the recovery time constant, It represents the integration of time from 0 to t1.
[0106] According to another embodiment of the present application, a data-driven auditory perception model can be constructed by combining deep neural networks with traditional auditory models. This method first uses a large amount of subjective listening data to train a neural network to simulate human auditory perception characteristics. The output of the traditional auditory model is then used as a constraint to guide network learning. This method is particularly suitable for capturing complex auditory phenomena that are difficult to mathematically describe with traditional models, such as individual differences in sound quality preferences and timbre perception.
[0107] Step 2.2, perception-driven data enhancement;
[0108] In the self-supervised learning task of step 1.2, optimize the distortion introduction method:
[0109] For noise interference distortion, the noise intensity is determined by the signal masking threshold T mask (f, t2) is dynamically adjusted. It should be noted that the adjustment relationship can be expressed as:
[0110] N perceptual (f, t2) = N base (f, t2)·scale(T mask (f, t2));
[0111] where N perceptual (f, t2) is the noise intensity at frequency f and time t2 under perception driving, N base (f, t2) is the basic noise, T mask (f, t2) is the masking threshold at frequency f and time t2, and scale(·) is the scaling function that dynamically adjusts the noise intensity according to the masking threshold;
[0112] For the frequency distortion type, the degree of distortion is dynamically adjusted based on the difference in auditory sensitivity at different frequencies;
[0113] Step 2.3, perception-oriented loss function;
[0114] Replace L in step 1.5 recon Reconstruction loss, introducing a loss function based on auditory perception:
[0115]
[0116] Among them L perceptual is the perceptual loss, D spectral (·,·|HAS) is the spectral distance function calculated in the HAS transform domain of the human auditory system, X orig is the original spectrum, To reconstruct the spectrum graph; |HAS means calculation in the auditory model domain, | means calculation under certain conditions, and HAS represents the human auditory system.
[0117] Step 2.4, training the optimized and improved representation model;
[0118] Based on the improvements in steps 2.2 and 2.3, retrain the model in step 1, and the optimized loss function is:
[0119]
[0120] in is the improved total loss function, L class is the distortion classification loss, L reg is the distortion degree regression loss, L perceptual is the perceptual loss; α 1,2 , α 2,2 , α 3,2 They represent the weight coefficients of distortion classification loss, regression loss, and perceptual loss, respectively, and are used to balance the contribution of each loss to the total loss. The specific values can be determined through experimental tuning according to actual task requirements.
[0121] Through these steps, we develop an audio quality characterization model that incorporates the characteristics of human auditory perception, enabling it to capture audio quality characteristics that are more consistent with human subjective perception. Therefore, compared to traditional methods based solely on signal features, this model's evaluation results are more consistent with human subjective hearing, improving the accuracy and practicality of quality assessment.
[0122] Step 3: Based on the optimization results of the audio quality characterization model, personalized fine-tuning is performed using physiological signal feedback to adapt to the user's individual hearing characteristics;
[0123] According to another embodiment of the present application, this step introduces the user's physiological signals as implicit feedback to perform personalized fine-tuning on the model obtained in step 2 to adapt it to the individual auditory characteristics of a specific user. The specific implementation is as follows:
[0124] Step 3.1, physiological signal acquisition and preprocessing;
[0125] The user wears a simple physiological signal collection device to collect the following signals:
[0126] EEG signals: Collect brain waves from the frontal and temporal lobe areas;
[0127] Galvanic skin response signal: measures changes in skin conductivity;
[0128] It should be noted that the collected original physiological signals are preprocessed through bandpass filtering, artifact removal and other preprocessing steps to obtain the preprocessed physiological signal sequence S physio (t3).
[0129] According to another embodiment of the present application, different types of user physiological signals may be collected, including:
[0130] Heart rate variability: Photoplethysmography (PPG) or electrocardiogram (ECG) is used to measure the pattern of heart rate changes.
[0131] Pupillometry: The dynamic changes in pupil size are monitored using eye tracking equipment to reflect auditory cognitive load.
[0132] Eye tracking: When users simultaneously watch related visual content, their eye movement patterns can reflect the distribution of attention to the audio content;
[0133] In some embodiments, especially for portable device scenarios, non-contact signal acquisition methods can be used, such as implementing remote photoplethysmography (rPPG) through a front camera to extract heart rate information, or using a microphone array to analyze the user's breathing pattern as an alternative indicator of emotional feedback, which greatly reduces the complexity and invasiveness of the system.
[0134] Step 3.2, constructing a physiological signal feedback model;
[0135] Train a small neural network f feedback , learn from S physio (t3) Mapping to audio perception state:
[0136] P state =f feedback (Sphysio (t3));
[0137] Among them, P state is the probability distribution of the perception state (such as “good perception”, “beginning to get tired”, “loss of details”, etc.), f feedback (·) is the neural network mapping function from physiological signals to sensory states, S physio (t3) is the input physiological signal feature sequence.
[0138] Alternatively, in another embodiment, a multimodal fusion architecture can be used to process different types of physiological signals. This architecture includes a dedicated encoder for each signal type (electroencephalogram signal, galvanic skin response signal, heart rate variability, etc.), each encoder independently extracts signal features, and then performs feature fusion through an attention mechanism:
[0139]
[0140] Among them F fused is the fused multimodal feature, For the i signal,1 The characteristics of the signal type, For the i signal,1 The attention weight of the signal, For all signal types i signal,1 Summation.
[0141] In some implementations, to address the issue of insufficient initial user data, a meta-learning approach can be used to train the feedback model. This approach first learns how to learn from multiple user data points, enabling the model to quickly adapt to new users from minimal individual data. Specifically, this approach utilizes the Model-Agnostic Meta-Learning (MAML) algorithm, achieving over 75% accuracy in perceptual state recognition after just 5 to 10 feedback samples, significantly reducing the burden of user labeling during system initialization.
[0142] Step 3.3, personalized fine-tuning method;
[0143] Based on the physiological signal feedback model, the audio quality characterization model obtained in step 2 is personalized fine-tuned:
[0144] When the user listens to the audio, the system simultaneously records the audio features and the quality representation of the model output. perceptual , and the user's physiological signal sequence S physio (t3);
[0145] Compute the loss of consistency between the model evaluation results and the physiological signal feedback:
[0146] L feedback =D KL (M SSL (x1;θ SSL,1 ), f feedback (S physio ));
[0147] Among them L feedback is the consistency loss, D KL (·,·) is the Kullback-Leibler divergence function, M SSL (x1;θ SSL,1 ) is the self-supervised learning model M SSL For input x1 and parameter θ SSL,1 The output probability distribution, f feedback (S physio ) is the output probability distribution of the physiological signal feedback model.
[0148] Construct the total loss function for fine-tuning:
[0149] L finetune =L SSL +λ balance,1 ·L feedback ;
[0150] Among them L finetune is the total loss in the fine-tuning stage, L SSL is the self-supervision loss, λ balance,1 is the balance coefficient (hyperparameter, controlling the impact of consistency loss on total loss), L feedback is the consistency loss.
[0151] Only update the parameters of the model's terminal layers, keeping the front-end feature extraction layers fixed to avoid overfitting;
[0152] Step 3.4, continuous learning and adaptation;
[0153] The system continuously collects physiological signal feedback data during daily use and periodically performs the following operations:
[0154] Aggregate accumulated personalized data;
[0155] Perform the fine-tuning process in step 3.3 based on the new data;
[0156] Update personalized model parameters;
[0157] Through these steps, a personalized audio quality assessment model tailored to the specific user's auditory characteristics is obtained, enabling sound optimization to better align with the individual user's actual auditory state. Furthermore, this method uses objective physiological signals as feedback, avoiding the burden of frequent subjective evaluations and improving the efficiency and accuracy of personalized adaptation.
[0158] Step 4: Use the personalized fine-tuned audio quality characterization model to evaluate the quality of the played audio in real time, and dynamically adjust the sound effect parameters based on the evaluation results to perform adaptive sound effect calibration;
[0159] According to another embodiment of the present application, this step utilizes the personalized audio quality assessment model constructed in the previous step to evaluate the quality of the played audio in real time, and dynamically adjusts the sound effect parameters accordingly to perform adaptive calibration. The specific implementation is as follows:
[0160] Step 4.1, real-time audio quality assessment;
[0161] Real-time quality assessment of playing audio:
[0162] Split the audio stream into frames (e.g. 20-50ms per frame);
[0163] Perform feature extraction on each frame of audio;
[0164] Input the extracted features into the personalized quality assessment model obtained in step 3;
[0165] Output the quality score Q of the audio frame score And the probability distribution of quality problem type P issue .
[0166] Step 4.2, quality problem identification and mapping;
[0167] Based on the quality assessment results, identify existing quality issues:
[0168] Set the problem identification threshold τ issue , when the probability P of a certain type of problem issue (i3)>τ issue When , it is determined that such a problem exists;
[0169] where τ issue is the probability threshold for problem identification, P issue (i3) is the probability of the i3th type of problem.
[0170] Mapping relationship between construction quality issues and sound effect parameters M issue→param ,For example:
[0171] Insufficient low frequency response → low frequency gain adjustment;
[0172] High frequency harshness → high frequency reduction and smoothing;
[0173] Excessive dynamic range compression → Adjust expander parameters;
[0174] Too much reverberation → adjust the reverberation intensity;
[0175] Among them, M issue→paramis a mapping function from quality problem type to sound effect parameter, and the arrow “→” represents the mapping relationship.
[0176] Step 4.3, adaptive adjustment of sound effect parameters;
[0177] Dynamically adjust audio processing parameters based on identified quality issues:
[0178] For each identified quality problem i3, through the mapping relationship M issue→param Determine the set of parameters that need to be adjusted
[0179] The amount of parameter adjustment is proportional to the confidence level of the problem and can be expressed as:
[0180]
[0181] in is the parameter adjustment corresponding to the i3 type problem, Adjust the scaling factor for the parameters of the i3-th problem, P issue (i3) is the probability of the i3th type of problem, τ issue is the recognition threshold.
[0182] Combining the parameter adjustment results of multiple problems, we can get the final parameter adjustment vector ΔP total,1 ;
[0183] where ΔP total,1 The resultant vector of adjustments to all problem parameters.
[0184] According to another embodiment of the present application, an adaptive adjustment method based on user historical preferences can be adopted. This method maintains a user preference model U pref , records the user's historical response to different types of adjustments. The parameter adjustment amount can be expressed as:
[0185]
[0186] Among them U pref (i3) is the user's sensitivity coefficient to the i3th type of problem; is the parameter adjustment corresponding to the i3 type problem; Adjust the scaling factor for the parameters of type i3 problem; P issue (i3) is the probability of the i3th type of problem; τ issue is the recognition threshold.
[0187] In some embodiments, parameter adjustment can use a gradient descent optimization method, with the quality score as the objective function and the parameter change as the optimization variable:
[0188]
[0189] Where ΔP1 is the parameter adjustment amount, η1 is the learning rate (hyperparameter), is the gradient of the parameter P1.
[0190] Step 4.4, smooth transition processing;
[0191] To avoid auditory discomfort caused by sudden parameter changes, a smooth transition process is introduced:
[0192] Set the parameter change rate limit r max,1 ;
[0193] When the calculated parameter changes exceed the limit, clipping is performed:
[0194] ΔP limited,1 =clip(ΔP total,1 , -r max,1 , r max,1 );
[0195] where ΔP limited,1 is the parameter adjustment amount after clipping, clip(·, -r max,1 , r max,1 ) is a clipping function that limits the input to (-r max,1 , r max,1 ) interval, -r max,1 and r max,1 are the lower and upper limits of parameter changes respectively; ΔP total,1 The resultant vector representing all problem parameter adjustments.
[0196] Exponential smoothing filtering is introduced to slow down the rate of parameter change. It should be noted that this process can be expressed as:
[0197]
[0198] in is the parameter value at the current moment, is the parameter value at the previous moment, ΔP limited,1 is the parameter adjustment after limiting, α smooth,1 is the smoothing coefficient (0<α smooth,1 <1), 1-α smooth,1 is the compensation coefficient.
[0199] Step 4.5, real-time sound processing;
[0200] Based on adaptively adjusted parameters, the audio stream is processed in real time:
[0201] Construct a sound processing pipeline that includes multiple stages of processing, including an equalizer, a dynamics processor, a spatial effects processor, etc. According to an embodiment of the present application, the processing pipeline specifically includes:
[0202] Parametric equalizer: A 10-band filter bank covers the audible frequency range of 20Hz-20kHz, with each filter independently adjustable for gain, Q, and center frequency.
[0203] Multi-band dynamics processor: divides the audio into three frequency bands (low, mid, and high) and independently compresses or expands each band. Adjustable parameters include threshold, ratio, attack time, and release time.
[0204] Stereo enhancement processor: Enhances the width and surround of the stereo sound field by adjusting the phase and amplitude relationship between channels;
[0205] Harmonic Generator: Add appropriate harmonic components through nonlinear processing to enhance the sound texture;
[0206] According to the parameters obtained in step 4.4 Update the parameters of each processor in the processing pipeline in real time. According to an embodiment of the present application, parameter mapping is achieved through a predefined conversion matrix, which maps the quality assessment dimensions to specific processor parameters, such as:
[0207]
[0208] in is the parameter vector at the current moment, T1 is the transformation matrix, is the current quality assessment result vector.
[0209] The input audio stream is processed and the optimized audio is output. According to the embodiment of the present application, the processing adopts a block processing method with a block size of 1024 samples. Adjacent blocks are smoothly connected using a 50% overlapping cosine window to ensure that no auditory discontinuity or distortion occurs during the processing.
[0210] Through these steps, adaptive sound calibration based on real-time quality assessment is achieved, enabling targeted optimization of specific quality issues across different audio content. Furthermore, this method utilizes smooth transitions to avoid the discomfort associated with sudden parameter changes, enhancing the continuity and naturalness of the user experience.
[0211] Step 5: Combine the real-time evaluation of the audio quality and the dynamic adjustment of sound effect parameters based on the evaluation results to diagnose and fine-tune multi-dimensional quality issues, and perform targeted optimization for quality issues in different dimensions.
[0212] This step further enhances the system's ability to diagnose quality issues, enabling multi-dimensional, refined quality assessment and sound optimization. Specific implementation is as follows:
[0213] Step 5.1, multi-dimensional quality representation decoding;
[0214] Build a dedicated quality decoding network, and characterize the quality features z obtained in step 3 quality Decode fine-grained quality information:
[0215] Construct a linear probe network set: It consists of multiple simple linear layers, each linear probe is responsible for quality According to the embodiment of the present application, each linear probe consists of two layers of fully connected networks. The first layer converts z quality Mapping to a 128-dimensional intermediate representation, using ReLU activation, the second layer maps the intermediate representation to the score value of the corresponding quality dimension, using different activation functions depending on the scoring type (e.g., linear activation for regression tasks and sigmoid activation for classification tasks);
[0216] The quality dimensions covered include:
[0217] Frequency balance (low frequency, mid frequency, high frequency);
[0218] Dynamic characteristics (dynamic range, compression ratio);
[0219] Spatial characteristics (sound field width, sense of envelopment);
[0220] Time domain characteristics (transient clarity, temporal resolution);
[0221] Training method: Freeze the parameters of the quality characterization model and only train the linear probe network so that it can quality Accurately extract the quality information of the corresponding dimension. According to the embodiment of the present application, a multi-stage training strategy is adopted: first, the model infrastructure is trained on an audio quality dataset with expert annotations, and then fine-tuned on data of specific application scenarios to enable the model to adapt to the specific quality assessment requirements of the target environment;
[0222] In order to improve the generalization ability, data enhancement and regularization techniques are introduced during the training process, including adding Gaussian noise to z quality vector, randomly mask some feature dimensions, and use Dropout (0.2) to prevent overfitting.
[0223] Step 5.2: Refine the diagnosis of quality issues;
[0224] Based on multi-dimensional quality characterization, more detailed quality problem diagnosis can be performed:
[0225] For each quality dimension d1, score it With standard threshold Compare;
[0226] When the score of a dimension deviates from the threshold and exceeds the set range, it is determined that there is a quality problem in the dimension:
[0227]
[0228] in is the score of the d1th quality dimension, is the standard threshold of the d1th quality dimension, is the absolute difference between the score and the threshold, is the allowable deviation of the d1th dimension.
[0229] Count the quality problems in each dimension and form a quality problem diagnosis report quality,1 ;
[0230] Among them D quality,1 Provides a multi-dimensional quality problem diagnosis report.
[0231] Step 5.3, priority strategy formulation;
[0232] Determine optimization priorities based on the severity of quality issues and user physiological feedback:
[0233] Calculate the severity of each quality issue:
[0234]
[0235] in is the severity of the d1 quality problem, is the weight coefficient of the d1th dimension, is the absolute difference between the score and the threshold, is the allowable deviation of the d1th dimension.
[0236] Combined with the physiological signal feedback in step 3, adjust the weights of each dimension:
[0237]
[0238] in is the adjusted weight, is the original weight, sensitivityFactor(d1,P state ) is based on the current perception state P state and the sensitivity adjustment factor calculated for dimension d1.
[0239] Sort by severity after adjustment and determine optimization priority;
[0240] Optionally, in some implementations, the priority strategy can be dynamically adjusted based on the audio type and the user's current environment. The system identifies the audio type (e.g., speech, music, ambient sound, etc.) through content analysis and dynamically adjusts the priorities of different quality dimensions based on the ambient noise detection results. For example, when playing voice content in a noisy environment, the system will prioritize issues related to speech clarity; while when playing music in a quiet environment, it will focus more on optimizing dimensions such as dynamic range and spatial perception.
[0241] According to another embodiment of the present application, the priority strategy can be implemented through a multi-objective optimization framework, which regards multiple quality dimensions as goals that need to be met simultaneously:
[0242]
[0243] Among them, P * For the optimal parameter setting, It means finding the value of parameter P2 that minimizes the objective function. To sum over all mass dimensions d1, is the adjusted weight, is the loss function of parameter P2 for the d1th quality dimension.
[0244] Step 5.4, refine parameter optimization;
[0245] Perform refined parameter optimization for high-priority quality issues:
[0246] Expand the mapping relationship in step 4.2 to make it more fine-grained:
[0247]
[0248] Among them, M dim→param For quality diagnosis report D quality,1 A mapping function to a set of parameter target values, is the i4th parameter and its target value, and n1 represents the total number of parameter-target value pairs.
[0249] For multiple parameters that interact with each other, a multi-parameter joint optimization strategy is adopted to avoid optimization conflicts. According to the embodiment of the present application, this strategy is implemented by the Bayesian optimization algorithm. Each time multiple related parameters are adjusted and the comprehensive effect is observed, the optimal parameter combination is predicted by the Gaussian process regression model. The system establishes the correlation matrix C between parameters. param,1 , when adjusting the parameters When the relevant parameters are automatically calculated based on the correlation The co-adjustment amount:
[0250]
[0251] in is the coordinated adjustment amount of the j1th parameter, C param,1 (i4, j1) is the parameter and The correlation coefficient of is the adjustment amount of the i4th parameter.
[0252] Introducing parameter optimization constraints ensures that the optimization does not lead to degradation of other quality dimensions. According to an embodiment of the present application, the constraints are implemented by setting an objective function:
[0253]
[0254] Among them F opt,1 To optimize the objective function, For all target dimensions i dim,1 Sum, is the improvement weight of the target dimension, For the target dimension i dim,1 The improvement of penalty,1 is the penalty coefficient, To sum over all non-target dimensions j2, is the negative impact on the non-target dimension j2.
[0255] Step 5.5, feedback loop and continuous optimization;
[0256] Establish a feedback loop mechanism for optimization results:
[0257] Monitor changes in quality scores after parameter adjustments. According to an embodiment of the present application, the system uses a sliding window method to calculate the quality score comparison within 5 seconds before and after the adjustment, obtain the score change curve ΔQ(t5), and detect the effect of the adjustment;
[0258] Wherein ΔQ(t5) is the quality score change curve at time t5, Δ is the sign of the change, and Q(t5) is the quality score at time t5.
[0259] If the adjustment fails to effectively improve the problem or causes other problems, the strategy adjustment is performed. According to the embodiment of the present application, the system uses the reinforcement learning framework to establish a strategy adjustment model, which takes the current state, the parameter adjustment action taken, and the quality improvement obtained as a triplet. Storage, update the parameter adjustment strategy through the Q-learning algorithm, where is the state at time t5 (including the current quality score vector, user physiological feedback status, etc.), is the parameter adjustment action at time t5, is the reward at time t5 (i.e., quality improvement measure).
[0260] Maintain a historical record of parameter adjustment effects to optimize future parameter adjustment strategies. According to an embodiment of the present application, historical records are stored in a structured database containing four key fields: audio type, quality issue, parameter adjustment solution, and effect score. A mapping table is established from audio content features to the optimal parameter adjustment solution, enabling rapid retrieval and experience reuse.
[0261] Through the above steps, the system can perform multi-dimensional and refined diagnosis of audio quality issues, and based on the user's individual perception characteristics, perform targeted sound optimization, avoid "one-size-fits-all" optimization methods, and improve the overall listening experience.
[0262] A deep learning-based sound effect optimization system, used to execute the above-mentioned deep learning-based sound effect optimization method, comprising:
[0263] A no-reference audio quality representation learning module for extracting quality-sensitive feature representations from audio signals;
[0264] The human auditory perception fusion module is used to optimize the audio quality characterization module so that its quality assessment results are closer to human subjective feelings;
[0265] Physiological signal feedback and personalized fine-tuning module to adapt to the user's individual hearing characteristics;
[0266] Real-time quality assessment and adaptive calibration module, used to evaluate the quality of played audio and dynamically adjust sound effect parameters;
[0267] The multi-dimensional quality problem diagnosis and fine-tuning module is used to optimize quality problems in different dimensions in a targeted manner.
[0268] Here, the present invention provides an implementation example: an application scenario in high-end smart headphones;
[0269] The deep learning-based sound optimization method of this embodiment has been applied in a high-end smart headset product. The headset is equipped with biosensors that can collect user physiological signals (such as galvanic skin response) and has a built-in high-performance, low-power neural network processing chip that can perform lightweight deep learning model inference locally. The product connects to the user's smartphone via Bluetooth, and the accompanying application installed on the phone is responsible for more complex model training and personalized parameter updates.
[0270] These headphones are primarily targeted at music enthusiasts and audio professionals, who have high expectations for audio quality and distinct listening preferences. Traditional fixed sound presets cannot meet their personalized needs, while this approach dynamically optimizes the sound experience based on different audio content and individual hearing characteristics.
[0271] In practical applications, the following datasets and training strategies are used for training the no-reference audio quality characterization model:
[0272] First, we collected over 10,000 hours of multi-genre audio content from publicly available audio datasets (including the MUSDB18, LibriSpeech, and FMA datasets). These audios were categorized into five main categories: speech, classical music, pop music, jazz, and environmental sounds. Subsequently, we used audio processing tools to generate 12 different types of audio distortions with five different degrees of distortion, constructing approximately 600,000 hours of self-supervised learning training data.
[0273] Model training was divided into two phases: the first phase used only self-supervised pre-training and lasted 10 days on four NVIDIA V100 GPUs. The second phase incorporated human auditory model constraints and continued training for another 5 days. The final model achieved a distortion classification accuracy of 78.3% on five test sets, with a mean absolute error of 0.12 (normalized to a scale of 0-1) for distortion estimation.
[0274] To adapt to the resource limitations of mobile devices, the model was lightweighted through knowledge distillation and quantization technology. The number of parameters in the encoder part was reduced from the original 23.5M to 2.8M, and the inference delay was reduced to 12ms (on a mid-range mobile phone processor), meeting the needs of real-time processing.
[0275] In the actual application of this headset product, the implementation process of personalized adaptation is as follows:
[0276] First, the user undergoes a 15-minute auditory calibration process upon initial use. The system plays a series of specially designed audio clips, prompting the user to rate the audio quality and recording the corresponding galvanic skin response data. This data is used to initialize the physiological signal feedback model.
[0277] Subsequently, during daily use, the system continuously monitors the skin electrical response signal and records the user's active feedback (such as adjusting the volume, skipping songs, or manually adjusting the equalizer). This data is uploaded to the cloud server through the companion application every day, and incremental update training is performed. The updated personalized model parameters are synchronized to the headphones the next day.
[0278] After about two weeks of use, the system was able to accurately identify user sensitivity patterns to specific audio features. For example, for some users, the system identified that they were particularly sensitive to distortion in the 3-5kHz frequency band, and therefore prioritized addressing issues in this frequency band when optimizing sound effects. For users with a preference for low frequencies, the system retained more low-frequency energy during dynamic range compression.
[0279] Here's an example of the system's workflow when processing a piece of poorly compressed pop music:
[0280] The user plays a low-bitrate song (96kbps AAC format) downloaded from a streaming service through the headphones;
[0281] The system's real-time evaluation revealed the following major quality issues: severe loss of high-frequency (greater than 10kHz) information (problem confidence level 0.92), blurred transient details (problem confidence level 0.85), and narrow stereo sound field (problem confidence level 0.78);
[0282] Based on the user's personalized model, the system identifies that the user is highly sensitive to transient details (sensitivity coefficient 1.4) and relatively less sensitive to stereo sound field width (sensitivity coefficient 0.8), so transient issues are prioritized.
[0283] The system dynamically adjusts the attack time (shortened from the default 15ms to 5ms) and release time (extended from 150ms to 200ms) of the multi-band dynamics processor, and slightly enhances the harmonic content of the high frequency band to compensate for the high frequency loss;
[0284] Parameter adjustment adopts a smooth transition strategy, gradually reaching the target parameter value within 3 seconds to avoid auditory discomfort caused by sudden changes;
[0285] The system continuously monitors the user's galvanic skin response and finds that the heart rate variability index has improved, indicating that the user responds positively to the optimized audio.
[0286] This method has achieved good technical results in the actual application of this headphone product. The effectiveness of this method was verified by collecting objective measurement data and subjective evaluation data through a one-month usage test with 100 test users from different backgrounds.
[0287] The accuracy comparison results of no-reference audio quality assessment are shown in Table 1:
[0288] Table 1: Comparison of accuracy of no-reference audio quality assessment (consistency with professional equipment measurement results)
[0289]
[0290] The comparison results of personalized adaptation effects are shown in Table 2:
[0291] Table 2: Comparison of personalized adaptation effects (subjective mean opinion score (MOS), full score 5 points)
[0292]
[0293] The above data demonstrates that this method achieves an average 24.5% improvement in accuracy for no-reference audio quality assessment compared to traditional methods. In terms of personalized adaptation, compared to fixed preset optimization, initial performance improved by 0.29 points, and after two weeks of personalized learning, it increased by 0.76 points, approaching the highest subjective satisfaction score. This demonstrates that this method can effectively address core issues in audio optimization and deliver a positive user experience in real-world applications.
[0294] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.
Claims
1. A sound effect optimization method based on deep learning, characterized in that: include: Build a reference-free audio quality characterization learning model to extract quality-sensitive feature representations from audio signals; By integrating the human auditory perception model, the audio quality characterization model is optimized using quality-sensitive feature representation; Based on the optimization results of the audio quality characterization model, personalized fine-tuning is performed using physiological signal feedback to adapt to the user's individual hearing characteristics; The system uses a personalized, fine-tuned audio quality characterization model to evaluate the quality of the audio playback in real time, and dynamically adjusts sound effect parameters based on the evaluation results for adaptive sound effect calibration. By combining the real-time evaluation of the quality results of the played audio and the dynamic adjustment of the sound effect parameters based on the evaluation results, multi-dimensional quality problem diagnosis and fine-tuning are carried out, and targeted optimization is carried out for quality problems in different dimensions.
2. The method for optimizing sound effects based on deep learning according to claim 1, wherein: The step of constructing a no-reference audio quality characterization learning model includes: The input original audio signal is framed and windowed, and converted into time-frequency representation; Construct a self-supervised learning task to apply multiple known types of distortion to the input audio and train the model to predict the type and severity of distortion from the distorted audio; A convolutional neural network is used as the encoder backbone network to extract hierarchical features through multi-layer convolution and pooling operations; A multi-task learning framework is constructed to jointly optimize distortion classification, distortion degree regression, and feature reconstruction tasks.
3. The method for optimizing sound effects based on deep learning according to claim 2, wherein: The self-supervised learning task also includes a context prediction task and a contrastive learning task, wherein: The context prediction task segments the audio into adjacent segments and trains the model to predict the temporal relationship between the segments; The contrastive learning task considers different time-frequency transformations of the same audio as positive sample pairs, and transformations of different audio as negative sample pairs, and trains the model to learn invariant features related to perceptual quality.
4. The method for optimizing sound effects based on deep learning according to claim 1, wherein: The step of integrating the human auditory perception model comprises: Build a computational model of human hearing that includes auditory filter banks, loudness perception nonlinearities, and time-frequency masking effects; In self-supervised learning tasks, the distortion introduction method is dynamically adjusted according to the characteristics of human hearing; A loss function based on auditory perception is introduced to replace the conventional reconstruction loss.
5. The method for optimizing sound effects based on deep learning according to claim 4, wherein: The human auditory computational model also includes binaural hearing characteristics and auditory fatigue models, wherein: Binaural hearing characteristics are implemented through a binaural cross-correlation function model to evaluate the spatial perception quality of stereo audio; The auditory fatigue model simulates the dynamic changes of auditory sensitivity after long-term listening through a time-varying gain control function.
6. The method for optimizing sound effects based on deep learning according to claim 1, wherein: The step of using physiological signal feedback to perform personalized fine-tuning includes: Collecting electroencephalogram (EEG) signals and galvanic skin response (GSR) signals provided by the physiological signal acquisition device worn by the user; Train the physiological signal feedback model to learn the mapping from physiological signals to audio perception states; Based on the physiological signal feedback model, the audio quality characterization model is personalized and fine-tuned; Continuously collect physiological signal feedback data from users during daily use and periodically update personalized model parameters.
7. The method for optimizing sound effects based on deep learning according to claim 6, wherein: The physiological signals also include heart rate variability, pupil diameter changes and eye movement trajectories; The physiological signal feedback model adopts a multimodal fusion architecture, including a dedicated encoder for each signal type, and performs feature fusion through an attention mechanism.
8. The method for optimizing sound effects based on deep learning according to claim 1, wherein: The steps of evaluating the quality of the played audio in real time and dynamically adjusting the sound effect parameters based on the evaluation results to perform adaptive sound effect calibration include: Split the audio stream into frames, perform feature extraction on each frame, and output quality scores and probability distribution of quality problem types; Based on the quality assessment results, identify existing quality issues and determine the parameters that need to be adjusted through mapping relationships; Dynamically adjust sound processing parameters, with the amount of parameter adjustment proportional to the confidence level of the problem; Introducing smooth transition processing to avoid auditory discomfort caused by sudden parameter changes; Perform real-time processing on the input audio stream and output optimized audio.
9. The method for optimizing sound effects based on deep learning according to claim 1, wherein: The steps of multi-dimensional quality problem diagnosis and fine-tuning include: Build a dedicated quality decoding network to decode fine-grained quality information from quality feature representations; Conduct refined quality problem diagnosis based on multi-dimensional quality characterization; Determine optimization priorities based on the severity of quality issues and user physiological feedback; Perform refined parameter optimization for high-priority quality issues; Establish a feedback loop mechanism for optimization effects and dynamically adjust optimization strategies.
10. A sound effect optimization system based on deep learning, characterized in that: A method for optimizing sound effects based on deep learning, for executing any one of claims 1 to 9, comprising: A no-reference audio quality representation learning module for extracting quality-sensitive feature representations from audio signals; The human auditory perception fusion module is used to optimize the audio quality characterization module so that its quality assessment results are closer to human subjective feelings; Physiological signal feedback and personalized fine-tuning module to adapt to the user's individual hearing characteristics; Real-time quality assessment and adaptive calibration module, used to evaluate the quality of played audio and dynamically adjust sound effect parameters; The multi-dimensional quality problem diagnosis and fine-tuning module is used to optimize quality problems in different dimensions in a targeted manner.