Sound box sound effect intelligent adjustment method and system based on data analysis, and storage medium
By building a three-dimensional fusion optimizer of user, environment and content features and combining it with deep learning technology, we have solved the problem that traditional sound adjustment methods cannot take into account user, environment and content features, achieved personalized sound adjustment and system optimization, and improved the sound experience and operating efficiency.
Patent Information
- Application Number
- CN202511202655.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-10-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional sound adjustment methods cannot simultaneously consider individual user differences, environmental acoustic characteristics, and audio content features, resulting in a poor sound experience.
Build user feature modeling, environmental feature analysis and content feature extraction engines, and generate personalized sound effect parameters through three-dimensional fusion optimizer and multi-objective optimization algorithm combined with deep learning technology.
It realizes personalized sound adjustment for different users, environments and content, improves the quality of auditory experience, simplifies the system structure, improves operational efficiency and maintainability, and continuously optimizes performance through enhanced learning mechanisms.
Smart Images

Figure CN120751310A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio signal processing, and more specifically, to a method, system and storage medium for intelligently adjusting sound effects based on data analysis. Background Art
[0002] With the rapid development of audio technology, people's requirements for sound quality are increasing, and they are no longer satisfied with simple volume adjustment and fixed sound effect presets. Traditional sound effect adjustment methods mainly rely on preset equalizer parameters or simple environmental mode switching. They fail to consider the complex interaction between three key dimensions: individual user differences, environmental acoustic characteristics, and audio content characteristics. As a result, the sound experience in practical applications is often unsatisfactory.
[0003] The main challenge facing current sound effect adjustment technology lies in achieving true intelligence and personalization in complex and ever-changing scenarios. On the one hand, different users have different auditory characteristics and preferences, and the same sound effect parameters can significantly affect the user experience. On the other hand, changes in environmental acoustic conditions (such as background noise and reverberation characteristics) can significantly affect the perceived quality of audio. Furthermore, different types of audio content (such as speech, music, and film) have different requirements for the optimal configuration of sound effect parameters. The complex interactions between these three dimensions make sound effect adjustment a high-dimensional, nonlinear optimization problem.
[0004] Therefore, a method for intelligently adjusting sound effects that comprehensively considers user characteristics, environmental characteristics, and content characteristics is needed. Through data analysis and machine learning, a unified model integrating these three factors can be established to automatically optimize sound effect parameters, providing users with the best auditory experience. This approach requires not only precise capture of feature expressions in each dimension but also a deep understanding of their interactions. Based on this understanding, an efficient optimization algorithm can be constructed to adapt to the needs of various complex scenarios. Summary of the Invention
[0005] The present invention provides a method, system and storage medium for intelligent adjustment of sound effects based on data analysis, which solves the technical problem in related technologies that the complex interactive relationship between the three key dimensions of user individual differences, environmental acoustic characteristics and audio content characteristics cannot be considered simultaneously, resulting in a poor sound experience.
[0006] The present invention provides a method for intelligently adjusting sound effects based on data analysis, comprising the following steps: Build a user feature modeling engine, extract user auditory feature data, and generate user feature vectors; Build an environmental feature analysis engine, collect environmental acoustic characteristic data, and generate environmental feature vectors; Build a content feature extraction engine to analyze audio content semantic data and generate content feature vectors; A three-dimensional fusion optimizer was constructed, which input the user feature vector, the environment feature vector, and the content feature vector into the auditory scene fusion model. The auditory experience score was calculated through tensor fusion operations, and a multi-objective optimization algorithm was applied to solve the optimal sound effect parameters. Build a parameter generation controller to convert the optimal sound effect parameters into specific audio processing parameters to achieve intelligent adjustment of sound effects.
[0007] In a preferred embodiment, the step of building a user feature modeling engine includes: Collect user explicit characteristic data, including age, gender, hearing test results, and known hearing status; Collect user implicit characteristic data, including volume adjustment habits, frequency response preferences, sound effect setting history, and other usage behavior data; Analyze user data using auditory psychology models, extract key feature parameters, and generate user auditory feature vectors; A deep neural network is applied to construct an individual response function, which describes the perceptual response of the user feature vector under a given sound effect parameter vector.
[0008] In a preferred embodiment, the step of building an environmental feature analysis engine includes: Deploy microphone arrays to collect environmental acoustic data, including ambient noise, reverberation characteristics, and sound field distribution; Execute the acoustic feature extraction algorithm, perform time-frequency analysis on the collected audio signals, and extract key acoustic parameters; Generate a three-dimensional sound field model using spatial sound field reconstruction technology; An adaptive filter is applied to construct an environmental response function, which describes the influence of the environmental feature vector on a given sound effect parameter vector.
[0009] In a preferred embodiment, the step of building a content feature extraction engine includes: Extract multi-scale features from the input audio, including spectral features, time domain features, prosodic features, and other underlying acoustic features; Use deep convolutional neural networks to extract semantic features of audio content; Use recurrent neural networks to analyze the temporal structure of audio and identify the rhythm, emotional changes, narrative structure, and other dynamic features of the audio content; A multi-task learning framework is applied to construct a content response function, which describes the performance quality of the content feature vector under a given sound effect parameter vector.
[0010] In a preferred embodiment, the step of constructing a three-dimensional fusion optimizer includes: Construct an auditory scene fusion model to represent the auditory experience as a three-dimensional interaction function; Implement tensor fusion operations. Tensor fusion operations are not simple tensor multiplications, but complex nonlinear mappings that consider the mutual influence between three dimensions. A multi-objective optimization algorithm is applied to solve the optimal parameters. The system goal is to find the best sound effect parameters to maximize the auditory experience score. The reinforcement learning mechanism is introduced to achieve long-term optimization. The system continuously adjusts and improves the parameters of the fusion model by recording user feedback.
[0011] In a preferred embodiment, the step of constructing a parameter generation controller includes: Design a parameter mapping network to map the optimized abstract parameter space to the parameter space of the actual audio processing system; Implement a smooth parameter transition algorithm to avoid auditory discomfort caused by sudden parameter changes; Build a multi-channel output control system to provide differentiated audio output for different locations in multi-user scenarios; Apply self-diagnosis and self-calibration mechanisms to continuously monitor the actual effects of parameter adjustments and promptly detect and correct deviations.
[0012] In a preferred embodiment, the steps for implementing the tensor fusion operation include: Calculate the initial output values of the three response functions; Construct a cross-dimensional attention network and calculate the interaction relationship matrix between different response dimensions; Generate interaction features between dimensions based on attention weights; The original response value is fused with the interaction feature and processed through residual connection and multi-layer perceptron network; The final fusion score is generated through weighted combination.
[0013] In a preferred embodiment, the reinforcement learning mechanism uses a deep Q network, including: Define the state space, which includes the current user model, environment model, content features, and historical interaction data; Define the action space, which is the adjustment direction and amplitude of the sound effect parameters; Design the network structure and adopt the double Q learning architecture to reduce the estimation bias; Design a reward function that is a weighted combination of user satisfaction score and system resource consumption; Experience replay and priority sampling techniques are used to enhance the learning effect of key samples.
[0014] In a preferred embodiment, a sound effect intelligent adjustment system based on data analysis is used to perform a sound effect intelligent adjustment method based on data analysis, including: User feature modeling engine, used to extract user auditory feature data and generate user feature vectors; Environmental feature analysis engine, used to collect environmental acoustic characteristic data and generate environmental feature vectors; Content feature extraction engine, used to analyze audio content semantic data and generate content feature vectors; A three-dimensional fusion optimizer, which inputs the user feature vector, the environment feature vector, and the content feature vector into the auditory scene fusion model, calculates the auditory experience score through tensor fusion operations, and applies a multi-objective optimization algorithm to solve for the optimal sound effect parameters; The parameter generation controller is used to convert the optimization results into specific sound effect parameters to achieve intelligent adjustment of the sound effects.
[0015] In a preferred embodiment, a computer-readable storage medium is used to store computer-readable instructions, which, when read by a computer, can run an intelligent sound effect adjustment system based on data analysis.
[0016] The beneficial effects of the present invention are: It breaks through the limitations of the traditional separate processing paradigm, establishes a theoretical framework for auditory scene fusion, realizes deep interaction of three-dimensional features through tensor fusion operations, and solves the mutual interference and conflict problems caused by independent optimization of each dimension in existing technologies. It enables the system to find the globally optimal sound effect parameter combination, greatly improving the quality of the auditory experience.
[0017] It cleverly balances the contradiction between universality and personalization. The individual response function, environmental response function and content response function constructed through deep learning can accurately capture the feature expression and response relationship of different dimensions. While maintaining the versatility of the system, it provides highly personalized sound experience for different users listening to different content in different environments, significantly improving user satisfaction.
[0018] The adoption of a unified three-dimensional fusion optimizer replaces the traditional multi-system collaborative architecture, significantly simplifying the system structure, reducing parameter redundancy and computational load, and improving system efficiency, enabling efficient operation on resource-constrained devices. This also enhances the system's maintainability and scalability, laying the foundation for widespread application of intelligent audio and sound effect adjustment technology. Furthermore, the reinforcement learning mechanism introduced in this invention enables the system to continuously learn and optimize from user feedback. As usage increases, system performance will continue to improve, providing users with increasingly accurate sound effect adjustment services. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a flow chart of a method for intelligently adjusting sound effects based on data analysis of the present invention; Figure 2 It is a bar chart comparing the performance improvements of the sound effect adjustment method of the present invention; Figure 3 is a bar graph showing the system complexity reduction effect of the present invention; Figure 4 It is a radar chart for the performance evaluation of the three-dimensional fusion optimizer of the present invention; Figure 5 is a line graph comparing the adaptability of the present invention under different scenarios; Figure 6 It is an area graph showing the long-term performance improvement trend of the present invention; Figure 7 It is a box plot comparing the consistency of satisfaction of multiple user scenarios of the present invention. DETAILED DESCRIPTION
[0020] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.
[0021] At least one embodiment of the present invention discloses a method for intelligently adjusting sound effects based on data analysis, such as Figure 1 As shown, the following steps are included: Step 1: Build a user feature modeling engine, extract user auditory feature data, and generate a user feature vector; It includes the following sub-steps: Step 1.1, collect user explicit feature data; This includes basic information such as age, gender, hearing test results, known hearing conditions, etc., which can be input by the user or obtained from the user profile.
[0022] It should be noted that, in some implementations, the system may also obtain richer user features through additional information sources such as the user's social media data or application usage habits.
[0023] Step 1.2, collect user implicit feature data; This includes usage behavior data such as volume adjustment habits, frequency response preferences, and sound effect setting history. These data are obtained by monitoring users' daily operations of the audio system.
[0024] Optionally, the system can also record users' emotional reactions and satisfaction ratings to further enrich the implicit feature dataset.
[0025] In step 1.3, the user data is analyzed using the auditory psychology model to extract key feature parameters, such as auditory sensitivity curve, masking threshold, dynamic range preference, etc., to generate the user auditory characteristic vector U.
[0026] It should be understood that the dimension of the vector can be flexibly adjusted according to the actual application scenario, and usually contains 50 to 200 eigenvalues to fully express the user's auditory characteristics.
[0027] Step 1.4: Apply deep neural network to construct individual response function , which describes the user feature vector Given a sound effect parameter vector The perceptual response below.
[0028] According to the embodiments of the present application, the network adopts a multi-layer perceptron structure, which is trained with a large amount of user feedback data and can accurately predict different users' subjective evaluations of various sound effect parameter combinations. The deep neural network specifically includes the following structure and implementation: The input layer receives two types of feature vectors: User auditory feature vector , including the user's hearing sensitivity curve (expressed by frequency segment), dynamic range preference, sound quality preference and other characteristics; Sound effect parameter vector , including equalizer parameters, dynamic processing parameters, space processing parameters and other sound effect adjustment related parameters.
[0029] The feature fusion layer performs a preliminary fusion of user features and sound effect parameters, employing an attention mechanism to identify the relative importance of different sound effect parameters for a specific user, and generates a fused feature representation. This layer contains 64 neurons and uses the ReLU activation function. In some embodiments, this layer may also employ adaptive normalization techniques to ensure consistent distribution of data across different users.
[0030] The hidden layer consists of three fully connected layers with 128, 256, and 128 neurons, respectively. Each layer is followed by batch normalization and dropout (ratio 0.3) to enhance model generalization. The hidden layer uses LeakyReLU as the activation function to avoid the vanishing gradient problem.
[0031] The output layer generates a user's perceptual response score for a given sound parameter, including multiple dimensions such as clarity score, comfort score, spatial sense score, and overall satisfaction score, each dimension is represented by a scalar.
[0032] During the training process, the mean square error was used as the loss function, the Adam optimizer was used, the initial learning rate was set to 0.001, and a learning rate decay mechanism was introduced.
[0033] To enhance the performance of the model in data-sparse areas, a transfer learning mechanism based on user similarity is introduced, allowing the model to draw on feedback data from similar users in the absence of specific user data.
[0034] like Figure 2 The data shows the percentage improvement in performance of three key indicators of the three-dimensional fusion method compared to the traditional separation processing method. The data shows that the method improved user satisfaction by 46%, sound clarity by 38%, and sound effect adjustment accuracy by 65%.
[0035] Step 2: Build an environmental feature analysis engine, collect environmental acoustic characteristic data, and generate an environmental feature vector; It includes the following sub-steps: Step 2.1: deploy a microphone array to collect ambient acoustic data; Including environmental noise, reverberation characteristics, sound field distribution, etc., multiple microphones can be distributed in different positions to obtain spatial acoustic information.
[0036] It should be noted that, in some embodiments, the number and layout of the microphone array can be optimized according to the application scenario, such as using more microphones in a large space, or adopting a higher-density layout in an environment with complex acoustic characteristics.
[0037] Step 2.2, executing the acoustic feature extraction algorithm; Perform time-frequency analysis on the collected audio signals to extract key acoustic parameters such as noise spectral density, reverberation time (RT60), acoustic transfer function (RTF), etc.
[0038] In addition, the system can also calculate parameters such as acoustic center frequency, clarity (C50), definition (D50), etc. to fully characterize the acoustic characteristics of the environment.
[0039] Step 2.3, generating a three-dimensional sound field model using spatial sound field reconstruction technology; The model accurately represents the propagation characteristics of sound waves in space, including physical phenomena such as reflection, diffraction, and absorption.
[0040] Optionally, the model can use beamforming techniques to enhance spatial resolution or employ microphone array self-calibration algorithms to improve positioning accuracy.
[0041] Step 2.4: Apply adaptive filter to construct environmental response function , which describes the environmental feature vector For a given sound effect parameter vector impact.
[0042] The function predicts the actual changes that the environment will cause to the audio signal by analyzing the modulation effects of the environment on different frequencies, dynamic ranges, and spatial characteristics.
[0043] It should be understood that the adaptive filter can be of various types such as a Kalman filter, a least mean square (LMS) filter, or a recursive least square (RLS) filter, and can be selected according to actual needs.
[0044] like Figure 3 The data shows that the unified theoretical framework and fusion optimization method of this application can reduce system complexity compared to traditional multi-system collaborative architectures. The data shows that this method reduces parameter redundancy by 70%, reduces computational load by 60%, and shortens system response time by 45%.
[0045] Step 3: Build a content feature extraction engine to analyze the audio content semantic data and generate a content feature vector; It includes the following sub-steps: In step 3.1, multi-scale feature extraction is performed on the input audio, including underlying acoustic features such as spectral features, time domain features, and prosodic features. These features are obtained through traditional signal processing methods such as short-time Fourier transform (STFT) and Mel-frequency cepstral coefficients (MFCC).
[0046] In addition, the system can also extract auxiliary features such as chromaticity features, harmonic noise ratio (HNR), zero crossing rate, etc. to enhance the ability to characterize audio content.
[0047] Step 3.2: Use a deep convolutional neural network (CNN) to extract semantic features of the audio content. According to one embodiment of the present application, the network is pre-trained with a large amount of labeled data and is capable of recognizing different types of audio content, such as speech, music, and ambient sound, and extracting their high-level semantic representations. The specific structure and implementation of this convolutional neural network are as follows: The input layer receives the audio Mel spectrogram and converts the audio signal into The two-dimensional feature map of is the number of time frames.
[0048] The feature extraction backbone consists of five convolutional blocks, each of which contains two convolutional layers, a batch normalization layer, a ReLU activation layer, and a max pooling layer.
[0049] The convolution kernel size is 3×3, 32 filters are used in the first layer, and then doubled to 512 filters in each subsequent layer.
[0050] The global context module is used to capture long-term dependencies. It uses one-dimensional convolution and global average pooling combined with an attention mechanism to ensure that the network can understand the long-range semantic associations of audio content.
[0051] The classification head contains a global average pooling layer and two fully connected layers with 256 and 128 neurons respectively, which are used to generate content type classification and feature representation.
[0052] The training adopts a two-stage strategy: first, pre-training is performed on large-scale audio datasets (such as AudioSet) to learn general audio features; then fine-tuning is performed on specific task data to optimize the feature extraction capabilities for sound effects.
[0053] In step 3.3, a recurrent neural network (RNN) is used to analyze the temporal structure of the audio. According to another embodiment of the present application, the network recognizes dynamic features such as rhythm, emotional changes, and narrative structure of the audio content, and generates a temporal semantic representation.
[0054] The specific structure and implementation of the recurrent neural network are as follows: The network adopts a bidirectional long short-term memory network (Bi-LSTM) structure, which consists of 3 layers with 256 hidden units in each layer, and can simultaneously consider the previous and next contextual information of the audio content.
[0055] The temporal attention mechanism is used to identify key time points in the audio stream, such as scene transitions and emotion changes, and assigns importance weights to each time point through soft attention calculation.
[0056] The sequence modeling layer uses a gated recurrent unit (GRU) with 128 units to model the long-term temporal characteristics of audio content, such as rhythmic patterns and emotional arcs.
[0057] The state aggregation layer integrates the temporal features into a fixed-dimensional representation vector, uses the self-attention mechanism to calculate the weighted average of features at different time points, and generates the final temporal semantic representation.
[0058] Step 3.4: Apply the multi-task learning framework to build the content response function , which describes the content feature vector Given a sound effect parameter vector The performance quality of the following.
[0059] This function takes into account multiple aspects such as content clarity, emotional expression, spatial positioning, etc., and optimizes content performance through joint optimization.
[0060] Optionally, in some implementations, the framework can also introduce knowledge distillation technology to migrate the knowledge of large pre-trained models to lightweight models to improve reasoning efficiency.
[0061] like Figure 4 The figure shows a comparison of the performance of our three-dimensional fusion method with traditional separation methods across five dimensions: user satisfaction, sound clarity, sound quality, computational efficiency, and adaptability. The chart clearly demonstrates the comprehensive advantages of our method across all dimensions, with particular strengths in adaptability and sound quality.
[0062] Step 4: Build a three-dimensional fusion optimizer, input the user feature vector, environment feature vector, and content feature vector into the auditory scene fusion model, calculate the auditory experience score through tensor fusion operations, and apply a multi-objective optimization algorithm to solve the optimal sound effect parameters; It includes the following sub-steps: Step 4.1: Build an auditory scene fusion model and represent the auditory experience as a three-dimensional interaction function: ; in, Score the listening experience to indicate the overall listening experience quality of the user under specific environment and content conditions; is a sound effect parameter vector, which includes adjustable audio processing parameters such as equalizer settings, dynamic range compression ratio, and spatial sound effect parameters; The user feature vector contains personal feature data such as user age, hearing status, and sound effect preferences; is the environmental feature vector, which includes acoustic environment characteristics such as ambient noise level, reverberation time, and sound field distribution; is the content feature vector, which contains content semantic information such as audio type, spectrum characteristics, and emotional characteristics; is the individual response function, which describes the perceptual response of a specific user to a given sound effect parameter; is the environmental response function, which describes the acoustic modulation effect of a specific environment on a given sound parameter; is the content response function, which describes the performance quality of specific content under given sound effect parameters; It is a tensor fusion operation and a nonlinear fusion method that considers the complex interactions between three dimensions.
[0063] Tensor fusion operations It is not a simple multiplication operation, but a nonlinear operation that considers the complex interrelationships between multi-dimensional features.
[0064] Specifically, this operation first performs weighted fusion on the feature representations of each dimension, then dynamically adjusts the importance of different dimensions through the attention mechanism, and finally learns the interaction patterns between dimensions through a multi-layer neural network to generate the final fusion score.
[0065] This fusion approach can effectively capture the nonlinear interactive relationship between the three dimensions of user, environment, and content, and achieve a more accurate overall evaluation than simple combination.
[0066] Step 4.2 implements the tensor fusion operation, which is not a simple tensor product, but a complex nonlinear mapping that considers the mutual influence between the three dimensions.
[0067] Through the attention mechanism, the system can dynamically adjust the relative importance of the three dimensions and flexibly balance the needs of user experience, environmental adaptation and content expression for different scenarios.
[0068] The specific implementation of the tensor fusion operation is as follows: First, calculate the initial output values of the three response functions: User Response Value: ; in, Indicates the user's perceived response value to a specific sound effect parameter; represents the individual response function; Represents the user feature vector containing personal characteristics such as user age, hearing condition, and sound effect preference; Represents a vector of adjustable audio processing parameters, including equalizer settings, dynamic range compression ratios, and spatial sound parameters.
[0069] Environmental response value: ; in, Indicates the response value of the environment to specific sound effect parameters; represents the environmental response function; Represents the environmental feature vector containing acoustic environment characteristics such as ambient noise level, reverberation time, and sound field distribution; Represents a vector of adjustable audio processing parameters, including equalizer settings, dynamic range compression ratios, and spatial sound parameters.
[0070] Content response value: ; in, Indicates the response value of the content to a specific sound effect parameter; Represents the content response function; Represents the content feature vector containing semantic information of the content such as audio type, spectral features, and emotional features; Represents a vector of adjustable audio processing parameters, including equalizer settings, dynamic range compression ratios, and spatial sound parameters.
[0071] A cross-dimensional attention network is constructed, which receives three response values and calculates the interaction matrix between them.
[0072] First, the system calculates an attention weight matrix for the user and environment response values. This is done by concatenating the user and environment response values, transforming them using a learnable weight matrix, and finally applying a softmax function to obtain normalized attention weights. Similarly, the system calculates attention weight matrices for the relationship between user and content response values, and between environment and content response values. These attention weights reflect the relative importance and interaction strength between different response dimensions.
[0073] Based on the calculated attention weights, the system generates interaction features between each dimension. Specifically, the system concatenates each pair of response values and multiplies them by the corresponding attention weight matrix to obtain a feature representation reflecting the interaction information between the two dimensions. In this way, the system generates user-environment interaction features, user-content interaction features, and environment-content interaction features.
[0074] The original response value is fused with the interaction feature and processed through residual connection and multi-layer perceptron network.
[0075] First, the user response value is added to the user-related interaction features (user environment and user content), and then a nonlinear transformation is performed through a multi-layer perceptron to obtain an enhanced user fusion representation.
[0076] The same processing is also applied to the environment and content dimensions, resulting in enhanced environment fusion representation and content fusion representation, respectively. This processing approach not only preserves the original response information but also integrates cross-dimensional interaction information, ensuring the integrity and richness of the information flow.
[0077] Finally, a fusion score is generated through weighted combination. The system first applies weight coefficients to each of the three fusion representations for linear combination, then concatenates the three representations and performs an overall nonlinear transformation through a multi-layer perceptron. Finally, the two results are combined to obtain the final auditory experience score.
[0078] The weight coefficients here are dynamically adjusted according to the current task type and user preferences, enabling the system to flexibly adapt to the optimization focus in different scenarios.
[0079] Step 4.3, apply the multi-objective optimization algorithm to solve the optimal parameters. The system goal is to find the best sound effect parameters , making the listening experience score maximize: ; in, The optimal sound parameter combination is the final goal of the system, including the optimal values of adjustable audio processing parameters such as equalizer settings, dynamic range compression ratio, and spatial sound parameters. Represents the sound effect parameter vector, which is a variable that can be adjusted by the system. It includes equalizer settings (such as low frequency, mid-frequency, and high frequency gain), dynamic range compression ratio, spatial sound effect parameters (such as reverberation depth and sound field width), and other adjustable audio processing parameters; Represents the user feature vector, which contains personal characteristic data such as user age, gender, hearing condition, and sound effect preference, and is used to describe the user's auditory characteristics and personalized needs; Represents the environmental feature vector, which includes acoustic environment characteristics such as ambient noise level, reverberation time, and sound field distribution. It is used to describe the impact of the current acoustic environment on audio playback. Represents the content feature vector, which contains content semantic information such as audio type, spectral characteristics, and emotional characteristics, and is used to describe the characteristics of the audio content being played; It represents the auditory experience scoring function, which is an evaluation function that comprehensively considers the three dimensions of user, environment, and content. Its output value represents the overall auditory experience quality of the user under specific environment and content conditions; It is a mathematical function that represents the independent variable that makes the function reach its maximum value.
[0080] This formula means that the system needs to find the combination of all possible sound effect parameters that can make the auditory experience score The set of parameters that reaches the maximum value To solve this optimization problem, the system first defines the objective function , which is the auditory experience scoring function defined above, which receives the sound effect parameter vector , user feature vector , environmental feature vector and content feature vector as input.
[0081] The system then searches for the optimal solution in the parameter space using an improved particle swarm optimization algorithm. This algorithm simulates the coordinated movement of a swarm of particles in the search space, with each particle representing a set of candidate sound effect parameters. These particles continuously adjust their positions based on their individual and group historical optimal positions, gradually approaching the global optimal solution.
[0082] The algorithm incorporates adaptive inertia weighting, allowing particles to explore a wider range in the early stages of the search, while gradually refining the search later. A dynamic convergence strategy also applies, adding appropriate perturbations as particles aggregate to avoid falling into local optima. Through iterative optimization, the system ultimately finds the sound parameter combination that maximizes the auditory experience score.
[0083] In addition, in some embodiments, alternative methods such as a multi-objective evolutionary algorithm (NSGA-II) or Bayesian optimization may be used to achieve better performance in specific scenarios.
[0084] In step 4.4, a reinforcement learning mechanism is introduced to achieve long-term optimization. The system records user feedback and continuously adjusts and improves the parameters of the fusion model, so that the system can continuously improve its performance as the usage time increases.
[0085] According to one embodiment of the present application, reinforcement learning uses a deep Q-network (DQN) to use user satisfaction as a reward signal to guide model optimization. The specific structure and implementation of the deep Q-network are as follows: The state space includes the current user model, environment model, content features, and historical interaction data, and the state vector dimension is 512.
[0086] The action space is defined as the adjustment direction and amplitude of the sound effect parameters, discretized into 256 possible operations.
[0087] The network structure consists of three fully connected layers, containing 256, 512 and 256 neurons respectively, and adopts a double Q learning architecture to reduce estimation bias.
[0088] The reward function is designed as a weighted combination of user satisfaction score and system resource consumption, encouraging maintaining system efficiency while improving user experience.
[0089] Experience replay and priority sampling techniques are used in the learning process to enhance the learning effect of key samples, while the target network mechanism is used to stabilize the training process.
[0090] like Figure 5 The data shows a comparison of user satisfaction scores between the proposed method and traditional methods in five different application scenarios (home environment, public place, in-vehicle environment, mobile device, and assisted hearing). The data shows that the proposed method performs well in all scenarios, with a particularly significant advantage in the complex and changing public places and in-vehicle environments.
[0091] Step 5: Build a parameter generation controller to convert the optimal sound effect parameters into specific audio processing parameters to achieve intelligent adjustment of sound effects; It includes the following sub-steps: Step 5.1 designs a parameter mapping network to map the optimized abstract parameter space to the parameter space of the actual audio processing system, including equalizer parameters, dynamic range processing parameters, spatial sound effect parameters, etc. It should be noted that in some embodiments, this network can adopt an autoencoder structure to ensure the continuity and smoothness of the mapping through the encoding and decoding process.
[0092] Step 5.2 implements a smooth parameter transition algorithm to avoid auditory discomfort caused by sudden parameter changes. The algorithm uses a dynamic time window and nonlinear interpolation method to ensure the continuity and naturalness of parameter changes.
[0093] Optionally, the system can also dynamically adjust the transition rate based on the content type, such as using faster parameter changes in music transitions and more gradual changes in speech content.
[0094] Step 5.3, build a multi-channel output control system. In a multi-user scenario, the system can provide differentiated audio output for different locations at the same time to meet the needs of different users while maintaining the coordination of the overall sound field.
[0095] In addition, the system also supports dynamic adjustment of sound field partitions, automatically adjusting the sound effect area division according to changes in user position.
[0096] In step 5.4, by applying the self-diagnosis and self-calibration mechanism, the system can continuously monitor the actual effect of parameter adjustment, detect and correct deviations in a timely manner, and ensure long-term stable operation.
[0097] When the environment, users, or content change significantly, the system automatically triggers a re-optimization process.
[0098] In addition, the system can also perform periodic calibration tests to ensure stability and accuracy during long-term operation.
[0099] like Figure 6 The data shows the performance trends of the proposed 3D fusion method and the static optimization system over long-term use. The data demonstrates that, thanks to the reinforcement learning mechanism, the proposed method continues to improve over time, with a 32% performance increase (from 72 to 95 points) after 12 months of use, while the performance of the static optimization system remains virtually unchanged.
[0100] like Figure 7 The data shows the distribution of user satisfaction between our approach and traditional methods in a multi-user sharing scenario. The data demonstrates that our approach not only achieves higher average satisfaction (approximately 90 points vs. approximately 60 points), but also exhibits a narrower range of fluctuations in user satisfaction, demonstrating that our approach effectively balances the needs of diverse users and improves the consistency of user satisfaction.
[0101] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.
Claims
1. A method for intelligently adjusting sound effects based on data analysis, characterized in that: The following steps are involved: Build a user feature modeling engine, extract user auditory feature data, and generate user feature vectors; Build an environmental feature analysis engine, collect environmental acoustic characteristic data, and generate environmental feature vectors; Build a content feature extraction engine to analyze audio content semantic data and generate content feature vectors; A three-dimensional fusion optimizer was constructed, which input the user feature vector, the environment feature vector, and the content feature vector into the auditory scene fusion model. The auditory experience score was calculated through tensor fusion operations, and a multi-objective optimization algorithm was applied to solve the optimal sound effect parameters. Build a parameter generation controller to convert the optimal sound effect parameters into specific audio processing parameters to achieve intelligent adjustment of sound effects.
2. The method for intelligently adjusting sound effects based on data analysis according to claim 1, characterized in that: The steps of constructing the user feature modeling engine include: Collect user explicit characteristic data, including age, gender, hearing test results, and known hearing status; Collect user implicit characteristic data, including volume adjustment habits, frequency response preferences, sound effect setting history, and other usage behavior data; Analyze user data using auditory psychology models, extract key feature parameters, and generate user auditory feature vectors; A deep neural network is applied to construct an individual response function, which describes the perceptual response of the user feature vector under a given sound effect parameter vector.
3. The method for intelligently adjusting sound effects based on data analysis according to claim 1, characterized in that: The steps of constructing the environmental feature analysis engine include: Deploy microphone arrays to collect environmental acoustic data, including ambient noise, reverberation characteristics, and sound field distribution; Execute the acoustic feature extraction algorithm, perform time-frequency analysis on the collected audio signals, and extract key acoustic parameters; Generate a three-dimensional sound field model using spatial sound field reconstruction technology; An adaptive filter is applied to construct an environmental response function, which describes the influence of the environmental feature vector on a given sound effect parameter vector.
4. The method for intelligently adjusting sound effects based on data analysis according to claim 1, characterized in that: The steps of constructing a content feature extraction engine include: Extract multi-scale features from the input audio, including spectral features, time domain features, prosodic features, and other underlying acoustic features; Use deep convolutional neural networks to extract semantic features of audio content; Use recurrent neural networks to analyze the temporal structure of audio and identify the rhythm, emotional changes, narrative structure, and other dynamic features of the audio content; A multi-task learning framework is applied to construct a content response function, which describes the performance quality of the content feature vector under a given sound effect parameter vector.
5. The method for intelligently adjusting sound effects based on data analysis according to claim 1, characterized in that: The steps of constructing a three-dimensional fusion optimizer include: Construct an auditory scene fusion model to represent the auditory experience as a three-dimensional interaction function; Implement tensor fusion operations. Tensor fusion operations are not simple tensor multiplications, but complex nonlinear mappings that consider the mutual influence between three dimensions. A multi-objective optimization algorithm is applied to solve the optimal parameters. The system goal is to find the best sound effect parameters to maximize the auditory experience score. The reinforcement learning mechanism is introduced to achieve long-term optimization. The system continuously adjusts and improves the parameters of the fusion model by recording user feedback.
6. The method for intelligently adjusting sound effects based on data analysis according to claim 1, characterized in that: The step of constructing a parameter generation controller includes: Design a parameter mapping network to map the optimized abstract parameter space to the parameter space of the actual audio processing system; Implement a smooth parameter transition algorithm to avoid auditory discomfort caused by sudden parameter changes; Build a multi-channel output control system to provide differentiated audio output for different locations in multi-user scenarios; Apply self-diagnosis and self-calibration mechanisms to continuously monitor the actual effects of parameter adjustments and promptly detect and correct deviations.
7. The method for intelligently adjusting sound effects based on data analysis according to claim 1, characterized in that: The steps to implement the tensor fusion operation include: Calculate the initial output values of the three response functions; Construct a cross-dimensional attention network and calculate the interaction relationship matrix between different response dimensions; Generate interaction features between dimensions based on attention weights; The original response value is fused with the interaction feature and processed through residual connection and multi-layer perceptron network; The final fusion score is generated through weighted combination.
8. The method for intelligently adjusting sound effects based on data analysis according to claim 5, characterized in that: The reinforcement learning mechanism uses a deep Q network, including: Define the state space, which includes the current user model, environment model, content features, and historical interaction data; Define the action space, which is the adjustment direction and amplitude of the sound effect parameters; Design the network structure and adopt the double Q learning architecture to reduce the estimation bias; Design a reward function that is a weighted combination of user satisfaction score and system resource consumption; Experience replay and priority sampling techniques are used to enhance the learning effect of key samples.
9. A sound and sound effect intelligent adjustment system based on data analysis, used to execute the sound and sound effect intelligent adjustment method based on data analysis according to any one of claims 1 to 8, characterized in that: include: User feature modeling engine, used to extract user auditory feature data and generate user feature vectors; Environmental feature analysis engine, used to collect environmental acoustic characteristic data and generate environmental feature vectors; Content feature extraction engine, used to analyze audio content semantic data and generate content feature vectors; A three-dimensional fusion optimizer, which inputs the user feature vector, the environment feature vector, and the content feature vector into the auditory scene fusion model, calculates the auditory experience score through tensor fusion operations, and applies a multi-objective optimization algorithm to solve for the optimal sound effect parameters; The parameter generation controller is used to convert the optimization results into specific sound effect parameters to achieve intelligent adjustment of the sound effects.
10. A computer-readable storage medium, characterized in that It is used to store computer-readable instructions, and when the computer-readable instructions are read by a computer, it can run the sound effect intelligent adjustment system based on data analysis as described in claim 9.
Citation Information
Cited By
Adaptive scene sound effect adjusting method and device and storage medium
CN121284478A
Adaptive scene sound effect adjustment method and device, and storage medium
CN121284478B
Sound effect enhancement processing method and system and storage medium
CN121438852A