A method for dynamic optimization of audio parameters based on discrete auditory responses
Patent Information
- Application Number
- CN202610887554.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-09-29
AI Technical Summary
这种方式主要针对听力损失补偿,无法优化用户的主观音质偏好(如喜欢更明亮或更温暖的声音)
本发明通过采用离散听觉响应采集方法,向用户呈现一组离散的听觉刺激样本并收集主观评分,将复杂的音质偏好问题转化为用户易于完成的评分任务,解决了现有技术中优化目标模糊、用户需要专业知识的问题,降低了用户的使用门槛。
Smart Images

Figure CN122842604A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio signal processing and personalized auditory optimization technology, specifically to a method for dynamic optimization of audio parameters based on discrete auditory response. Background Technology
[0002] With the increasing popularity of personal audio devices, users' demand for personalized sound quality is growing. Different users may have drastically different experiences with the same audio parameter settings due to differences in hearing characteristics, aesthetic preferences, and usage scenarios.
[0003] Traditional audio parameter adjustment methods mainly include the following: Manual equalizer adjustment: Users manually drag the gain sliders of each frequency band on a graphical equalizer to adjust the sound quality. This method requires users to have some audio expertise, and the adjustment process is cumbersome, making it difficult to find the optimal parameter combination. Preset sound effect modes: The device provides several preset sound effects (such as pop, classical, rock, etc.) for users to choose from. While this method is simple, preset modes cannot cover all users' personalized needs and lack flexibility. Automatic compensation based on hearing tests: By playing test sounds and recording the user's response (such as a hearing threshold test), a compensation curve is automatically generated. This method mainly targets hearing loss compensation and cannot optimize users' subjective sound quality preferences (such as preferring a brighter or warmer sound).
[0004] Existing technologies suffer from the following main problems: Ambiguous optimization objectives: Users' subjective auditory experiences are difficult to quantify using a single objective indicator, and traditional methods based on physical measurements cannot accurately reflect users' true preferences. Vast parameter space: The combination space of audio parameters (such as multi-band equalizer and compressor parameters) is enormous, making exhaustive search infeasible and requiring efficient optimization algorithms. Heavy user burden: Existing methods require users to perform numerous and repetitive listening tests and adjustments, easily leading to auditory fatigue and a poor user experience. Lack of adaptability: Once the optimization results are fixed, they cannot dynamically adjust to changes in user preferences or usage scenarios.
[0005] In view of this, we propose a dynamic optimization method for audio parameters based on discrete auditory response. Summary of the Invention
[0006] The purpose of this invention is to provide a method for dynamic optimization of audio parameters based on discrete auditory response, so as to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: A method for dynamic optimization of audio parameters based on discrete auditory response, the method comprising: S1. Through an interactive user interface, a set of discrete auditory stimulus samples are presented to the user. Each sample corresponds to a set of preset audio processing parameters. The user subjectively rates the auditory experience of each sample, forming a discrete auditory response dataset. S2. Associate the audio parameter vector corresponding to each sample in the discrete auditory response dataset with the user rating, and construct a continuous mapping model from the audio parameter space to the user's auditory perception rating through a nonlinear mapping function. S3. Based on the continuous mapping model, define a parameter optimization objective function. The objective function is used to quantify the expected user score under the current audio parameters, and introduce an exploration-utilization balance term to encourage the exploration of unrated parameter regions during the optimization process. S4. Using a sequential optimization algorithm, iteratively search for the parameter combination that maximizes the objective function value in the audio parameter space. After each iteration, present the audio sample corresponding to the newly generated parameter combination to the user for scoring, add the new scoring data to the discrete auditory response dataset, and update the continuous mapping model until the convergence condition is met. S5. Apply the final optimized audio parameters to the real-time audio processing link, perform corresponding parameterization processing on the input audio signal, and output the optimized audio signal.
[0008] Preferably, the preset audio processing parameters in step S1 include at least the gain values of each frequency band of the multi-band equalizer, the threshold and compression ratio of the dynamic range compressor, and the loudness normalization target level; the auditory stimulus sample is a standard audio segment processed by the audio processing parameters, and the standard audio segment is selected from a speech, music or environmental noise sample library.
[0009] Preferably, the nonlinear mapping function in step S2 adopts a Gaussian process regression model. This model takes the audio parameter vector as input, the user rating as output, measures the similarity between different points in the parameter space through a kernel function, and outputs the predicted rating and its uncertainty estimate.
[0010] Preferably, the Gaussian process regression model is used in a given observation dataset. Below, for a new parameter vector Its predicted score The calculation formula is: ; in, Let the covariance vector be the sum of the variances between the new point and all observed points. is the transpose operator, K is the covariance matrix between observation points, and y is the observation score vector; Forecast uncertainty The calculation formula is: ; in, This is the kernel function.
[0011] Preferably, the objective function for parameter optimization in step S3 is the upper confidence boundary acquisition function, and its calculation formula is as follows: ; in, The mean score predicted by the Gaussian process regression model. To predict the standard deviation, To explore and utilize the balance coefficient; this function takes larger values in the parameter range where the prediction score is high (utilization) and the prediction uncertainty is large (exploration), thus guiding the search process.
[0012] Preferably, the sequential optimization algorithm in step S4 is a Bayesian optimization algorithm. In each iteration, the next parameter point to be evaluated is selected by maximizing the acquisition function, the audio sample corresponding to the point is presented to the user for rating, the new data is added to the observation set and the Gaussian process regression model is updated, and this process is repeated until the preset maximum number of iterations is reached or the change in the acquisition function value is less than the convergence threshold.
[0013] Preferably, the method further includes: before formal optimization begins, selecting a set of initial sample points in the audio parameter space using Latin hypercube sampling or uniform design methods, having users rate the samples, and establishing an initial Gaussian process regression model to provide prior knowledge for subsequent sequential optimization.
[0014] Preferably, the method further includes: when the system serves a new user, using historical rating data of the existing user group as prior information, initializing the hyperparameters of the Gaussian process regression model through transfer learning or meta-learning, reducing the number of ratings required for the new user, and accelerating the personalized parameter optimization process.
[0015] Preferably, the method further includes: the user rating includes multiple dimensions, including at least sound quality clarity rating, auditory comfort rating, and naturalness rating; the parameter optimization objective function is extended to a multi-objective acquisition function, and multiple rating dimensions are optimized simultaneously through Pareto front method or weighted summation method.
[0016] Preferably, the method is implemented in an embedded manner in an audio playback device or mobile terminal, the interactive user interface is a touch screen or physical buttons, and the audio parameter optimization process is executed asynchronously in a background thread without affecting the normal playback of audio.
[0017] By employing the above technical solution, this invention provides a method for dynamic optimization of audio parameters based on discrete auditory response. It possesses at least the following beneficial effects: This invention presents a set of discrete auditory stimulus samples to users and collects subjective ratings by adopting a discrete auditory response acquisition method. It transforms the complex problem of sound quality preference into a rating task that is easy for users to complete, solving the problems of vague optimization goals and the need for professional knowledge in existing technologies, and lowering the user's usage threshold.
[0018] By using a Gaussian process regression model to construct a continuous mapping model from the audio parameter space to the user's auditory perception score, and outputting the predicted score and its uncertainty estimate, the problem of the large parameter space and inefficient modeling in the existing technology is solved, and a probabilistic surrogate model is provided for subsequent optimization.
[0019] By defining an objective function based on an upper confidence bound acquisition function and introducing an exploration-utilization balance term, the optimization process both favors selecting parameter regions with high prediction scores and encourages the exploration of parameter regions that have not yet been fully scored. This solves the problem of optimization easily getting trapped in local optima in existing technologies and improves global search efficiency.
[0020] By employing a Bayesian optimization algorithm in a sequential iteration, the model is updated by selecting the parameter points most likely to improve the score in each iteration for user evaluation. This approach approximates the globally optimal parameters with the fewest user interactions, solving the problems of heavy user burden and the need for extensive listening and speaking adjustments in existing technologies, and significantly reducing the number of times users need to score.
[0021] By applying the final optimized audio parameters to the real-time audio processing link, personalized automatic optimization of audio parameters is achieved, solving the problems of fixed optimization results and lack of adaptability in existing technologies. Users can enjoy sound quality that matches their preferences without manual adjustment.
[0022] By using an initial model warm-up step, initial sample points are selected through Latin hypercube sampling to establish an initial model before formal optimization. This solves the problems of lack of prior knowledge and low initial search efficiency in the existing technology during cold start, and accelerates the optimization convergence speed.
[0023] By using the personalized preference migration step to initialize the model for new users with historical data from the existing user group, the problem of repetitive work in existing technologies, which require each user to be optimized from scratch, is solved. This significantly reduces the number of ratings required for new users and improves system deployment efficiency.
[0024] By optimizing multiple scoring dimensions such as sound clarity, auditory comfort, and naturalness through multi-objective optimization extension steps, the problem of existing technologies failing to fully reflect user needs with a single scoring dimension is solved, enabling more refined personalized tuning.
[0025] By embedding the method into the audio playback device and executing the optimization process asynchronously in the background, the problem of the optimization process affecting normal audio playback in the prior art is solved, and a seamless personalized optimization experience is achieved. Attached Figure Description
[0026] The accompanying drawings, which are provided to further illustrate the invention, constitute a part of this application: Figure 1 This is a schematic diagram of the overall process of a dynamic optimization method for audio parameters based on discrete auditory response according to the present invention. Detailed Implementation
[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] Please see Figure 1 , Figure 1 This is a flowchart illustrating a dynamic optimization method for audio parameters based on discrete auditory response according to the present invention. The present invention provides a dynamic optimization method for audio parameters based on discrete auditory response, comprising: Step S1: Through an interactive user interface, a set of discrete auditory stimulus samples are presented to the user. Each sample corresponds to a set of preset audio processing parameters. The user subjectively rates the auditory experience of each sample, forming a discrete auditory response dataset. It's important to note that this step forms the foundation for data acquisition, its core being the transformation of users' "sound quality preferences," which are difficult to describe verbally, into quantifiable scoring data. The interactive user interface can be a mobile app, webpage, or embedded device screen, guiding users to complete a scoring task by playing audio clips processed with different parameters. Each auditory stimulus sample corresponds to a specific set of audio processing parameters (such as equalizer gain across frequency bands, compressor thresholds, etc.). Users simply need to intuitively score on a scale of 1-5 or 1-10, without needing to understand the underlying technical details. This design significantly lowers the barrier to entry for users, allowing even non-professional users to participate in the personalized optimization of audio parameters. The resulting discrete auditory response dataset forms the basis for all subsequent modeling and optimization.
[0029] Step S2: Associate the audio parameter vector corresponding to each sample in the discrete auditory response dataset with the user rating, and construct a continuous mapping model from the audio parameter space to the user's auditory perception rating through a nonlinear mapping function. It's important to note that this step involves modeling from discrete rating points to a continuous perceptual surface. Since users only rate a limited number of discrete samples, we cannot directly know their preferences for unrated parameter combinations. The Gaussian process regression model assumes that the output of any finite number of input points follows a joint Gaussian distribution and uses a kernel function to measure the similarity between different points in the parameter space. This allows it to predict the mean rating (expected preference) and uncertainty (predicted confidence) for any unrated parameter point. This continuous mapping model essentially constructs a user preference topography, enabling us to search for optimal solutions across the entire parameter space, rather than being limited to the finite number of rated samples.
[0030] Step S3: Based on the continuous mapping model, define a parameter optimization objective function. The objective function is used to quantify the expected user score under the current audio parameters, and introduce an exploration-utilization balance term to encourage the exploration of unrated parameter regions during the optimization process. It's important to note that this step defines the objective of searching for the optimal parameters. Focusing solely on "utilizing" already discovered high-scoring regions might miss the global optimum; focusing only on "exploring" unknown regions is inefficient. The upper confidence bound acquisition function cleverly balances these two aspects: it takes on larger values for parameter regions with high prediction scores (utilization) and high prediction uncertainty (exploration). By maximizing this function, the system can intelligently select the next most valuable evaluation point—which could be a point near a known good region or a region with high potential but still uncertain in the model. This balancing mechanism is key to the efficient convergence of Bayesian optimization.
[0031] Step S4: Using a sequential optimization algorithm, iteratively search for the parameter combination that maximizes the objective function value in the audio parameter space. After each iteration, present the audio sample corresponding to the newly generated parameter combination to the user for scoring, add the new scoring data to the discrete auditory response dataset, and update the continuous mapping model until the convergence condition is met. It's important to note that this step is the core iterative process of the method. The Bayesian optimization algorithm executes a closed loop of "selection-evaluation-update" at each step: first, it finds the maximum value of the acquisition function under the current model, which is then used as the next parameter combination to be evaluated; next, it generates corresponding audio samples and presents them to the user for rating; finally, it adds the new data to the observation set and updates the Gaussian process regression model. As the iterations proceed, the model's understanding of user preferences becomes increasingly accurate, prediction uncertainty gradually decreases, and the search region gradually converges to the parameter region where the user truly prefers it. Typically, only 10-20 iterations are needed to find satisfactory parameters, significantly improving efficiency compared to traditional manual adjustment methods.
[0032] Step S5: Apply the final optimized audio parameters to the real-time audio processing link, perform corresponding parameterization processing on the input audio signal, and output the optimized audio signal. It's important to note that this step is the practical application of the optimization results. After the iteration is complete, the system selects the parameter combination with the highest score (or the point with the largest acquisition function value) from all evaluated points as the final optimization result. This parameter is written to the configuration register of the audio processing link, and all subsequent audio signals passing through this link (whether music, speech, or game sound effects) will be automatically processed using this parameter. Users can enjoy sound quality that matches their preferences without any additional operation. If the user's preferences change, the optimization process can be restarted at any time, and the system will continue to optimize based on the existing model, rather than starting from scratch.
[0033] The preset audio processing parameters in step S1 include at least the gain values of each frequency band of the multi-band equalizer, the threshold and compression ratio of the dynamic range compressor, and the loudness normalization target level; the auditory stimulus sample is a standard audio segment processed by the audio processing parameters, and the standard audio segment is selected from a speech, music or environmental noise sample library. It's important to note that the selection of audio processing parameters covers the main dimensions affecting sound quality. The multi-band equalizer controls the loudness across different frequency ranges, influencing the brightness and warmth of the timbre; the dynamic range compressor controls the contrast between loud and soft sounds, affecting the fullness and impact of the sound; and the loudness normalization target level controls the overall volume. These parameters, combined, can produce a rich variety of sound quality variations. The selection of standard audio clips is also crucial: speech clips optimize vocal clarity, music clips optimize the aesthetics of the sound, and ambient noise clips optimize noise reduction. Users can choose appropriate test materials based on their primary usage scenarios.
[0034] The nonlinear mapping function mentioned in step S2 adopts a Gaussian process regression model. This model takes the audio parameter vector as input and the user rating as output. It measures the similarity between different points in the parameter space through a kernel function and outputs the predicted rating and its uncertainty estimate. It's important to note that Gaussian process regression is suitable for this task because it inherently handles small samples and nonlinear relationships, and it can output prediction uncertainty. The choice of kernel function is crucial: the squared exponential kernel assumes a smooth function, suitable for most audio quality preference scenarios; the Matérn kernel allows for more flexible function variations. The hyperparameters of the kernel function (such as length scale) determine how far apart two points in the parameter space the model considers "similar," and these hyperparameters can be automatically learned from the data by maximizing marginal likelihood. The estimation of prediction uncertainty is the basis for "exploration" in Bayesian optimization—regions with high uncertainty warrant further exploration.
[0035] The Gaussian process regression model is used in a given observation dataset. Below, for a new parameter vector Its predicted score The calculation formula is: ; in, Let the covariance vector be the sum of the variances between the new point and all observed points. is the transpose operator, K is the covariance matrix between observation points, and y is the observation score vector; Forecast uncertainty The calculation formula is: ; in, For kernel functions; It should be noted that the first formula calculates the predicted mean. It is actually a weighted average of the observed scores y, with the weights determined by the similarity between the new point and each observed point (via the covariance vector). Inverse of the covariance matrix (Decision) Determined. The more similar the observation point is to the new point, the greater its score's impact on the prediction. The second formula calculates the prediction variance. It equals the prior variance. Subtract the information from the observation data The more observation data there is, and the more similar it is to the new point, the greater the reduction in uncertainty. Together, these two formulas provide a complete probabilistic prediction for points with unknown parameters.
[0036] The objective function for parameter optimization in step S3 adopts the upper confidence boundary acquisition function, and its calculation formula is as follows: ; in, The mean score predicted by the Gaussian process regression model. To predict the standard deviation, To explore and utilize the balance coefficient; this function takes larger values in the parameter range where the prediction score is high (utilization) and the prediction uncertainty is large (exploration), thus guiding the search process; It should be noted that when When =0, the function degenerates into pure exploitation, always selecting the point with the highest current predicted score; when When the value is large, the function tends towards pure exploration, always choosing the point with the greatest uncertainty. Typically, it takes... =2 or 3, corresponding to an upper confidence bound of approximately 95% or 99.7%. The choice of this coefficient affects the convergence speed and the final result: If it's too small, it might get stuck in a local optimum. If the value is too large, convergence will be slow. In practical applications, an adaptive strategy can be adopted, using a larger value in the early stages of optimization. Encourage exploration, reduce later Strengthen utilization.
[0037] The sequential optimization algorithm described in step S4 is a Bayesian optimization algorithm. In each iteration, the next parameter point to be evaluated is selected by maximizing the acquisition function, the audio sample corresponding to the point is presented to the user for rating, the new data is added to the observation set and the Gaussian process regression model is updated, and this process is repeated until the preset maximum number of iterations is reached or the change in the acquisition function value is less than the convergence threshold. It should be noted that Bayesian optimization is an effective framework for solving black-box function optimization problems. In each iteration, maximizing the acquisition function is itself an optimization subproblem. Since the acquisition function is analytical and easily differentiated, it can be solved efficiently using algorithms such as gradient descent or L-BFGS. The computational complexity of updating the Gaussian process model after adding new data is O(N). 3 ), where N is the current number of observation points. Updates are fast when N is small, but the computational cost increases as N increases. In practical applications, the model usually stops after a certain number of iterations (e.g., 20-30), at which point it has converged to the user preference region. The convergence condition can be set as the maximum value of the acquisition function changing less than a threshold for several consecutive times, or the user rating not improving for several consecutive times.
[0038] The method further includes: before the formal optimization begins, selecting a set of initial sample points in the audio parameter space through Latin hypercube sampling or uniform design methods, having users score the data, and establishing an initial Gaussian process regression model to provide prior knowledge for subsequent sequential optimization; It's worth noting that the choice of initial samples significantly impacts the efficiency of Bayesian optimization. Uneven distribution of initial samples can lead to large blind spots in certain areas, resulting in inefficient subsequent searches. Latin hypercube sampling, a stratified sampling method, ensures uniform coverage of each parameter dimension, offering better space-filling than random sampling. Typically, 5-10 initial samples are sufficient to build a preliminary model. The scores of these initial samples also allow users to understand the testing process, preparing them for more refined scoring later.
[0039] The method further includes: when the system serves new users, using historical rating data of existing user groups as prior information, initializing the hyperparameters of the Gaussian process regression model through transfer learning or meta-learning, reducing the number of ratings required for new users, and accelerating the personalized parameter optimization process; It's worth noting that transfer learning is an effective means of solving the cold start problem. While each user's preferences differ, the human auditory system shares commonalities—for example, most people dislike harsh high frequencies or muddy low frequencies. By analyzing a large amount of historical user rating data, a "general preference prior" can be learned as the starting point for the new user model. Meta-learning goes a step further, learning how to quickly adapt to new users, allowing the model to adjust to individual preferences with only a small amount of new data. This method can reduce the number of ratings required for new users from 10-20 to 3-5, significantly improving the user experience.
[0040] The method further includes: the user rating includes multiple dimensions, including at least sound quality clarity rating, auditory comfort rating, and naturalness rating; the parameter optimization objective function is extended to a multi-objective acquisition function, and multiple rating dimensions are optimized simultaneously through Pareto front method or weighted summation method; It's worth noting that a single-dimensional rating may not fully reflect a user's true needs. For example, some users may prioritize clarity and be willing to sacrifice some comfort, while others may prioritize clarity at the expense of comfort. Multi-objective optimization allows the system to track multiple rating dimensions simultaneously and find a set of "non-dominated" compromises using the Pareto front method—that is, solutions where no single dimension can be further improved without harming other dimensions. Users can then choose the one that best suits their preferences from this set of solutions. The weighted summation method is even simpler and more direct; users can adjust the weights of each dimension using a slider, and the system automatically finds the parameter with the highest weighted total score.
[0041] The method is implemented in an embedded manner in an audio playback device or mobile terminal. The interactive user interface is a touch screen or physical buttons. The audio parameter optimization process is executed asynchronously in a background thread and does not affect the normal playback of audio. It's worth noting that the optimization process is executed asynchronously in a background thread to ensure uninterrupted or stuttering audio playback while the user is rating the track. The computational load for each iteration is minimal (primarily matrix operations of the Gaussian process model), and can be completed within a few hundred milliseconds on mainstream mobile chips. The user interface design is simple and intuitive: a loop of play-rating-next track, accompanied by a progress bar to indicate the optimization progress. Once optimization is complete, the device can automatically switch to the optimized parameter mode and provide a "comparison preview" function, allowing users to intuitively experience the differences before and after optimization.
[0042] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0043] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for dynamic optimization of audio parameters based on discrete auditory response, characterized in that, The method includes: S1. Through an interactive user interface, a set of discrete auditory stimulus samples are presented to the user. Each sample corresponds to a set of preset audio processing parameters. The user subjectively rates the auditory experience of each sample, forming a discrete auditory response dataset. S2. Associate the audio parameter vector corresponding to each sample in the discrete auditory response dataset with the user rating, and construct a continuous mapping model from the audio parameter space to the user's auditory perception rating through a nonlinear mapping function. S3. Based on the continuous mapping model, define a parameter optimization objective function. The objective function is used to quantify the expected user score under the current audio parameters, and introduce an exploration-utilization balance term to encourage the exploration of unrated parameter regions during the optimization process. S4. Using a sequential optimization algorithm, iteratively search for the parameter combination that maximizes the objective function value in the audio parameter space. After each iteration, present the audio sample corresponding to the newly generated parameter combination to the user for scoring, add the new scoring data to the discrete auditory response dataset, and update the continuous mapping model until the convergence condition is met. S5. Apply the final optimized audio parameters to the real-time audio processing link, perform corresponding parameterization processing on the input audio signal, and output the optimized audio signal.
2. The method for dynamic optimization of audio parameters based on discrete auditory response according to claim 1, characterized in that, The preset audio processing parameters in step S1 include at least the gain values of each frequency band of the multi-band equalizer, the threshold and compression ratio of the dynamic range compressor, and the loudness normalization target level; the auditory stimulus sample is a standard audio segment processed by the audio processing parameters, and the standard audio segment is selected from a speech, music or environmental noise sample library.
3. The method for dynamic optimization of audio parameters based on discrete auditory response according to claim 1, characterized in that, The nonlinear mapping function mentioned in step S2 adopts a Gaussian process regression model. This model takes the audio parameter vector as input and the user rating as output. It measures the similarity between different points in the parameter space through a kernel function and outputs the predicted rating and its uncertainty estimate.
4. The method for dynamic optimization of audio parameters based on discrete auditory response according to claim 1, characterized in that, The Gaussian process regression model is used in a given observation dataset. Now, for a new parameter vector Its predicted score The calculation formula is: ; in, Let the covariance vector be the sum of the variances between the new point and all observed points. is the transpose operator, K is the covariance matrix between observation points, and y is the observation score vector; Forecast uncertainty The calculation formula is: ; in, This is the kernel function.
5. The method for dynamic optimization of audio parameters based on discrete auditory response according to claim 1, characterized in that, The objective function for parameter optimization in step S3 adopts the upper confidence boundary acquisition function, and its calculation formula is as follows: ; in, The mean score predicted by the Gaussian process regression model. To predict the standard deviation, To explore and utilize the balance coefficient; this function takes larger values in the parameter range where the prediction score is high (utilization) and the prediction uncertainty is large (exploration), thus guiding the search process.
6. The method for dynamic optimization of audio parameters based on discrete auditory response according to claim 1, characterized in that, The sequential optimization algorithm described in step S4 is a Bayesian optimization algorithm. In each iteration, the next parameter point to be evaluated is selected by maximizing the acquisition function, the audio sample corresponding to the point is presented to the user for rating, the new data is added to the observation set and the Gaussian process regression model is updated, and this process is repeated until the preset maximum number of iterations is reached or the change in the acquisition function value is less than the convergence threshold.
7. The method for dynamic optimization of audio parameters based on discrete auditory response according to claim 1, characterized in that, The method further includes: before formal optimization begins, selecting a set of initial sample points in the audio parameter space using Latin hypercube sampling or uniform design methods, having users rate the samples, and establishing an initial Gaussian process regression model to provide prior knowledge for subsequent sequential optimization.
8. The method for dynamic optimization of audio parameters based on discrete auditory response according to claim 1, characterized in that, The method further includes: when the system serves new users, using historical rating data of existing user groups as prior information, initializing the hyperparameters of the Gaussian process regression model through transfer learning or meta-learning, reducing the number of ratings required for new users, and accelerating the personalized parameter optimization process.
9. The method for dynamic optimization of audio parameters based on discrete auditory response according to claim 1, characterized in that, The method further includes: the user rating includes multiple dimensions, including at least sound quality clarity rating, auditory comfort rating, and naturalness rating; the parameter optimization objective function is extended to a multi-objective acquisition function, and multiple rating dimensions are optimized simultaneously through Pareto front method or weighted summation method.
10. The method for dynamic optimization of audio parameters based on discrete auditory response according to claim 1, characterized in that, The method is implemented in an embedded manner in an audio playback device or mobile terminal. The interactive user interface is a touch screen or physical buttons. The audio parameter optimization process is executed asynchronously in a background thread and does not affect the normal playback of audio.