Sound parameter determination method and system
By real-time perception of the audio usage environment and analyzing user auditory preferences and optimizing audio parameters, the problem of traditional audio systems being unable to cope with environmental changes and neglecting nonlinear perception of the human ear in real time, realizing personalized audio adjustment and better user experience.
Patent Information
- Application Number
- CN202510182753.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The audio parameter adjustment of traditional audio systems relies on manual adjustments, and cannot cope with environmental changes in real time, and ignores the nonlinear characteristics of the human ear's perception of sound, which cannot meet the auditory needs of different users.
By obtaining multimodal sensor data of the audio, sensing the audio usage environment in real time, combining user audio control data, analyzing user auditory preferences, and optimizing audio parameters to achieve personalized audio adjustment.
It realizes automatic adaptation and personalized adjustment of the audio system in complex environments, avoids subjective errors in manual adjustment and untimely response to environmental changes, and improves the performance and user experience of the audio.
Smart Images

Figure CN120050570A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sensor recognition, and particularly to a method and system for determining audio parameters. Background Art
[0002] With the continuous development of audio technology, audio systems have been widely used in various places, including home entertainment, professional audio, cinemas, and public broadcasting, etc. However, traditional audio systems often rely on manual adjustment and subjective judgment to determine audio parameters, such as volume, frequency response, echo control, etc. This method has certain limitations in practical applications and cannot fully consider environmental factors and individual differences in user needs.
[0003] In early audio adjustment methods, the parameters of audio systems were usually set through empirical adjustment and simple manual measurement. Technicians adjusted the audio settings according to the specifications of audio equipment, the results of auditory perception, and intuitive understanding of audio signals. For example, the adjustment of the equalizer (EQ) of audio signals, volume setting, and bass enhancement, etc., were all set manually. However, this traditional manual adjustment method has many problems. Firstly, it relies heavily on auditory perception, which may lead to large individual differences, different subjective evaluations of different listeners, and it is difficult to unify the standards. Secondly, manual settings cannot respond to environmental changes in real time, especially in places with complex acoustic environments. The performance of audio systems may be affected by external factors such as room layout, wall reflection, and air-conditioning noise.
[0004] With the development of digital signal processing (DSP) technology, more and more audio systems have begun to adopt algorithm-based automatic adjustment methods. These methods automatically adjust audio parameters by analyzing the spectral and time-domain characteristics of audio signals and changes in the acoustic environment. Although these methods have improved the deficiencies in traditional audio adjustment to a certain extent, there are still some problems. Existing audio parameter adjustment methods often ignore the non-linear characteristics of human ear's perception of sound, that is, the response differences of the human ear to different frequencies and different sound pressure levels. This makes a single adjustment method often unable to meet the auditory needs of different users. Summary of the Invention
[0005] Based on this, it is necessary for the present invention to provide a method and system for determining audio parameters to solve at least one of the above technical problems.
[0006] To achieve the above object, an audio parameter determination method includes the following steps:
[0007] Step S1: Obtain audio multi-modal sensing data, and perform audio usage environment perception based on the audio multi-modal sensing data to obtain audio usage environment data; perform multi-dimensional feature vector combination on the audio usage environment data to obtain a high-dimensional feature data set;
[0008] Step S2: Perform environmental pattern recognition on the high-dimensional feature dataset to obtain audio usage environment pattern data, and perform convolutional feature mapping of the audio parameter response pattern based on the audio usage environment pattern data to obtain an audio response feature map;
[0009] Step S3: Determine the audio parameter influence factors based on the audio response feature map to obtain the audio parameter influence factors; analyze the environmental dynamic parameter adjustment strategy according to the audio usage environment data and the audio parameter influence factors to obtain the environmental dynamic parameter adjustment strategy, and transmit it to the audio control platform to execute the parameter adjustment task;
[0010] Step S4: Obtain the user audio control data through the audio control platform, and perform user audio control mode recognition according to the user audio control data to obtain the user audio control mode data; obtain the real-time audio multi-modal sensing data, and perform user auditory preference analysis on the real-time audio multi-modal sensing data and the user audio control mode data to obtain the user auditory preference data;
[0011] Step S5: Optimize the user personalized audio adjustment feedback of the environmental dynamic parameter adjustment strategy according to the user auditory preference data to obtain an optimized feedback audio dynamic adjustment strategy, and transmit it to the audio control platform to execute the parameter adjustment task.
[0012] By obtaining the multi-modal sensing data of the sound system, the present invention can perceive the changes in the usage environment of the sound system in real time, including factors such as noise, reflection, and spatial layout. This process enables the sound system to automatically adapt to different acoustic environments, improving the performance of the sound system and avoiding the subjective errors in traditional manual adjustments and the problem of untimely response to environmental changes. By combining the multi-dimensional feature vectors of these environmental data, a high-dimensional data set can be constructed, from which useful information can be extracted, and then environmental pattern recognition can be carried out to identify the adaptive requirements of the sound system in different environments. This method not only avoids the limitations of single-parameter adjustment but also enables optimized adjustment under different environmental conditions. Based on the analysis of the sound response feature map and the sound parameter influence factors, the response patterns of various sound parameters under specific environmental conditions can be deeply understood, and thus a dynamic environmental parameter adjustment strategy can be formulated. Different from the traditional fixed parameter settings, this method can flexibly respond to the real-time changes in the environment, enabling the sound system to maintain the best performance in a changing environment. For example, parameters such as the frequency response, volume, and bass boost of the sound system can be adjusted in a timely manner according to the environmental noise and acoustic reflection characteristics, thus improving the auditory experience and avoiding problems such as over-adjustment or under-adjustment. The personalized needs of users are also fully considered. By obtaining the sound control data of users, the sound control modes of different users in different environments can be identified, which not only enables the sound system to respond to the user settings in real time but also adjusts the sound quality settings according to the user preferences. At the same time, by combining the real-time multi-modal sensing data and the user control modes, the auditory preferences of users can be analyzed and extracted, and then a more accurate personalized sound experience can be provided. By optimizing the feedback mechanism, the sound parameters can be continuously adjusted to ensure the best sound effect in a changing environment, thus meeting the auditory needs of different users. In summary, each link of this technical step helps to break through the limitations of traditional sound adjustment, and through automated and intelligent adjustment methods, it can respond to environmental changes and user needs in real time, thus achieving a more personalized and accurate sound adjustment effect.
[0013] Optionally, step S1 is specifically as follows:
[0014] Step S11: Obtain the multi-modal sensing data of the sound system, and perform data preprocessing on the multi-modal sensing data of the sound system to obtain the multi-modal sensing data of the sound system to be analyzed;
[0015] Step S12: Perform statistical analysis on the spatial distribution of sensors for the multi-modal sensing data of the sound system to be analyzed to obtain the spatial distribution data of the sound sensors, and perform sound sensing data fusion on the multi-modal sensing data of the sound system to be analyzed according to the spatial distribution data of the sound sensors to obtain the sound usage sensing network;
[0016] Step S13: Extract the sensing time-frequency features using the audio sensing network to obtain the audio usage sensing time-domain feature data and the audio usage sensing frequency-domain feature data;
[0017] Step S14: Perform environmental perception based on the audio usage sensing time-domain feature data and the audio usage sensing frequency-domain feature data to obtain the audio usage environment data;
[0018] Step S15: Combine the multi-dimensional feature vectors of the audio usage environment data to obtain a high-dimensional feature data set.
[0019] Through the acquisition and processing of multi-modal sensing data, the present invention effectively improves the adaptability and personalized adjustment ability of the audio system in complex environments. After obtaining the multi-modal sensing data of the audio and performing preprocessing, the solution ensures the high quality and accuracy of the data, providing a reliable basis for subsequent analysis. By statistically analyzing the spatial distribution data of the audio sensors, the solution can understand the layout and working status of the sensors in the environment, thus laying an important foundation for data fusion. During the data fusion process, the construction of the audio usage sensing network enables the system to integrate data from different sensors in real time, comprehensively reflecting the changes in the audio environment and improving the accuracy of environmental perception. Through time-frequency feature extraction, the solution can extract the time-domain and frequency-domain feature data of the audio usage environment from different dimensions, providing rich reference bases for subsequent environmental perception and parameter adjustment. This process can not only analyze the dynamic changes of the environment but also identify the complex relationships between the audio system and the environment, thus ensuring the optimized performance of the audio system in a changing environment. Finally, the solution obtains a high-dimensional feature data set through the multi-dimensional feature combination of the audio usage environment data, providing strong data support for the further intelligent adjustment and personalized optimization of the system.
[0020] Optionally, step S14 is specifically:
[0021] Perform envelope analysis of the sensing signal based on the audio usage sensing time-domain feature data to obtain the sensing signal waveform time-series data;
[0022] Perform statistical analysis of the sensing signal frequency distribution according to the audio usage sensing frequency-domain feature data to obtain the sensing signal frequency distribution data;
[0023] Perform wavelet transform time-frequency analysis on the sensing signal waveform time-series data and the sensing signal frequency distribution data to obtain the sensing signal time-frequency feature data set;
[0024] Perform environmental noise pattern recognition according to the sensing signal time-frequency feature data set to obtain the environmental noise pattern data set;
[0025] Extract the audio time-frequency feature according to the time-frequency feature dataset of the sensing signal, so as to obtain the time-frequency feature data of the audio signal;
[0026] Denoise the audio signal time-frequency feature data according to the environmental noise pattern dataset, so as to obtain the time-frequency data of the audio signal to be analyzed, and analyze the acoustic wave reflection characteristics of the time-frequency data of the audio signal to be analyzed, so as to obtain the environmental acoustic wave reflection characteristic data;
[0027] Integrate the environmental dynamic changes based on the time-frequency feature dataset of the sensing signal, so as to obtain the environmental dynamic change data;
[0028] Perform multi-dimensional environmental perception data fusion on the environmental noise pattern dataset, the environmental dynamic change data, and the environmental acoustic wave reflection characteristic data, so as to obtain the audio usage environment data.
[0029] Through in-depth analysis of the time-domain and frequency-domain feature data of the audio usage, the present invention significantly improves the environmental adaptability and sound quality optimization ability of the audio system. The envelope analysis of the sensing signal can effectively extract the timing change characteristics in the audio signal, ensuring that the system can perceive the tiny changes in the environment and respond in a timely manner. This analysis method provides a basis for subsequent time-frequency feature extraction, enabling the audio system to more accurately capture the multi-dimensional characteristics of environmental noise and audio signals in a complex environment. The frequency distribution statistics of the frequency-domain feature data further help the system identify the frequency characteristics of the signal, thereby optimizing the signal processing method. The introduction of wavelet transform time-frequency analysis can organically combine the time-domain and frequency-domain information, and the obtained time-frequency feature dataset provides key support for identifying and processing environmental noise patterns and optimizing audio signals. Through environmental noise pattern recognition, the solution can accurately separate noise from useful signals, providing an effective basis for subsequent audio signal denoising, and ensuring that the audio system can still maintain high-quality sound effects in a noisy environment. In addition, through acoustic wave reflection characteristic analysis, the reflection characteristics of acoustic waves in different environments can be identified, so as to better adjust the output of the audio device and optimize the sound quality. The integrated analysis of environmental dynamic changes can reflect the changes in the environment in real time, enabling the audio system to still maintain stable performance when facing a dynamic environment. Finally, the integration of multi-dimensional environmental perception data enhances the audio system's comprehensive understanding of the environment, providing comprehensive data support for personalized adjustment and dynamic optimization.
[0030] Optionally, step S2 is specifically as follows:
[0031] Step S21: Perform principal component feature dimensionality reduction on the high-dimensional feature dataset to obtain the principal component feature dataset;
[0032] Step S22: Classify the audio behavior according to the principal component feature dataset to obtain the audio behavior data;
[0033] Step S23: Perform environmental acoustic pattern recognition on the audio behavior data based on the dynamically changing environmental data, so as to obtain the audio usage environment pattern data;
[0034] Step S24: Obtain the audio rated parameter set, and perform parameter combinations of the audio usage mode on the audio rated parameter set, so as to obtain the audio usage mode parameter group;
[0035] Step S25: Perform audio response mode mapping according to the audio usage environment pattern data and the audio usage mode parameter group, so as to obtain the audio response feature map.
[0036] Through principal component feature dimensionality reduction, the present invention reduces the redundant information in high-dimensional data, enabling the audio system to efficiently process complex environmental features with fewer computing resources while maintaining the integrity of important information. Next, the audio behavior classification technology helps the system identify different audio usage modes, enabling it to automatically adjust parameters according to different application scenarios (such as home entertainment, professional audio, etc.), thus ensuring the best sound effect output. Combining the dynamically changing environmental data, performing environmental acoustic pattern recognition on the audio behavior data helps to accurately understand the continuously changing acoustic features in the environment, thus providing real-time reference for the adjustment of audio equipment. The acquisition of the audio rated parameter set and its parameter combination with the usage mode effectively ensure that the audio equipment can be accurately adjusted according to the preset performance standards and requirements in different environments, improving the stability and consistency of the audio. Finally, through the response mode mapping of the audio usage environment pattern data and the usage mode parameter group, the system can accurately generate the audio response feature map and dynamically optimize the audio parameters based on this feature map. The whole process endows the audio system with a powerful environmental perception ability, enabling it to achieve automated and personalized audio optimization under various environmental conditions, thus enhancing the user's auditory experience.
[0037] Optionally, step S25 is specifically:
[0038] Step S251: Perform magnitude standardization on the audio usage environment pattern data and the audio usage mode parameter group, so as to obtain the standardized audio usage environment pattern data and the standardized audio usage mode parameter group;
[0039] Step S252: Perform audio usage mode feature selection on the standardized audio usage mode parameter group, so as to obtain the audio usage mode feature data;
[0040] Step S253: Perform convolutional layer audio response mode feature mapping according to the standardized audio usage environment pattern data and the audio usage mode feature data, so as to obtain the audio response mode feature map;
[0041] Step S254: Optimize the acoustic response feature map of the pooling layer for the acoustic response mode feature map to obtain the acoustic response feature map.
[0042] In the present invention, by performing magnitude standardization on the acoustic usage environment mode data and the acoustic usage mode parameter set, the magnitude differences between different data dimensions can be eliminated, ensuring that all data is processed on the same scale, thereby improving the accuracy of subsequent analysis and calculation. This step effectively reduces the noise interference caused by inconsistent data ranges, enabling the audio system to more accurately understand environmental changes and usage patterns, thus providing a more stable basis for optimizing audio parameters. Feature selection on the standardized acoustic usage mode parameter set helps to select feature data closely related to the audio performance, thereby reducing unnecessary data processing, improving the operation efficiency of the system, and avoiding overfitting problems, ensuring that the audio system can respond to the most valuable features. Through the convolutional layer, feature mapping of the acoustic response mode is performed, enabling the system to effectively capture the complex non-linear relationship between audio parameters and the environment, making the audio system more accurately adapt to different environmental conditions and enhancing its adaptive adjustment ability. Optimizing the acoustic response feature map by the pooling layer can reduce redundant information and further highlight important features, thereby improving the response accuracy and stability of the audio system.
[0043] Optionally, step S3 is specifically as follows:
[0044] Step S31: Extract the acoustic parameter feature weights from the acoustic response feature map to obtain the acoustic parameter weight data set;
[0045] Step S32: Perform normalization of the parameter magnitude differences according to the acoustic parameter weight data set to obtain the acoustic parameter influence data set;
[0046] Step S33: Perform environmental acoustic behavior analysis according to the acoustic parameter influence data set to obtain the acoustic behavior mode data;
[0047] Step S34: Perform regression analysis on the acoustic behavior mode data and the acoustic usage environment mode data to obtain the acoustic parameter influence factor;
[0048] Step S35: Select the environmental condition acoustic parameter combination according to the acoustic parameter influence factor and the acoustic usage environment data to obtain the environmental condition acoustic parameter combination data set;
[0049] Step S36: Develop a dynamic parameter adjustment strategy for the environmental condition acoustic parameter combination data set to obtain the environmental dynamic parameter adjustment strategy and transmit it to the audio control platform to execute the parameter adjustment task.
[0050] Through the extraction of the acoustic parameter feature weights from the acoustic response feature map, the present invention can identify the specific contributions of acoustic parameters to the system performance, thereby forming a clear acoustic parameter weight data set, laying a foundation for subsequent precise adjustment. The normalization of the parameter order of magnitude ensures that different parameters are analyzed on the same scale, eliminates the bias caused by the order of magnitude difference, improves the adjustment accuracy of the audio system, and ensures that each parameter can reasonably affect the system behavior. The analysis of the environmental acoustic behavior of the acoustic parameter influence data set enables the system to identify and understand the acoustic behavior patterns in different environments, further optimize the acoustic performance to adapt to complex environmental conditions. This process can effectively reveal the influencing factors of acoustic parameters in a specific environment through regression analysis of the acoustic behavior pattern data and the acoustic usage environment pattern data, thereby providing a theoretical basis for refined adjustment. Through the analysis of these influencing factors, the audio system can more accurately select the best combination of acoustic parameters to ensure the best auditory experience under different environmental conditions. By formulating a dynamic parameter adjustment strategy, the audio system can respond to environmental changes and user needs in real time, ensure that the acoustic performance is always in the best state, further enhance the user experience, and execute this adjustment task through the audio control platform to achieve the adaptive adjustment function of the system.
[0051] Optionally, step S35 is specifically as follows:
[0052] Step S351: Quantify the response difference of the acoustic behavior pattern data for the acoustic behavior pattern to obtain the acoustic behavior pattern response error data;
[0053] Step S352: Correlate the response error of the acoustic parameter influencing factor according to the acoustic behavior pattern response error data to obtain the response error-influencing factor correlation data;
[0054] Step S353: Construct an acoustic parameter response influence model according to the response error-influencing factor correlation data;
[0055] Step S354: Construct an acoustic usage environment parameter space based on the acoustic usage environment data;
[0056] Step S355: Select the parameter combination with the lowest response error for the environmental conditions in the acoustic usage environment parameter space through the acoustic parameter response influence model to obtain the environmental condition acoustic parameter combination data set.
[0057] The present invention quantifies the response differences of the audio behavior pattern data for the audio behavior pattern, which can effectively identify and quantify the response errors of the audio system under different environments or different settings. This process provides important error data for subsequent optimization, enabling the audio system to make targeted adjustments. Based on the correlation analysis between these response error data and the audio parameter influence factors, it can be further revealed which parameters have a greater impact on the errors, thereby providing a key basis for parameter adjustment for the fine-tuning of the audio system. This analysis provides data support for constructing an audio parameter response influence model, and through the model, it is possible to more accurately predict the impact of audio parameters on the system performance, thereby optimizing the overall performance of the system. By constructing an audio usage environment parameter space based on the audio usage environment data, it is possible to more comprehensively understand and analyze the impact of different environmental factors on the audio effect, providing a more accurate parameter range for dynamic adjustment. On this basis, by selecting the parameter combination with the minimum response error for the audio usage environment parameter space through the audio parameter response influence model, it can help the audio system select the optimal audio parameter combination under different environmental conditions, thereby maximizing the system performance and ensuring that the system can provide the best sound effects and user experience in various complex environments.
[0058] Optionally, step S4 is specifically as follows:
[0059] Step S41: Obtain user audio control data through the audio control platform and perform data preprocessing on the user audio control data to obtain the user audio control data to be analyzed;
[0060] Step S42: Extract the time-frequency characteristics of the user audio control for the user audio control data to be analyzed to obtain the time-frequency characteristic data of the user audio control, and perform user audio control mode recognition based on the time-frequency characteristic data of the user audio control to obtain the user audio control mode data;
[0061] Step S43: Obtain real-time multimodal sensing data and perform data preprocessing on the real-time multimodal sensing data to obtain the real-time multimodal sensing data to be analyzed;
[0062] Step S44: Perform real-time audio usage environment perception based on the real-time multimodal sensing data to be analyzed to obtain the real-time audio usage environment data;
[0063] Step S45: Perform user auditory preference analysis based on the real-time audio usage environment data and the user audio control mode data to obtain the user auditory preference data.
[0064] Through obtaining user audio control data from the audio control platform and performing data preprocessing on it, the present invention can remove noise and ensure the accuracy of the data, laying a solid foundation for subsequent analysis. Next, by extracting time-frequency features from these data, it helps to understand more meticulously the user's control methods and preferences for the audio system, and then identify different control modes. This provides data support for personalized audio adjustment, enabling the audio system to adapt to the specific needs of each user. Obtaining and preprocessing real-time multi-modal sensing data helps to capture the dynamic changes of the environment while ensuring the efficient utilization of the data. Further perception of the audio usage environment can evaluate the current environmental conditions in real time, such as noise level, spatial layout, etc., to ensure that the audio system automatically adjusts according to environmental changes and enhances the user experience. By combining the user audio control mode data and the real-time environmental data for user auditory preference analysis, it can accurately grasp the auditory needs of different users, thereby providing personalized adjustment feedback for the audio system.
[0065] Optionally, step S45 is specifically as follows:
[0066] Step S451: Extract the time-frequency features of the audio signal from the real-time audio usage environment data, thereby obtaining the real-time audio signal time-frequency feature data;
[0067] Step S452: Extract the time-frequency features of the audio control mode for the user audio control mode data, thereby obtaining the user audio control mode time-frequency feature data;
[0068] Step S453: Perform multi-modal neural network feature fusion based on the real-time audio signal time-frequency feature data and the user audio control mode time-frequency feature data, thereby obtaining the user behavior-environment audio feature space;
[0069] Step S454: Build a user auditory fitness model based on the user behavior-environment audio feature space, thereby obtaining the user auditory fitness evaluation model;
[0070] Step S455: Evaluate the user auditory fitness for the real-time audio usage environment data and the user audio control data to be analyzed through the user auditory fitness evaluation model, thereby obtaining the audio environment user auditory fitness data;
[0071] Step S456: Map the user auditory preferences according to the audio environment user auditory fitness data, thereby obtaining the user auditory preference data.
[0072] The present invention extracts the time-frequency characteristics of audio signals from the real-time audio usage environment data, which can capture in detail the time-domain and frequency-domain characteristics of the audio signals, providing accurate environmental audio data for subsequent audio adjustments. At the same time, extracting the time-frequency characteristic data of the user's audio control mode helps to analyze the user's dynamic response to audio control and provides a deep understanding of the user's control habits. After combining these data, through multi-modal neural network feature fusion, a feature space that comprehensively considers user behavior and environmental audio characteristics can be constructed, thus more accurately capturing the interaction effect between the user and the environment. Establishing a user auditory fitness model based on this feature space helps to quantify the user's adaptation degree to the audio environment and provides a personalized adjustment basis for the audio system. Through this evaluation model, the audio control data and environmental data of the user can be comprehensively analyzed in real time to obtain the user's auditory fitness, thereby guiding the audio system to make more accurate responses. Using the user auditory fitness data for user auditory preference mapping of environmental conditions can provide the user with personalized audio adjustments that better match their auditory preferences, thereby optimizing the user's auditory experience and ensuring that the audio system can provide the best sound effects in various environments.
[0073] Optionally, this specification also provides an audio parameter determination system for performing the audio parameter determination method described above. The audio parameter determination system includes:
[0074] An audio usage environment perception module for obtaining multi-modal sensing data of the audio, and performing audio usage environment perception based on the multi-modal sensing data of the audio to obtain audio usage environment data; performing a multi-dimensional feature vector combination on the audio usage environment data to obtain a high-dimensional feature data set;
[0075] An environmental mode recognition module for performing environmental mode recognition on the high-dimensional feature data set to obtain audio usage environmental mode data, and performing convolutional feature mapping of the audio parameter response mode based on the audio usage environmental mode data to obtain an audio response feature map;
[0076] An audio parameter adjustment module for determining the audio parameter influence factor based on the audio response feature map to obtain the audio parameter influence factor; performing an analysis of the environmental dynamic parameter adjustment strategy based on the audio usage environment data and the audio parameter influence factor to obtain the environmental dynamic parameter adjustment strategy, and transmitting it to the audio control platform to perform the parameter adjustment task;
[0077] A user auditory preference analysis module, which is used to obtain user audio control data through an audio control platform, identify the user's audio control mode based on the user audio control data, and thus obtain user audio control mode data; obtain real-time audio multi-modal sensing data, and perform user auditory preference analysis on the real-time audio multi-modal sensing data and the user audio control mode data, so as to obtain user auditory preference data;
[0078] An audio adjustment feedback optimization module, which is used to perform user personalized audio adjustment feedback optimization on the environmental dynamic parameter adjustment strategy according to the user auditory preference data, so as to obtain an optimized feedback audio dynamic adjustment strategy, and transmit it to the audio control platform to execute the parameter adjustment task.
[0079] The audio parameter determination system of the present invention can implement any audio parameter determination method of the present invention, and is used as a medium for coordinating the operations and signal transmissions between various modules to complete the audio parameter determination method. The internal modules of the system cooperate with each other, thus effectively avoiding the subjective errors in traditional manual adjustments and the problem of untimely response to environmental changes. Brief Description of the Drawings
[0080] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, objectives and advantages of the present invention will become more obvious:
[0081] Figure 1 It is a schematic flow chart of the steps of the audio parameter determination method of the present invention;
[0082] Figure 2 It is a detailed schematic flow chart of step S1 in the present invention;
[0083] Figure 3 It is a detailed schematic flow chart of step S2 in the present invention;
[0084] The realization, functional characteristics and advantages of the objectives of the present invention will be further described with reference to the embodiments and the drawings. Detailed Embodiments
[0085] The technical method of the present invention will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0086] In addition, the attached drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.
[0087] It should be understood that although the terms "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.
[0088] To achieve the above object, please refer to Figures 1 to 3 , the present invention provides a method for determining audio parameters, and the method includes the following steps:
[0089] Step S1: Obtain audio multi-modal sensing data, and perform audio usage environment perception based on the audio multi-modal sensing data to obtain audio usage environment data; perform multi-dimensional feature vector combination on the audio usage environment data to obtain a high-dimensional feature data set;
[0090] In this embodiment, by integrating multi-modal sensors (such as acoustic sensors, temperature sensors, and humidity sensors) in the audio device, real-time data of the audio device usage environment is obtained. The data collected by these sensors includes environmental factors such as sound intensity, frequency, air humidity, and temperature. Based on these data, the system analyzes the environmental characteristics through algorithms and combines the real-time working state of the audio device to form a comprehensive audio usage environment data set. Then, data processing methods (such as principal component analysis, feature engineering, etc.) are applied to perform multi-dimensional feature combination on these environmental data to generate a high-dimensional feature data set. For example, when the environmental temperature rises, the system will automatically adjust the output power or audio effect of the audio to cope with possible environmental changes.
[0091] Step S2: Perform environmental mode recognition on the high-dimensional feature data set to obtain audio usage environment mode data, and perform audio parameter response mode convolutional feature mapping based on the audio usage environment mode data to obtain an audio response feature map;
[0092] In this embodiment, based on the collected high-dimensional feature dataset, environmental pattern recognition technology is adopted, and the sound usage environment is modeled through machine learning algorithms such as clustering analysis and neural networks to identify environmental patterns (such as quiet environment, noisy environment, etc.). The identified environmental patterns provide a basis for subsequent sound parameter adjustment. Then, according to the sound usage environment pattern data, a convolutional neural network (CNN) is used to perform convolutional feature mapping of the sound response pattern to generate a sound response feature map.
[0093] Step S3: Determine the sound parameter influence factors based on the sound response feature map to obtain the sound parameter influence factors; analyze the environmental dynamic parameter adjustment strategy according to the sound usage environment data and the sound parameter influence factors to obtain the environmental dynamic parameter adjustment strategy, and transmit it to the sound control platform to execute the parameter adjustment task;
[0094] In this embodiment, the sound parameter influence factors can be determined through the sound response feature map. These factors include factors such as the volume, frequency response, and bass boost of the sound, and these parameters are closely related to factors such as environmental noise and room layout. Combining real-time environmental data, such as humidity and temperature, analyze the adjustment strategy of the sound parameters. For example, when the humidity is high, resulting in unclear bass effects of the sound, the bass output will be adjusted according to the influence factors to make the sound performance clearer. At the same time, the environmental dynamic parameter adjustment strategy will also be transmitted to the sound control platform to execute specific parameter adjustment tasks to ensure that the performance of the sound always conforms to environmental changes.
[0095] Step S4: Obtain the user sound control data through the sound control platform, and perform user sound control mode recognition according to the user sound control data to obtain the user sound control mode data; obtain the real-time sound multi-modal sensing data, and perform user auditory preference analysis on the real-time sound multi-modal sensing data and the user sound control mode data to obtain the user auditory preference data;
[0096] In this embodiment, the sound control platform obtains the user sound control data, which includes operations such as the user's volume adjustment and sound effect settings. Use algorithms such as support vector machine (SVM) or deep learning models to analyze the user's control habits and identify the user's sound control mode (such as preferring bass, higher volume, etc.). At the same time, real-time environmental perception is performed through the real-time sound multi-modal sensing data, the audio characteristics and noise level of the current environment are analyzed, and combined with the user's sound control mode data, the user's auditory preference is analyzed.
[0097] Step S5: Optimize the user personalized sound adjustment feedback of the environmental dynamic parameter adjustment strategy according to the user auditory preference data to obtain the optimized feedback sound dynamic adjustment strategy, and transmit it to the sound control platform to execute the parameter adjustment task.
[0098] In this embodiment, the dynamic parameter adjustment strategy is optimized personalized by combining the user's auditory preference data and real-time environmental changes. If it is recognized that the user prefers low-frequency enhancement in a noisy environment and the real-time environmental noise is relatively high, the low-frequency gain and high-frequency noise reduction strategies will be adjusted accordingly, and the genetic algorithm is used to optimize the low-frequency response and reduce the interference of external noise. The optimized feedback adjustment strategy will be transmitted to the audio control platform, and the device will automatically adjust the audio parameters (such as volume, low frequency, echo, etc.) to adapt to the user's preferences and ensure the best auditory experience is always provided in a dynamic environment.
[0099] Optionally, step S1 is specifically as follows:
[0100] Step S11: Obtain the audio multi-modal sensing data and perform data preprocessing on the audio multi-modal sensing data to obtain the audio multi-modal sensing data to be analyzed;
[0101] In this embodiment, the multi-modal sensing data of the environment where the audio system is located is obtained by integrating multiple sensors (such as temperature and humidity sensors, noise sensors, infrared sensors, etc.). Specifically in implementation, an integrated sensor module can be used to monitor the temperature, humidity, noise level in the environment and the object distance data detected by the infrared sensor in real time. The output signal of the sensor is amplified by using a low-noise amplifier, and high-frequency noise is removed through a filter to obtain clear multi-modal sensing data. These sensing data will be further preprocessed, such as removing outliers, normalizing the data range, etc., to obtain the audio multi-modal sensing data to be analyzed.
[0102] Step S12: Perform statistical analysis on the spatial distribution of the audio sensors for the audio multi-modal sensing data to be analyzed to obtain the audio sensor spatial distribution data, and perform audio sensing data fusion on the audio multi-modal sensing data to be analyzed according to the audio sensor spatial distribution data to obtain the audio usage sensing network;
[0103] In this embodiment, after the data preprocessing is completed, the spatial distribution of the audio multi-modal sensing data to be analyzed is then statistically analyzed. Specifically, the layout of the sensors can be analyzed, such as arranging multiple acoustic sensors in a room, and using the spatial distribution characteristics of the acoustic wave sensors to evaluate their response conditions at different positions. By performing spatial statistics on these sensor data (for example, calculating the response intensity distribution of the sensors, the spatial distance relationship between the sensors, etc.), the spatial distribution data of the audio sensors can be obtained. Based on these distribution data, fusion processing is performed on the audio multi-modal sensing data. For example, algorithms such as weighted average or Kalman filtering are used to fuse the data from different sensors to obtain a comprehensive audio usage sensing network for subsequent analysis and calculation.
[0104] Step S13: Extract the sensing time-frequency features using the audio sensing network to obtain the audio usage sensing time-domain feature data and the audio usage sensing frequency-domain feature data;
[0105] In this embodiment, after the audio sensing network is established, the next step is to extract the sensing time-frequency features. Specifically, by performing a Fourier transform on the real-time data collected in the audio sensor network, the time-domain signal is converted into a frequency-domain signal, and then the time-domain features (such as waveform, amplitude, time delay, etc.) and frequency-domain features (such as spectrum, center frequency, bandwidth, etc.) of the audio system are extracted. For example, in an audio system, if a microphone array is used for sound collection, by performing a short-time Fourier transform (STFT) on the sound signal, the frequency-domain features at each moment can be obtained, and further analyze the change characteristics of the audio signal, so as to obtain the audio usage sensing time-domain feature data and frequency-domain feature data.
[0106] Step S14: Perform environmental perception based on the audio usage sensing time-domain feature data and the audio usage sensing frequency-domain feature data to obtain the audio usage environment data;
[0107] In this embodiment, based on the audio usage sensing time-domain feature data and frequency-domain feature data, the next goal is to perform environmental perception. Specifically, machine learning algorithms (such as support vector machines, decision trees, neural networks, etc.) are used to train the feature data obtained from the sensor network, and the state of the environment is identified through the algorithm, such as detecting the noise level, sound source location, or the presence of echo in the environment. Taking noise detection as an example, the system can compare the background noise collected by the audio sensor with a set threshold to determine whether the current environment is in a quiet, noisy, or other state, and output the corresponding audio usage environment data.
[0108] Step S15: Perform multi-dimensional feature vector combination on the audio usage environment data to obtain a high-dimensional feature dataset.
[0109] In this embodiment, after obtaining the audio usage environment data, multi-dimensional feature vector combination is performed to construct a high-dimensional feature dataset. Specifically, it can be realized by splicing or weighted merging of multiple sensor data in the audio environment to form a dataset containing multi-dimensional information. For example, various features such as time-domain data, frequency-domain data, and environmental noise data can be vectorized, and then dimensionality reduction algorithms such as PCA (principal component analysis) are used to perform dimensionality reduction on these high-dimensional feature data to obtain the final high-dimensional feature dataset. These feature datasets will become the basis for subsequent analysis, modeling, and optimization.
[0110] Optionally, step S14 is specifically:
[0111] Based on the sensing time-domain characteristic data of the audio, perform envelope analysis on the sensing signal to obtain the waveform time-series data of the sensing signal;
[0112] In this embodiment, perform envelope analysis on the sensing signal using the sensing time-domain characteristic data of the audio. By collecting the audio sensor data, use the signal envelope analysis method (such as Hilbert transform) to extract the envelope characteristics of the audio signal. Specifically, for a segment of audio signal, by performing Hilbert transform on the audio signal, convert the signal into a complex form, and then extract its envelope curve, so as to obtain the amplitude change of the signal, generate the waveform time-series data of the sensing signal, and further use it for the analysis of audio performance.
[0113] Based on the sensing frequency-domain characteristic data of the audio, perform statistical analysis on the frequency distribution of the sensing signal to obtain the frequency distribution data of the sensing signal;
[0114] In this embodiment, perform Fourier transform on the frequency-domain data collected by the audio sensor to obtain the frequency distribution information. At this time, by calculating the amplitude and distribution of each frequency component in the frequency-domain data, the frequency distribution data of the sensing signal can be obtained. For example, by statistically analyzing the energy distribution of the signal within a specified frequency range, the main frequency band of the signal and other related frequency components can be identified, which is very helpful for the identification of noise sources and the localization of sound sources.
[0115] Perform wavelet transform time-frequency analysis on the waveform time-series data of the sensing signal and the frequency distribution data of the sensing signal to obtain the time-frequency characteristic data set of the sensing signal;
[0116] In this embodiment, use wavelet transform (such as Morlet wavelet or Daubechies wavelet) to perform joint time-frequency analysis on the time-series signal and frequency-domain data collected by the audio sensor. Wavelet transform obtains the joint characteristics of the signal in the time domain and frequency domain by decomposing different frequency bands of the signal and combining the time-domain information. In this way, it can simultaneously capture the changes of the signal at different time points and different frequency bands, so as to generate the time-frequency characteristic data set, which is helpful for the further analysis of the audio signal.
[0117] Based on the time-frequency characteristic data set of the sensing signal, perform environmental noise pattern recognition to obtain the environmental noise pattern data set;
[0118] In this embodiment, use machine learning algorithms (such as K-means clustering or support vector machine) to analyze the time-frequency characteristic data, and distinguish different types of environmental noise by identifying different noise patterns. For example, by analyzing the frequency characteristics, amplitude distribution, etc. of the noise, traffic noise, human voice noise, or mechanical noise in the environment can be identified, and the corresponding environmental noise pattern data set can be output, providing a basis for subsequent denoising processing.
[0119] Extract the time-frequency features of the audio based on the time-frequency feature dataset of the sensing signal, so as to obtain the time-frequency feature data of the audio signal;
[0120] In this embodiment, by further analyzing the time-frequency data of the audio signal, features of the audio signal are extracted, such as the pitch, loudness, duration, etc. of the audio signal. Methods such as short-time Fourier transform (STFT) or wavelet transform can be used to extract the time-frequency features of the audio signal. These time-frequency features can help judge the nature of the audio signal, such as music, language, noise, etc., and obtain the time-frequency feature data of the audio signal, providing support for subsequent noise removal and environmental analysis.
[0121] Denoise the time-frequency feature data of the audio signal according to the environmental noise pattern dataset, so as to obtain the time-frequency data of the audio signal to be analyzed, and analyze the acoustic wave reflection characteristics of the time-frequency data of the audio signal to be analyzed, so as to obtain the environmental acoustic wave reflection characteristic data;
[0122] In this embodiment, through a noise suppression algorithm (such as Wiener filtering or adaptive filtering), the noise components in the audio signal are removed according to the identified environmental noise pattern. By segmenting the audio signal and analyzing the noise pattern and audio signal characteristics of each segment respectively, the background noise is removed, and only the clear audio components related to the sound signal are retained. The time-frequency data of the denoised audio signal will be more suitable for subsequent acoustic characteristic analysis. Combining the time-frequency feature data of the audio signal and the model of the environmental space, the reflection analysis of the acoustic wave is carried out. An acoustic wave propagation model is used to simulate the reflection characteristics of acoustic waves on different surfaces (such as walls, windows, etc.), and the time delay and amplitude changes of the reflected acoustic waves are detected through the time-frequency data, so as to obtain the environmental acoustic wave reflection characteristic data. These data help to further understand the indoor acoustic environment, especially the acoustic wave propagation mode in complex environments.
[0123] Integrate the environmental dynamic changes based on the time-frequency feature dataset of the sensing signal, so as to obtain the environmental dynamic change data;
[0124] In this embodiment, the environmental dynamic changes are integrated based on the time-frequency feature dataset of the sensing signal. By performing a time series analysis on the continuous sensor data, the dynamic factors in the environment (such as noise, temperature and humidity changes, etc.) are monitored. Data fusion methods (such as Kalman filtering or particle filtering) are used to integrate the data at different time points to obtain more accurate environmental dynamic change data. These dynamic data help to understand the environmental state changes in real time and provide support for the adaptive adjustment of the sound system.
[0125] Perform multi-dimensional environmental perception data fusion on the environmental noise pattern dataset, the environmental dynamic change data, and the environmental acoustic wave reflection characteristic data, so as to obtain the sound use environment data.
[0126] In this embodiment, methods such as Kalman filtering and weighted average method are adopted to integrate the environmental data from the above multiple sources to obtain a comprehensive environmental perception data set. Through these fused data, the audio system can more accurately perceive the current environmental state, including noise level, environmental reflection characteristics and dynamic changes, providing accurate environmental data support for optimizing audio performance and providing a personalized audio experience.
[0127] Optionally, step S2 is specifically as follows:
[0128] Step S21: Perform principal component feature dimensionality reduction on the high-dimensional feature data set to obtain a principal component feature data set;
[0129] In this embodiment, principal component feature dimensionality reduction is performed on the high-dimensional feature data set, and the principal component analysis (PCA) method is used to reduce the dimensionality of the high-dimensional data set of the audio usage environment. Through PCA, redundant data in the original feature space is compressed into a lower-dimensional principal component space, retaining most of the data variability and information. For example, in the sensor data collected by the audio system, there are multiple variables (such as temperature, humidity, noise, etc.). PCA can calculate the covariance matrix of each variable and find the most important principal components, thereby transforming the original data set into a low-dimensional principal component feature data set, improving data processing efficiency and reducing computational complexity.
[0130] Step S22: Classify audio behaviors according to the principal component feature data set to obtain audio behavior data;
[0131] In this embodiment, audio behavior classification is performed according to the principal component feature data set, and the support vector machine (SVM) or K-means clustering algorithm is used to classify the principal component feature data. By learning and training the audio behavior data, different types of audio behaviors can be identified, such as different volume adjustments, audio switch operations, audio mode switching, etc. Specifically, the dimension-reduced principal component data is used as input features, and different audio behaviors are labeled through an SVM classifier. For example, the data is divided into categories such as "volume increase", "volume decrease", and "audio switch", and finally an accurate audio behavior data set is obtained, which is very crucial for subsequent intelligent adjustment and optimization.
[0132] Step S23: Perform environmental acoustic mode recognition on the audio behavior data based on the environmental dynamic change data to obtain audio usage environment mode data;
[0133] In this embodiment, environmental acoustic pattern recognition is performed on the audio behavior data based on the dynamically changing environmental data. By combining the audio behavior data with the real-time environmental change data (such as noise level, spatial temperature and humidity, echo, etc.), clustering algorithms (such as K-means or DBSCAN) are used to identify different environmental acoustic patterns. For example, in a noisy environment, acoustic patterns such as "high-noise environment" or "low-noise environment" can be identified and corresponding to specific audio behavior data. In this way, the audio output can be dynamically adjusted according to different environmental changes, so as to obtain the audio usage environment pattern data, thereby improving the adaptability and user experience of the audio system.
[0134] Step S24: Obtain the audio rated parameter set and perform parameter combination of the audio usage mode on the audio rated parameter set to obtain the audio usage mode parameter group;
[0135] In this embodiment, the audio rated parameter set is obtained through the audio control platform. These parameters include information such as the power output, frequency response range, and volume adjustment range of the audio. According to these rated parameter sets, parameter combination of the audio usage mode is performed through the algorithm in the system to obtain the audio usage mode parameter group. For example, when the audio works in different environmental modes, such as a low-noise or high-noise environment, it will be optimally combined according to the preset audio parameters (such as frequency response, volume output, etc.) to select the best parameter configuration to adapt to different auditory needs. These optimal combinations ensure that the audio exhibits the best sound effects in a specific environment.
[0136] Step S25: Perform audio response mode mapping according to the audio usage environment pattern data and the audio usage mode parameter group to obtain the audio response feature map.
[0137] In this embodiment, according to the previous recognition results, combining the environmental mode data and the audio usage mode parameter group, the mapping of the audio response mode is performed through a neural network model (such as a convolutional neural network). Specifically, by associating and mapping the audio behavior with the environmental feature data, an audio response feature map is generated. For example, the audio response feature map will show the best response of the audio output under different environmental conditions (such as in a quiet, noisy or echo environment), so as to ensure that the audio system can be adaptively adjusted to provide the optimal sound effect and user experience in different usage scenarios.
[0138] Optionally, step S25 is specifically:
[0139] Step S251: Perform magnitude normalization on the audio usage environment pattern data and the audio usage mode parameter group to obtain the normalized audio usage environment pattern data and the normalized audio usage mode parameter group;
[0140] In this embodiment, the audio usage environment mode data and the audio usage mode parameter group are first subjected to magnitude normalization processing. Specifically, using normalization methods (such as Z-score normalization or Min-Max normalization), the data is transformed into a unified dimension range, so that the values of all input audio environment mode data and audio mode parameter groups are within the same scale. For example, parameters such as the volume and frequency response range of the audio may have different magnitudes. Through normalization, we transform them into data with zero mean and unit variance, thus avoiding certain large-value parameters dominating the overall analysis during subsequent processing and ensuring the balanced contribution of each feature in the model.
[0141] Step S252: Perform audio usage mode feature selection on the normalized audio usage mode parameter group to obtain audio usage mode feature data;
[0142] In this embodiment, through a feature selection algorithm, such as LASSO (Least Absolute Shrinkage and Selection Operator) or decision tree method, the parameters most relevant to the audio system performance are identified. For example, from various parameters such as the volume adjustment, frequency response, and echo feedback of the audio, several of the most important features affecting the audio performance and user experience are selected, redundant and irrelevant features are removed, and a more representative and effective audio usage mode feature data set is obtained. For example, through feature selection, only the volume adjustment and echo feedback are retained as the important determinants of audio performance, thus simplifying the calculation of the subsequent model.
[0143] Step S253: Perform convolutional layer audio response mode feature mapping based on the normalized audio usage environment mode data and the audio usage mode feature data to obtain an audio response mode feature map;
[0144] In this embodiment, the feature extraction of the audio response mode is realized through a convolutional neural network (CNN). The convolutional layer performs convolutional operations on the input audio environment data and mode feature data through a series of filters (convolution kernels) to extract effective spatial or temporal features. For example, the input data includes the time-domain and frequency-domain features of the audio signal of the audio, as well as information such as the noise pattern of the environment. Through convolutional operations, high-dimensional features in the data are gradually extracted, thereby generating an audio response mode feature map, which shows the response mode of the audio under different environmental conditions. This feature map can effectively characterize the reaction characteristics of the audio system in various usage scenarios.
[0145] Step S254: Perform pooling layer audio response feature optimization on the audio response mode feature map to obtain an audio response feature map.
[0146] In this embodiment, the sound response feature map of the audio is optimized in the pooling layer. The pooling layer reduces the size of the feature map through dimensionality reduction operations while retaining its most important features. Specifically, the maximum pooling (MaxPooling) or average pooling (Average Pooling) method is used to process the feature map output by the convolutional layer. For example, a 2x2 pooling window is used to downsample the feature map, and the maximum eigenvalue or average value in the local area is selected to further reduce the computational amount and avoid overfitting. The sound response feature map output by the pooling layer can represent the key response characteristics of the audio system under different environmental conditions and has a smaller size, which is convenient for subsequent model analysis and inference.
[0147] Optionally, step S3 is specifically as follows:
[0148] Step S31: Extract the sound parameter feature weights from the sound response feature map to obtain a sound parameter weight data set;
[0149] In this embodiment, the sound parameter feature weights are extracted from the sound response feature map. Specifically, deep learning-based feature extraction techniques are used. For example, for the sound response feature map extracted by a convolutional neural network (CNN), the feature map is mapped to a set of important sound parameters through a fully connected layer or a weighted pooling layer. These sound parameters may include the frequency response, echo, volume, time delay, etc. of the audio. By applying a weighting algorithm such as L1 regularization or gradient boosting algorithm, the weights of each sound parameter are extracted to generate a data set containing sound parameter feature weights. For example, if the frequency response of a certain audio system has a greater impact on the audio effect, its weight will be relatively high, and the system can prioritize the optimization of the audio performance based on this data.
[0150] Step S32: Perform normalization of the parameter magnitude differences according to the sound parameter weight data set to obtain a sound parameter influence data set;
[0151] In this embodiment, normalization of the parameter magnitude differences is performed according to the sound parameter weight data set. This step aims to eliminate the differences in the numerical ranges of different sound parameters so that the parameters can be compared on the same magnitude. Specifically, the Min-Max normalization or Z-score standardization method is used to normalize the values of all sound parameters to the interval [0,1] to ensure that the influence of each parameter on subsequent analysis is equal. For example, if the frequency response range of an audio is between 20 Hz and 20 kHz, and the value of the echo is between 0 and 1, through the normalization operation, these parameters will be converted to a unified range, thus avoiding some parameters with larger magnitudes from dominating the entire model.
[0152] Step S33: Perform environmental sound behavior analysis according to the sound parameter influence data set to obtain sound behavior pattern data;
[0153] In this embodiment, environmental sound behavior analysis is performed based on the audio parameter impact dataset. Machine learning algorithms, such as K-means clustering or support vector machine (SVM), are used to analyze the audio parameter impact dataset to identify different environmental sound behavior patterns. For example, according to the performance of the audio device under different environmental conditions (such as open space, enclosed space, high-noise environment), the model can identify and classify different behavior patterns, such as "low-noise optimization mode", "dynamic volume adjustment mode", etc. These patterns can reveal the adaptability of the audio device in different environments and provide a basis for subsequent optimization.
[0154] Step S34: Perform a regression analysis of the audio parameter impact factors on the audio behavior pattern data and the audio usage environment pattern data to obtain the audio parameter impact factors;
[0155] In this embodiment, a regression analysis of the audio parameter impact factors is performed on the audio behavior pattern data and the audio usage environment pattern data. Through a regression analysis model, such as linear regression or multiple regression analysis, the relationship between the audio usage environment and the audio parameters is determined, and the audio parameter impact factors are extracted. These factors can quantify the parameter adjustments required for the audio device in different environments. For example, the model shows that in a high-noise environment, the impact factor of the echo cancellation function is larger, while in a quiet environment, the impact factor of the volume control is higher. Through regression analysis, the most suitable audio parameters can be selected for each environmental condition.
[0156] Step S35: Select the audio parameter combinations for environmental conditions based on the audio parameter impact factors and the audio usage environment data to obtain the audio parameter combination dataset for environmental conditions;
[0157] In this embodiment, the audio parameter combinations for environmental conditions are selected based on the audio parameter impact factors and the audio usage environment data. According to the identified impact factors, the appropriate parameter combinations are selected for the audio system under different environmental conditions. This can be achieved through optimization algorithms, such as genetic algorithms or particle swarm optimization algorithms, to automatically select the optimal parameter combinations. For example, in a closed environment, higher low-frequency response and echo reduction settings are selected, while in an open environment, the high-frequency response is adjusted to ensure clear sound. Ultimately, these parameter combinations will help improve the performance of the audio in specific environments.
[0158] Step S36: Develop a dynamic parameter adjustment strategy for the audio parameter combination dataset for environmental conditions to obtain the environmental dynamic parameter adjustment strategy and transmit it to the audio control platform to execute the parameter adjustment task.
[0159] In this embodiment, a dynamic parameter adjustment strategy is formulated for the environmental condition audio parameter combination dataset. By analyzing the environmental condition audio parameter combination data, a real-time adjustment strategy is formulated to ensure that the audio system can dynamically adjust according to environmental changes. For example, if an increase in environmental noise is detected, the volume will be automatically adjusted or the noise cancellation function will be enabled according to a pre-set adjustment strategy. In addition, the strategy will transmit this adjustment information to the audio control platform to implement automated parameter adjustment tasks. The audio control platform executes real-time parameter adjustment of the audio system according to the received strategy instructions to ensure that the audio device always maintains the best state.
[0160] Optionally, step S35 is specifically as follows:
[0161] Step S351: Quantify the response difference of the audio behavior pattern for the audio behavior pattern data, so as to obtain audio behavior pattern response error data;
[0162] In this embodiment, the response difference of the audio behavior pattern is quantified for the audio behavior pattern data, and the performance difference of the audio system in different usage environments is quantified using an error analysis algorithm (such as mean square error MSE). Response data is collected from multiple audio behavior patterns, such as the reactions of the audio in the mute, play, volume adjustment, etc. modes, and the quality differences of the audio outputs in each mode are compared. Then, the error values of these differences are calculated to obtain an audio behavior pattern response error dataset, indicating in which modes the audio system has a larger response deviation under different environmental conditions. For example, in a high-noise environment, the volume control response of the audio shows a larger error, while it is relatively small in a quiet environment.
[0163] Step S352: Perform response error-influence factor association on the audio parameter influence factors according to the audio behavior pattern response error data, so as to obtain response error-influence factor association data;
[0164] In this embodiment, according to the audio behavior pattern response error data, combined with the audio parameter influence factors, response error-influence factor association analysis is carried out. The specific method is to associate the response error of the audio behavior pattern with various parameters in the audio system (such as volume, frequency response, echo suppression, etc.), and analyze how each parameter affects the error of the audio behavior pattern. By establishing a multiple regression model or a correlation analysis model, the parameters that most significantly affect the audio behavior response error are found. This process helps to understand which audio parameters contribute the most to the audio behavior pattern error under specific environmental conditions. For example, if the error of the audio behavior is large in an environment with insufficient high-frequency response, it indicates that the high-frequency response is an important influence factor for audio performance optimization.
[0165] Step S353: Construct an audio parameter response influence model according to the response error-influence factor association data;
[0166] In this embodiment, an audio parameter response impact model is constructed based on the response error - impact factor correlation data. Using machine learning methods (such as support vector machines, random forests, or neural networks), the audio parameters are correlated with their impact on the behavior pattern error to construct an audio parameter response impact model. This model will predict its impact on the audio behavior pattern according to different values of the audio parameters. Specifically, for example, by training a neural network, the model can learn how to dynamically adjust the audio parameters (such as gain, audio equalization, etc.) under different environmental conditions (such as open space, closed space, noisy environment, etc.) to minimize the response error, thereby achieving the best audio performance.
[0167] Step S354: Construct an audio usage environment parameter space based on the audio usage environment data;
[0168] In this embodiment, an audio usage environment parameter space is constructed based on the audio usage environment data. The environmental data includes factors such as temperature, humidity, and noise level, and the performance of the audio device is affected by these environmental factors. By inputting this environmental data into a multi - dimensional parameter space, each environmental factor represents a dimension. For example, temperature and humidity respectively correspond to two dimensions in the space. Using multi - dimensional interpolation or Kriging interpolation methods, a parameter space map of the audio usage environment is generated to provide a basis for subsequent audio parameter adjustment. This process helps to understand how the audio system should be adjusted to meet the actual needs under different environmental conditions.
[0169] Step S355: Select the parameter combination with the lowest response error for the environmental conditions in the audio usage environment parameter space through the audio parameter response impact model, so as to obtain an environmental condition audio parameter combination data set.
[0170] In this embodiment, the parameter combination with the lowest response error for the environmental conditions in the audio usage environment parameter space is selected through the audio parameter response impact model. By inputting the environmental data into the audio parameter response impact model, the model can output a set of optimal audio parameter configurations to minimize the response error. For example, in a high - noise environment, the audio system will increase echo cancellation and noise reduction settings, while in a low - noise environment, audio clarity is emphasized. Through this model, the most suitable audio parameter combination can be automatically selected to ensure the best audio effect under different environmental conditions, and finally an environmental condition audio parameter combination data set is generated, and these data will be used to adjust the actual operation of the audio system.
[0171] Optionally, step S4 is specifically as follows:
[0172] Step S41: Obtain user audio control data through the audio control platform and perform data pre - processing on the user audio control data to obtain the user audio control data to be analyzed;
[0173] In this embodiment, the audio control platform starts data processing by receiving user audio control data (such as volume adjustment, sound effect setting, playlist selection, etc.) from the audio system. First, the audio control platform preprocesses this control data, including removing noise, filling in missing values, and normalizing the data. For data with sparse timestamps, interpolation is used to supplement the gaps to ensure data integrity. The preprocessed data is convenient for subsequent analysis of user behavior to further understand the user's control habits and preferences for the audio system. In specific operations, the platform can format this data into a unified time series format for synchronous analysis of various control behaviors.
[0174] Step S42: Extract the time-frequency characteristics of the user audio control for the data to be analyzed to obtain the user audio control time-frequency characteristic data, and perform user audio control mode recognition based on the user audio control time-frequency characteristic data to obtain the user audio control mode data;
[0175] In this embodiment, the preprocessed user audio control data is used for feature extraction, and the features in the control time series are extracted through time-frequency analysis. For the time-domain characteristics of the control data, the fast Fourier transform (FFT) is used to extract the frequency characteristics, and the wavelet transform is used to analyze the time-frequency characteristics in different frequency bands. By extracting these feature data, the user's control habits and their changes can be accurately captured. For example, when the user adjusts the volume at a higher frequency, it indicates that they are in a state of frequent control, while lower-frequency control indicates that the user adjusts the settings more gently. Then, a clustering algorithm is used to analyze these time-frequency characteristic data to identify the audio control modes of the user in different situations, such as entertainment mode, music listening mode, or work mode, etc., and finally obtain the user audio control mode data.
[0176] Step S43: Obtain real-time multimodal sensing data and perform data preprocessing on the real-time multimodal sensing data to obtain the real-time multimodal sensing data to be analyzed;
[0177] In this embodiment, real-time multimodal sensing data is obtained through an integrated sensor group, and this data includes sensor data of the audio system and environmental noise monitoring data. To ensure the accuracy and reliability of the data, the sensor data is first preprocessed. For example, electromagnetic interference noise is filtered out and the timestamps of the sensors are synchronously calibrated to ensure that the data of each sensor is consistent with the time. By using multimodal data fusion technology, the data from different sensors is integrated to form a complete set of real-time multimodal sensing data to be analyzed, which represents the real environment and user interaction state of the current audio system.
[0178] Step S44: Perform real-time audio usage environment perception based on the real-time multi-modal sensing data to be analyzed, so as to obtain real-time audio usage environment data;
[0179] In this embodiment, based on the real-time multi-modal sensing data to be analyzed, audio usage environment perception is performed. First, machine learning methods are used to perform real-time perception and analysis of factors such as environmental noise, light, and temperature. By performing Fourier analysis on the environmental noise data, the noise spectrum is obtained. Combining with the environmental temperature and humidity data, the changing trend of the environment can be identified. For example, the change in the surrounding environmental volume can be detected by sensors. When the audio is played at a high volume in a low-noise environment, it will cause discomfort to the user, while in a noisy environment, the audio needs to increase the volume to ensure clear sound quality. Finally, through data fusion processing, these data constitute the real-time audio usage environment data, providing an important basis for subsequent user preference analysis.
[0180] Step S45: Perform user auditory preference analysis based on the real-time audio usage environment data and the user audio control mode data, so as to obtain user auditory preference data.
[0181] In this embodiment, based on the real-time audio usage environment data and the user audio control mode data, user auditory preference analysis is performed. By establishing an evaluation model for user auditory preference and considering the influence of factors such as environmental noise, temperature, and light, the model will analyze the user's auditory preference in different situations. For example, if the user likes to listen to low-frequency music in a quiet environment, the bass setting of the audio can be adjusted according to this preference. Conversely, in a noisy environment, the high-frequency sound effect will be increased to improve clarity. Through machine learning algorithms (such as decision trees, support vector machines, etc.), the user's auditory preference is quantified and data is generated, becoming an important basis for the audio system to optimize sound quality, adjust sound effects, and improve user experience.
[0182] Optionally, step S45 is specifically:
[0183] Step S451: Extract the time-frequency characteristics of the audio signal from the real-time audio usage environment data, so as to obtain the real-time audio signal time-frequency characteristic data;
[0184] In this embodiment, the extraction of the time-frequency features of the audio signal of the real-time audio usage environment data is mainly completed through wavelet transform and short-time Fourier transform (STFT). The audio signal, including environmental noise and played audio data, is collected from the audio sensor. To obtain the time-frequency features, the audio signal is divided into multiple time segments, and each segment undergoes STFT analysis to obtain the frequency-domain features of that segment of audio. At the same time, the high-frequency components of the signal are processed using wavelet transform to finely capture instantaneous audio changes. For example, by extracting the energy changes in specific frequency bands (such as low frequency, medium frequency, high frequency), the current environmental noise level and the sound characteristics of different frequencies emitted by the audio system can be identified. Finally, the time-frequency feature data is used to describe the relationship between the current environment and the audio playback.
[0185] Step S452: Extract the time-frequency features of the audio control mode for the user audio control mode data, so as to obtain the time-frequency feature data of the user audio control mode.
[0186] In this embodiment, the extraction of the time-frequency features of the audio control mode involves performing time-frequency analysis on the audio control data input by the user (such as volume, sound effect selection, playback mode, etc.). By segmenting the audio control commands input by the user, each command segment undergoes Fourier transform on the time axis to extract the frequency-domain features. At the same time, wavelet analysis is performed on each command segment to extract detailed frequency information. For example, when the user adjusts the volume, they prefer lower-frequency changes, or when changing the sound effect, they prefer the middle and high-frequency bands. Using these time-frequency features, the audio control behavior of the user at different time periods can be accurately captured, providing data support for subsequent behavior modeling and adaptive adjustment.
[0187] Step S453: Perform multi-modal neural network feature fusion based on the time-frequency feature data of the real-time audio signal and the time-frequency feature data of the user audio control mode, so as to obtain the user behavior-environment audio feature space.
[0188] In this embodiment, multi-modal neural network feature fusion is a process of combining the time-frequency feature data of the real-time audio signal and the time-frequency feature data of the user audio control mode. Through a deep learning framework (such as convolutional neural network CNN), the two types of features are fused, and the neural network can simultaneously process the time-frequency features from the environmental audio signal and the user control mode. During the fusion process, a multi-layer neural network is used to weight the two types of features, enabling the network to automatically learn the relationship between the audio environment and the user behavior. For example, it is found that at a higher volume setting, the user is more inclined to increase the output of the low-frequency band, and this pattern will form a unique behavior-environment pattern in the feature space. Finally, the fused feature data forms a user behavior-environment audio feature space.
[0189] Step S454: Model the user's auditory fitness based on the user behavior - environmental audio feature space to obtain a user auditory fitness evaluation model;
[0190] In this embodiment, the user auditory fitness modeling is to establish an adaptation model by analyzing the relationship between the user's audio control mode and environmental audio features. This model is based on the data in the user behavior - environmental audio feature space and uses machine learning methods (such as support vector machine SVM or neural network) to train the user's auditory preferences under different environmental conditions. For example, the user prefers enhanced bass sound effects in a noisy environment and clearer high - frequency performance in a quiet environment. Through multiple trainings, the model can predict the user's auditory adaptation needs under different environments according to the user control mode and environmental changes, providing a basis for the dynamic adjustment of the audio system.
[0191] Step S455: Evaluate the user's auditory fitness for the real - time audio usage environment data and the user audio control data to be analyzed through the user auditory fitness evaluation model to obtain the audio environment user auditory fitness data;
[0192] In this embodiment, the real - time audio usage environment data and the user audio control data to be analyzed are evaluated through the user auditory fitness evaluation model. Specifically, the real - time audio usage environment data (such as environmental noise, temperature, humidity, etc.) and the user's control mode (such as volume adjustment, sound effect selection, etc.) are input into the evaluation model. The model calculates a fitness score based on the user's historical behavior and current environmental characteristics to evaluate whether the user's current auditory needs are met. For example, the user prefers a higher sound quality setting in a lower - noise environment and tends to increase the volume in a noisy environment. Finally, the generated audio environment user auditory fitness data can guide the audio system to make necessary adjustments.
[0193] Step S456: Map the user's auditory preferences for environmental conditions according to the audio environment user auditory fitness data to obtain the user auditory preference data.
[0194] In this embodiment, the user's auditory preferences for environmental conditions are mapped according to the audio environment user auditory fitness data. By establishing the mapping relationship between environmental conditions and user auditory preferences, the audio parameters can be automatically adjusted according to real - time environmental changes. For example, if it is evaluated that the user prefers a high - volume setting and enhanced bass in a low - noise environment, and tends to improve the clarity of the high - frequency part in a high - noise environment, the audio settings will be automatically adjusted according to this preference. This mapping process is based on the audio environment user auditory fitness data and combines the established user preference model to ensure that the user can obtain the best audio experience in various environments.
[0195] Optionally, this specification also provides an audio parameter determination system for performing the audio parameter determination method described above. The audio parameter determination system includes:
[0196] An audio usage environment perception module, configured to obtain audio multi-modal sensing data, perform audio usage environment perception based on the audio multi-modal sensing data to obtain audio usage environment data; perform multi-dimensional feature vector combination on the audio usage environment data to obtain a high-dimensional feature data set;
[0197] An environment mode recognition module, configured to perform environment mode recognition on the high-dimensional feature data set to obtain audio usage environment mode data, and perform audio parameter response mode convolution feature mapping based on the audio usage environment mode data to obtain an audio response feature map;
[0198] An audio parameter adjustment module, configured to determine an audio parameter influence factor based on the audio response feature map to obtain an audio parameter influence factor; perform environmental dynamic parameter adjustment strategy analysis based on the audio usage environment data and the audio parameter influence factor to obtain an environmental dynamic parameter adjustment strategy, and transmit it to the audio control platform to execute a parameter adjustment task;
[0199] A user auditory preference analysis module, configured to obtain user audio control data through the audio control platform, perform user audio control mode recognition based on the user audio control data to obtain user audio control mode data; obtain real-time audio multi-modal sensing data, and perform user auditory preference analysis on the real-time audio multi-modal sensing data and the user audio control mode data to obtain user auditory preference data;
[0200] An audio adjustment feedback optimization module, configured to perform user personalized audio adjustment feedback optimization on the environmental dynamic parameter adjustment strategy according to the user auditory preference data to obtain an optimized feedback audio dynamic adjustment strategy, and transmit it to the audio control platform to execute a parameter adjustment task.
[0201] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application document are intended to be encompassed by the present invention.
[0202] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features invented herein.
Claims
1. A method for determining an acoustic parameter, characterized in that: The following steps are involved: Step S1: Acquire audio multimodal sensing data, and perform audio usage environment perception based on the audio multimodal sensing data, thereby obtaining audio usage environment data; perform multidimensional feature vector combination on the audio usage environment data, thereby obtaining a high-dimensional feature data set; Step S2: performing environmental pattern recognition on the high-dimensional feature data set to obtain audio environment pattern data, and performing audio parameter response pattern convolution feature mapping according to the audio environment pattern data to obtain an audio response feature map; Step S3: determining the acoustic parameter influencing factor based on the acoustic response characteristic diagram, thereby obtaining the acoustic parameter influencing factor; analyzing the environmental dynamic parameter adjustment strategy according to the acoustic use environment data and the acoustic parameter influencing factor, thereby obtaining the environmental dynamic parameter adjustment strategy, and transmitting it to the acoustic control platform to execute the parameter adjustment task; Step S4: obtaining user audio control data through the audio control platform, and performing user audio control mode recognition according to the user audio control data, thereby obtaining user audio control mode data; Acquire real-time audio multimodal sensor data, and perform user auditory preference analysis on the real-time audio multimodal sensor data and user audio control mode data, thereby obtaining user auditory preference data; Step S5: Optimize the user's personalized sound adjustment feedback on the environment dynamic parameter adjustment strategy according to the user's auditory preference data, so as to obtain an optimized feedback sound dynamic adjustment strategy, and transmit it to the sound control platform to execute the parameter adjustment task.
2. The method for determining sound parameters according to claim 1, characterized in that: Step S1 is specifically as follows: Step S11: Acquire the acoustic multimodal sensor data, and perform data preprocessing on the acoustic multimodal sensor data, so as to obtain the acoustic multimodal sensor data to be analyzed; Step S12: performing sensor spatial distribution statistics on the audio multimodal sensor data to be analyzed, thereby obtaining audio sensor spatial distribution data, and performing audio sensor data fusion on the audio multimodal sensor data to be analyzed according to the audio sensor spatial distribution data, thereby obtaining an audio usage sensor network; Step S13: extracting sensing time-frequency features according to the audio usage sensing network, thereby obtaining audio usage sensing time domain feature data and audio usage sensing frequency domain feature data; Step S14: performing environmental perception based on the audio usage sensing time domain feature data and the audio usage sensing frequency domain feature data, thereby obtaining audio usage environment data; Step S15: performing multi-dimensional feature vector combination on the audio usage environment data to obtain a high-dimensional feature data set.
3. The method for determining sound parameters according to claim 2, characterized in that: Step S14 is specifically as follows: Based on the acoustic sensor time domain characteristic data, the sensor signal envelope analysis is performed to obtain the sensor signal waveform time series data; Performing frequency distribution statistics of sensor signals according to the sensor frequency domain characteristic data of the audio, thereby obtaining frequency distribution data of sensor signals; Perform wavelet transform time-frequency analysis on the sensor signal waveform time series data and the sensor signal frequency distribution data, so as to obtain the sensor signal time-frequency feature data set; Performing environmental noise pattern recognition according to the sensor signal time-frequency feature data set, thereby obtaining an environmental noise pattern data set; Extracting audio time-frequency features based on the sensor signal time-frequency feature data set, thereby obtaining audio signal time-frequency feature data; Performing audio signal denoising on the audio signal time-frequency feature data according to the environmental noise pattern data set, thereby obtaining the time-frequency data of the audio signal to be analyzed, and performing sound wave reflection characteristic analysis on the audio signal time-frequency data to be analyzed, thereby obtaining environmental sound wave reflection characteristic data; Based on the sensor signal time-frequency feature data set, the dynamic changes of the environment are integrated to obtain the dynamic changes of the environment data; Multi-dimensional environmental perception data fusion is performed on the environmental noise pattern data set, environmental dynamic change data, and environmental sound wave reflection characteristic data to obtain the audio usage environment data.
4. The method for determining sound parameters according to claim 1, characterized in that: Step S2 is specifically as follows: Step S21: performing principal component feature dimensionality reduction on the high-dimensional feature data set to obtain a principal component feature data set; Step S22: classifying the sound behavior according to the principal component feature data set, thereby obtaining the sound behavior data; Step S23: performing environmental acoustic pattern recognition on the audio behavior data based on the environmental dynamic change data, thereby obtaining audio use environment pattern data; Step S24: obtaining an audio rated parameter set, and combining the audio usage mode parameters with the audio rated parameter set, thereby obtaining an audio usage mode parameter group; Step S25: performing an audio response pattern mapping according to the audio use environment pattern data and the audio use pattern parameter group, thereby obtaining an audio response characteristic graph.
5. The method for determining sound parameters according to claim 4, characterized in that: Step S25 is specifically as follows: Step S251: performing magnitude standardization on the audio usage environment pattern data and the audio usage pattern parameter group, thereby obtaining standardized audio usage environment pattern data and the standardized audio usage pattern parameter group; Step S252: performing audio usage pattern feature selection on the standardized audio usage pattern parameter group, thereby obtaining audio usage pattern feature data; Step S253: performing convolutional layer sound response pattern feature mapping according to the standardized sound usage environment pattern data and the sound usage pattern feature data, thereby obtaining a sound response pattern feature map; Step S254: optimizing the acoustic response features of the acoustic response pattern feature map at the pooling layer, thereby obtaining an acoustic response feature map.
6. The method for determining sound parameters according to claim 1, characterized in that: Step S3 is specifically as follows: Step S31: extracting the acoustic parameter feature weights from the acoustic response feature graph, thereby obtaining an acoustic parameter weight data set; Step S32: normalizing the parameter magnitude differences according to the sound parameter weight data set, thereby obtaining the sound parameter influence data set; Step S33: performing environmental sound behavior analysis according to the sound parameter impact data set, thereby obtaining sound behavior pattern data; Step S34: performing an audio parameter influencing factor regression analysis on the audio behavior pattern data and the audio use environment pattern data, thereby obtaining the audio parameter influencing factor; Step S35: selecting an environmental condition sound parameter combination according to the sound parameter influencing factor and the sound use environment data, thereby obtaining an environmental condition sound parameter combination data set; Step S36: formulating a dynamic parameter adjustment strategy for the environmental condition sound parameter combination data set, thereby obtaining an environmental dynamic parameter adjustment strategy, and transmitting it to the sound control platform to execute the parameter adjustment task.
7. The method for determining sound parameters according to claim 6, characterized in that: Step S35 is specifically as follows: Step S351: quantifying the sound behavior pattern response difference of the sound behavior pattern data, thereby obtaining the sound behavior pattern response error data; Step S352: performing response error-influence factor association on the sound parameter influencing factor according to the sound behavior pattern response error data, thereby obtaining response error-influence factor association data; Step S353: constructing an acoustic parameter response influence model according to the response error-influence factor correlation data; Step S354: constructing an audio usage environment parameter space based on the audio usage environment data; Step S355: selecting a parameter combination with the lowest response error of the environmental conditions in the acoustic use environment parameter space through the acoustic parameter response influence model, thereby obtaining an environmental condition acoustic parameter combination data set.
8. The method for determining sound parameters according to claim 1, characterized in that: Step S4 is specifically as follows: Step S41: obtaining user audio control data through the audio control platform, and performing data preprocessing on the user audio control data, thereby obtaining user audio control data to be analyzed; Step S42: extracting user audio control time-frequency features from the user audio control data to be analyzed, thereby obtaining user audio control time-frequency feature data, and performing user audio control pattern recognition based on the user audio control time-frequency feature data, thereby obtaining user audio control pattern data; Step S43: acquiring real-time multimodal sensing data, and performing data preprocessing on the real-time multimodal sensing data, thereby obtaining real-time multimodal sensing data to be analyzed; Step S44: performing real-time audio usage environment perception according to the real-time multimodal sensing data to be analyzed, thereby obtaining real-time audio usage environment data; Step S45: Performing a user auditory preference analysis based on the real-time audio usage environment data and the user audio control mode data, thereby obtaining the user auditory preference data.
9. The method for determining sound parameters according to claim 8, characterized in that: Step S45 is specifically as follows: Step S451: extracting the time-frequency features of the audio signal from the real-time audio usage environment data, thereby obtaining the time-frequency feature data of the real-time audio signal; Step S452: extracting the time-frequency feature of the user's audio control mode data, thereby obtaining the time-frequency feature data of the user's audio control mode; Step S453: performing multimodal neural network feature fusion according to the real-time audio signal time-frequency feature data and the user sound control mode control time-frequency feature data, thereby obtaining a user behavior-environment audio feature space; Step S454: Modeling the user's auditory fitness based on the user behavior-environmental audio feature space, thereby obtaining a user's auditory fitness evaluation model; Step S455: performing user auditory adaptability evaluation on the real-time audio usage environment data and the user audio control data to be analyzed by using a user auditory adaptability evaluation model, thereby obtaining audio environment user auditory adaptability data; Step S456: Perform environmental condition user auditory preference mapping according to the user auditory adaptability data of the sound environment, thereby obtaining the user auditory preference data.
10. A system for determining sound parameters, characterized in that: Used to execute the sound parameter determination method as claimed in claim 1, the sound parameter determination system comprises: The audio usage environment perception module is used to obtain audio multimodal sensor data, and perform audio usage environment perception based on the audio multimodal sensor data, thereby obtaining audio usage environment data; and perform multidimensional feature vector combination on the audio usage environment data, thereby obtaining a high-dimensional feature data set; An environmental pattern recognition module is used to perform environmental pattern recognition on a high-dimensional feature data set to obtain audio environment pattern data, and to perform audio parameter response pattern convolution feature mapping based on the audio environment pattern data to obtain an audio response feature map; The audio parameter adjustment module is used to determine the audio parameter influencing factor based on the audio response characteristic diagram, so as to obtain the audio parameter influencing factor; analyze the environment dynamic parameter adjustment strategy according to the audio use environment data and the audio parameter influencing factor, so as to obtain the environment dynamic parameter adjustment strategy, and transmit it to the audio control platform to execute the parameter adjustment task; A user auditory preference analysis module is used to obtain user audio control data through the audio control platform, and perform user audio control mode recognition based on the user audio control data, thereby obtaining user audio control mode data; obtain real-time audio multimodal sensor data, and perform user auditory preference analysis on the real-time audio multimodal sensor data and the user audio control mode data, thereby obtaining user auditory preference data; The sound adjustment feedback optimization module is used to perform user-personalized sound adjustment feedback optimization on the environmental dynamic parameter adjustment strategy according to the user's auditory preference data, so as to obtain the optimized feedback sound dynamic adjustment strategy and transmit it to the sound control platform to execute the parameter adjustment task.
Citation Information
Cited By
Sound control method and system based on AI Internet of Things technology
CN120729920A
Sound control method and system based on AI internet of things technology
CN120729920B
Audio control method and system for stage performance
CN120916095A
Intelligent sound equipment with protection structure
CN121568001A