An audio output device immersive sound interaction method and system
By acquiring and optimizing multi-dimensional information, the problem of inaccurate sound effect adjustment in existing technologies has been solved, achieving more precise sound effect adjustment and improving the accuracy and comfort of the user's immersive experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN BOBEITE TECH DEV CO LTD
- Filing Date
- 2026-04-07
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies struggle to accurately identify a user's physiological state, leading to inaccurate sound effect adjustments and impacting the immersive listening experience.
By acquiring the current sound effect parameters of the audio output device, the original parameter adjustment information, and the original media information, and combining the user's visual, physiological activities, and sound information, the sound effect parameters are adjusted and optimized in multiple dimensions. The adjustment effect is evaluated using the sound effect adjustment effect coefficient and correlation degree, and interactive optimization adjustments are performed.
It achieves more accurate sound effect adjustments, improves the accuracy and comfort of the user's immersive experience, adapts to the user's personalized needs, and avoids excessive or inappropriate sound effect adjustments caused by misjudgment.
Smart Images

Figure CN122120693A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and specifically to an immersive sound effect interaction method and system for audio output devices. Background Technology
[0002] Currently, in smart entertainment spaces designed for individuals, advanced multi-channel audio output systems provide users with an ultimate immersive listening experience and can intelligently adjust sound effects based on the user's real-time status. However, the complexity of real-world scenarios far exceeds the idealized initial settings. Before watching a movie, users may engage in light physical activity, such as doing a few simple stretches or quickly tidying up their room; the activity is not strenuous, but it is enough to cause subtle changes in the user's physiological reference points. For example, the user's resting heart rate may be slightly higher than when completely relaxed, and the skin conductivity may increase slightly due to slight sweating or muscle activity.
[0003] At this point, after the movie starts playing, real-time monitoring of the user's physiological data includes reference point shifts caused by external factors. Because physiological state recognition is trained and calibrated based on a completely resting state, the user's slightly elevated heart rate and skin conductance are misinterpreted as initial excitement or heightened alertness towards the content. This triggers preset sound effect adjustment strategies to enhance immersion or increase stimulation. For example, during the movie's calm opening, the low frequencies of the background music might be slightly amplified, or the dynamic range of ambient sound effects might be widened, even giving the sound a stronger sense of space and impact in mundane scenes. When the user attempts to manually correct these inappropriate adjustments, the physiological reactions caused by discomfort or frustration are misinterpreted again, leading to inaccurate sound effect adjustments and a cycle of over-adjustment, ultimately severely damaging the user's immersion and overall experience. Summary of the Invention
[0004] The purpose of this invention is to provide an immersive sound effect interaction method and system for audio output devices, which solves the problem that when audio output devices provide an immersive sound effect experience, existing technologies have difficulty in accurately identifying the user's physiological state, and are prone to misjudging the user's emotional or cognitive reactions, resulting in inaccurate sound effect adjustments.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: an immersive sound effect interaction method for an audio output device, comprising: Obtain the current sound effect parameters, original parameter adjustment information, and original media information of the current media content output by the audio output device; Using the parameters to adjust the original information and media original information, the current sound effect parameters are adjusted to obtain the adjusted sound effect parameters and user feedback information on the sound effect parameters within a certain period of time after the adjustment. The user feedback information includes the user's visual information, the user's physiological activity information, and the sound information emitted by the user. Based on the adjusted sound effect parameters and user feedback, the sound effect adjustment coefficient is determined; Based on the sound effect adjustment coefficient, assess the correlation between the adjusted sound effect parameters and user feedback information; Based on the correlation degree and the correlation degree threshold, the original parameter adjustment information is optimized to obtain parameter adjustment optimization information; Using the parameter adjustment and optimization information, the current sound effect parameters are optimized and adjusted to complete the interactive optimization and adjustment of the current sound effect parameters.
[0006] Preferably, in the step of adjusting the current sound effect parameters using the parameters to adjust the original information and media original information, and obtaining the adjusted sound effect parameters and user feedback information on the sound effect parameters within a certain period of time after adjustment, the step of obtaining user feedback information includes: Acquire raw user feedback information, current audio stream information, and ambient sound information regarding the user's adjustments to sound effect parameters over a period of time after the adjustments are made. Low-frequency features are extracted from the environmental sound information to obtain environmental low-frequency features; Based on the original media information, the low-frequency characteristics of the current audio stream information, and the low-frequency characteristics of the environment, the environmental interference characteristics are determined; By utilizing the environmental interference characteristics, the original user feedback information is denoised to obtain the user feedback information.
[0007] Preferably, the step of determining environmental interference characteristics based on the original media information, the low-frequency characteristics of the current audio stream information, and the low-frequency characteristics of the environment includes: Using the original media information, the low-frequency features of the current audio stream information are used to determine the sound quality characteristics of the media content in the current audio stream information. The audio quality features of the media content are compared with preset media content features to obtain comparison anomaly features; Based on the aforementioned anomaly characteristics and the aforementioned low-frequency environmental characteristics, environmental interference characteristics are determined.
[0008] Preferably, the step of determining environmental interference characteristics based on the contrasting anomaly characteristics and the low-frequency environmental characteristics includes: By performing a similarity analysis on the aforementioned anomaly features and the aforementioned low-frequency environmental features, similarity features are obtained; The contrasting abnormal features and the low-frequency environmental features are superimposed to obtain the original environmental interference features; By using the aforementioned comparative identical features, the original features of the environmental interference are corrected to obtain the environmental interference features.
[0009] Preferably, in the step of obtaining the original user feedback information, current audio stream information, and ambient sound information regarding the sound effect parameters within a certain period after adjustment, the step of obtaining the original user feedback information and current audio stream information includes: Obtain the user's location spatial information; The spatial information of the region is divided to obtain each sub-region; Based on each of the sub-regions, obtain the sub-region information of user feedback and the sub-region information of the current audio stream within each sub-region; Based on the sub-region information of user feedback within each divided sub-region and the sub-region information of the current audio stream, the original user feedback information and the current audio stream information are obtained.
[0010] Preferably, the step of obtaining the original user feedback information and the current audio stream information based on the sub-region information of user feedback within each divided sub-region and the sub-region information of the current audio stream includes: Denoising is performed on the sub-region information of user feedback and the sub-region information of the current audio stream within each sub-region, resulting in the denoised sub-region information of user feedback and the denoised sub-region information of the current audio stream within each sub-region. Information fusion processing is performed on the denoised user feedback sub-region information and the denoised current audio stream sub-region information within each divided sub-region to obtain the original user feedback information and the current audio stream information.
[0011] Preferably, the step of performing information fusion processing on the denoised user feedback sub-region information and the denoised current audio stream sub-region information within each divided sub-region to obtain the original user feedback information and the current audio stream information includes: Information fusion processing is performed on the denoised user feedback sub-region information and the denoised current audio stream sub-region information in each divided sub-region to obtain the initial user feedback information and the initial current audio stream information. The integrity of the initial user feedback information and the initial current audio stream information are checked separately to obtain the initial user feedback information and the initial current audio stream information that have passed the checks. The verified initial information of user feedback and the verified initial information of current audio stream are preprocessed to obtain the original information of user feedback and the current audio stream information.
[0012] Preferably, the step of determining the sound effect adjustment coefficient based on the adjusted sound effect parameters and user feedback information includes: Based on the user feedback information, determine the sound effect response parameters; The sound effect parameters and sound effect response parameters are analyzed and processed to obtain the adjustment effect parameters; The sound effect adjustment coefficient is obtained by comparing the adjustment effect parameters with the preset effect parameter thresholds.
[0013] Preferably, the step of evaluating the correlation between the adjusted sound effect parameters and user feedback information based on the sound effect adjustment coefficient includes: Based on the sound effect adjustment coefficients, determine the correlation evaluation model; Using the aforementioned correlation evaluation model, the correlation evaluation is performed on the adjusted sound effect parameters and user feedback information to obtain the degree of correlation between the adjusted sound effect parameters and user feedback information.
[0014] The present invention also provides an immersive audio effect interactive system for an audio output device, the system comprising: The information acquisition module is used to acquire the current sound effect parameters of the audio output device, the original parameter adjustment information, and the original media information of the current media content output by the audio output device; The parameter adjustment module is used to adjust the current sound effect parameters using the parameter adjustment original information and media original information, and to obtain the adjusted sound effect parameters and user feedback information on the sound effect parameters within a certain period of time after the adjustment. The user feedback information includes the user's visual information, the user's physiological activity information and the sound information emitted by the user. The coefficient determination module is used to determine the sound effect adjustment coefficient based on the adjusted sound effect parameters and user feedback information; The correlation evaluation module is used to evaluate the correlation between the adjusted sound effect parameters and user feedback information based on the sound effect adjustment coefficient. The information optimization module is used to optimize the original information of parameter adjustment based on the correlation degree and the correlation degree threshold to obtain parameter adjustment optimization information; The parameter optimization module is used to optimize and adjust the current sound effect parameters using the parameter adjustment and optimization information, so as to complete the interactive optimization and adjustment of the current sound effect parameters.
[0015] Compared with the prior art, the immersive sound effect interaction method and system for audio output devices of the present invention have the following advantages: This invention acquires the current sound effect parameters of the audio output device, the original parameter adjustment information, and the original media information, and adjusts the current sound effect parameters to obtain adjusted sound effect parameters and user feedback information. The user feedback information includes not only the user's visual and physiological activity information but also the user's vocal information, thus comprehensively capturing the user's true reaction to the sound effect adjustment. Based on the adjusted sound effect parameters and user feedback information, a sound effect adjustment effect coefficient is determined, and the correlation between the adjusted sound effect parameters and user feedback information is further evaluated. Finally, based on the correlation degree and a correlation threshold, the original parameter adjustment information is optimized to obtain optimized parameter adjustment information, which is then used to optimize the current sound effect parameters. By introducing multi-dimensional user feedback information and combining it with the original media information, the invention can more accurately determine the user's true feelings about the sound effect adjustment, avoiding potential misjudgments that may arise from relying solely on physiological data. Furthermore, through iterative optimization of the original parameter adjustment information, this invention can continuously learn and adapt to the user's personalized needs, making the sound effect adjustment more precise and personalized. Therefore, it can significantly improve the accuracy, comfort, and personalization of the user's immersive experience, providing a more natural and enjoyable auditory experience. Attached Figure Description
[0016] To more clearly illustrate the specific embodiments of the present invention, the accompanying drawings used in the specific embodiments will be briefly described below. In all the drawings, the elements or parts are not necessarily drawn to scale.
[0017] Figure 1 This is a flowchart of an immersive sound effect interaction method for an audio output device according to the present invention.
[0018] Figure 2 This is a structural block diagram of an immersive sound effect interactive system for an audio output device according to the present invention.
[0019] In the diagram: 210, Information Acquisition Module; 220, Parameter Adjustment Module; 230, Coefficient Determination Module; 240, Correlation Evaluation Module; 250, Information Optimization Module; 260, Parameter Optimization Module.
[0020] The implementation and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0021] The following drawings disclose several embodiments of the present invention. For clarity, many practical details will be described in the following description. However, it should be understood that these practical details are not intended to limit the invention. That is, in some embodiments of the invention, these practical details are not essential. Furthermore, for the sake of simplicity, some conventional structures and components will be shown in the drawings in a simple schematic manner.
[0022] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indication will also change accordingly.
[0023] Furthermore, in this invention, the use of terms such as "first" and "second" is for descriptive purposes only and does not specifically refer to any order or sequence, nor is it intended to limit the invention. They are merely used to distinguish components or operations described using the same technical terms, and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of various embodiments can be combined with each other, but only if they are feasible for those skilled in the art. If a combination of technical solutions is contradictory or impossible to implement, such a combination should be considered nonexistent and not within the scope of protection claimed by this invention.
[0024] To further understand the content, features, and effects of this invention, the following embodiments are provided, and detailed descriptions are given below in conjunction with the accompanying drawings: Please see Figure 1 This invention provides an immersive sound effect interaction method for an audio output device, comprising the following steps: S100: Obtain the current sound effect parameters, parameter adjustment raw information, and media raw information of the current media content output by the audio output device. Here, the audio output device refers to any device capable of playing audio content, such as smart speakers, headphones, and home theater systems. Current sound effect parameters refer to the sound effect settings currently being used by the audio output device, such as volume, equalizer settings, and surround sound effects. Parameter adjustment raw information refers to the raw data or instructions used to guide the initial adjustment of sound effect parameters, which may originate from user presets, device default settings, or preliminary intelligent analysis. Media raw information refers to the raw data of the currently playing media content (such as movies, music, and games), including its type, genre, audio track characteristics, and content description. Specifically, the current sound effect parameters can be obtained by reading the configuration of the digital signal processor (DSP) or audio driver inside the audio output device. Parameter adjustment raw information may originate from presets such as movie mode or music mode manually selected by the user in the device settings interface. The original media information can be obtained by analyzing the metadata of the currently playing audio stream (such as ID3 tags), the encoding information of the video file, or through content recognition services. For example, it can be identified that the currently playing video is an action movie, whose audio track contains rich low-frequency effects and surround sound information.
[0025] S200. Using the parameter adjustment original information and media original information, the current sound effect parameters are adjusted to obtain the adjusted sound effect parameters and user feedback information on the sound effect parameters over a period of time after adjustment. The user feedback information includes the user's visual information, physiological activity information, and vocal information emitted by the user. Specifically, user feedback information refers to the user's reaction to the sound effect in various ways over a period of time after the sound effect parameters are adjusted, including the user's visual information (such as facial expressions and eye movements), physiological activity information (such as heart rate, respiration, and skin conductance), and vocal information emitted by the user (such as speech, sighs, and laughter). Adjusting the sound effect parameters refers to the sound effect parameters after preliminary adjustment. Specifically, the parameter adjustment original information may instruct the system to adjust the sound effect to a bass-enhanced mode, while the media original information indicates that a science fiction movie is currently playing. Based on this information, the current sound effect parameters are initially adjusted, for example, increasing the bass gain by 3dB and slightly enhancing the surround sound effect. User feedback is continuously monitored over a period of time after adjustment. The user's visual information can be captured by a camera, for example, analyzing whether the user's facial expression is pleasant and whether their eyes are focused on the screen. Users' physiological activity information can be acquired through wearable devices or non-contact sensors, such as monitoring their heart rate, respiratory rate, and skin conductance. Voice information emitted by users can be collected through microphone arrays, for example, analyzing whether the user makes sounds of surprise or satisfaction.
[0026] S300. Based on the adjusted sound effect parameters and user feedback, determine the sound effect adjustment effect coefficient. The sound effect adjustment effect coefficient is a quantitative indicator of the quality of the sound effect adjustment. Specifically, if after adjusting the sound effect parameters, the user's visual information shows an increase in facial expression pleasure, physiological activity information shows a more stable heart rate, and no sounds indicating dissatisfaction are emitted, then the adjustment effect can be considered good, and the sound effect adjustment effect coefficient can be set to a higher value, such as 0.8. Conversely, if the user exhibits frowning, large heart rate fluctuations, or complains, the sound effect adjustment effect coefficient may be set to a lower value, such as 0.2.
[0027] S400. Based on the sound effect adjustment effect coefficient, evaluate the degree of correlation between the adjusted sound effect parameters and user feedback information. The degree of correlation refers to the closeness of the relationship between the adjusted sound effect parameters and user feedback information. Specifically, a high sound effect adjustment effect coefficient indicates a strong positive correlation between the adjusted sound effect parameters and the user's positive feedback, meaning the sound effect adjustment effectively improved the user experience; in this case, the degree of correlation will be evaluated as high. If the sound effect adjustment effect coefficient is low, it indicates a low degree of correlation, which may mean that the sound effect adjustment failed to achieve the expected effect or even had a negative impact.
[0028] S500. Based on the correlation degree and correlation degree threshold, the original parameter adjustment information is optimized to obtain optimized parameter adjustment information. The correlation degree threshold is a preset standard used to determine whether the correlation degree meets the expected level. The optimized parameter adjustment information refers to more precise and effective parameter information, after optimization, used to guide subsequent sound effect parameter adjustments.
[0029] S600. Using the parameter adjustment and optimization information, the current sound effect parameters are optimized and adjusted to complete the interactive optimization and adjustment of the current sound effect parameters. Specifically, if the evaluated correlation degree is lower than a preset correlation degree threshold (e.g., 0.6), it indicates that the original parameter adjustment information may have defects and fails to accurately capture user needs or media content characteristics. At this time, the original parameter adjustment information is optimized, for example, by adjusting its weights, modifying its internal logic, or introducing new reference factors, thereby obtaining parameter adjustment and optimization information. Using the optimized information, the current sound effect parameters are optimized and adjusted again. For example, based on the previous bass enhancement, the mid-high frequencies are further fine-tuned according to the optimization information, or the intensity of surround sound is adjusted, in order to achieve a better user experience. Through iterative optimization, it can continuously learn and adapt to user preferences and environmental changes, realizing interactive optimization and adjustment of sound effect parameters.
[0030] This invention not only acquires multi-dimensional user feedback information (visual, physiological activity, and sound), but also determines the actual effect of sound effect adjustments by determining the sound effect adjustment effect coefficient and evaluating the correlation between the adjusted sound effect parameters and user feedback information. This allows it to identify whether the sound effect adjustments truly meet the user's immersive experience needs, rather than simply inferring from superficial physiological or behavioral data. When the adjustment effect is found to be poor or the correlation is low, the original parameter adjustment information can be optimized. After initial sound effect adjustments, if user feedback indicates discomfort (such as frowning and heart rate fluctuations), a low sound effect adjustment effect coefficient and correlation are assessed. In this case, the current adjustment strategy is not reinforced; instead, the original parameter adjustment information is optimized. For example, sensitivity to certain physiological indicators is reduced, or environmental noise analysis is introduced to eliminate external interference, resulting in a more reasonable secondary adjustment. Through interactive optimization of the adjustment method, sound effect adjustments can more accurately adapt to the user's real-time state and media content, effectively avoiding over- or inappropriate sound effect adjustments caused by misjudgment in traditional methods, significantly improving the accuracy and comfort of the user's immersive experience.
[0031] In some embodiments of this application described above, the step of adjusting the current sound effect parameters using the parameters to adjust the original information and media original information, and obtaining the adjusted sound effect parameters and user feedback information on the sound effect parameters within a certain period of time after adjustment, includes the following steps for obtaining user feedback information: This system acquires raw user feedback information, current audio stream information, and ambient sound information within a certain period after the sound effect parameters have been adjusted. Specifically, raw user feedback information refers to the initial feedback data on the sound effect effect expressed directly or indirectly by the user through visual, physiological, or auditory means within a certain period after the sound effect parameters have been adjusted. Current audio stream information refers to the real-time audio data stream of the media content being output by the audio output device. Ambient sound information refers to the background sound data collected by sensors such as microphones in the user's environment.
[0032] Low-frequency features are extracted from the environmental sound information to obtain environmental low-frequency features. This extraction aims to identify and quantify prevalent low-frequency noise components in the environment, which are often major sources of environmental interference, such as air conditioning noise and traffic noise. Extracting these low-frequency features provides foundational data for subsequent determination of environmental interference characteristics.
[0033] Environmental interference characteristics are determined based on the original media information, the low-frequency characteristics of the current audio stream, and the low-frequency characteristics of the environment. The original media information provides baseline information about the currently playing content, while the low-frequency characteristics of the current audio stream reflect the low-frequency components of the media content itself. By comprehensively analyzing this information along with the environmental low-frequency characteristics, the true environmental interference characteristics can be identified more accurately.
[0034] By utilizing the aforementioned environmental interference characteristics, the original user feedback information is denoised to obtain the actual user feedback information. Specifically, the denoising process aims to effectively remove noise components caused by environmental interference from the original user feedback information. For example, if the user's voice feedback is masked by environmental noise, denoising can make the user's true feedback clearer, thereby obtaining pure and accurate user feedback information.
[0035] Specifically, when a user listens to music using an audio output device in a noisy coffee shop and attempts to adjust the sound effects through voice commands or physiological responses, the background music, conversations, and the sound of the coffee machine constitute the environmental sound information. This is achieved by acquiring the user's voice commands (original user feedback information), the music being played (current audio stream information), and the environmental sound information of the coffee shop. Subsequently, low-frequency features of the coffee shop's environmental sounds are extracted, identifying low-frequency humming from the coffee machine, the bass frequencies of the background music, and other environmental low-frequency features. Simultaneously, the original media information and low-frequency features of the currently playing music are analyzed. Through comparative analysis, the unique interference of the coffee shop environment and the inherent low-frequency effects of the music can be accurately determined, thus obtaining environmental interference features. Finally, using these environmental interference features, the user's voice commands are denoised, effectively filtering out background noise in the coffee shop, clearly and accurately identifying the user's actual voice commands, and then making precise sound effect adjustments.
[0036] This embodiment provides a comprehensive data foundation for subsequent interference identification and denoising by acquiring original user feedback information, current audio stream information, and ambient sound information. Next, low-frequency feature extraction is performed on the ambient sound information to capture prevalent low-frequency noise in the environment, which often significantly impacts the clarity of user feedback. Based on this, by combining the original media information and the low-frequency features of the current audio stream, the low-frequency components of the media content itself can be intelligently distinguished from the low-frequency components of ambient noise, thereby accurately identifying environmental interference characteristics. Finally, the accurately identified environmental interference characteristics are used to denoise the original user feedback information, effectively filtering out environmental noise and ensuring that the obtained user feedback information is authentic, accurate, and undisturbed, providing a reliable basis for subsequent sound effect adjustments and optimizations.
[0037] In some embodiments of this application described above, the step of determining environmental interference characteristics based on the original media information, the low-frequency characteristics of the current audio stream information, and the low-frequency characteristics of the environment includes: Using the original media information, the low-frequency characteristics of the current audio stream are assessed for sound quality to determine the sound quality characteristics of the media content within the current audio stream. This step aims to identify the inherent sound quality characteristics of the current media content itself. For example, media content may intentionally contain low-frequency effects (such as explosions and heavy bass music), which are inherent components of the media content rather than environmental noise. By assessing sound quality, the sound quality characteristics of the media content in the current audio stream can be determined, thereby distinguishing the low-frequency components of the media content itself from potential environmental interference.
[0038] The audio quality characteristics of the current media content are compared with preset media content characteristics to obtain contrast anomaly characteristics. Specifically, the preset media content characteristics are a series of known or typical audio quality models or benchmarks for media content, such as typical low-frequency characteristics of different types of media like high-quality music, movie soundtracks, and voice calls. By comparing the audio quality characteristics of the current media content with these preset characteristics, abnormal parts in the current audio stream that do not match the expected audio quality of the media content can be identified. These abnormal parts are more likely to be caused by environmental interference. For example, if the low-frequency characteristics of a piece of high-quality music suddenly exhibit abnormal fluctuations similar to environmental noise, then this fluctuation is identified as a contrast anomaly characteristic.
[0039] Environmental interference features are determined based on the contrasting abnormal features and the low-frequency features of the environment. Specifically, environmental interference is only identified when the media content itself exhibits abnormal features and corresponding low-frequency features also exist in the environment. This avoids misjudging the inherent low-frequency components of the media content as environmental noise, thereby improving the accuracy of environmental interference feature identification.
[0040] Specifically, when a user watches a movie through an audio output device, which includes an explosion scene accompanied by strong low-frequency sound effects, there is also a low-frequency humming sound from an air conditioner in the user's environment. First, using the original media information of the movie, the low-frequency features of the explosion scene in the current audio stream are assessed for sound quality. Analysis identifies these low-frequency features as inherent to the movie content's sound quality, rather than anomalies. Next, the media content sound quality features are compared with a pre-defined movie sound effect feature library. Since the explosion sound is part of the movie's intended sound, no significant contrast anomalies are detected. However, if additional low-frequency features matching the air conditioner hum are detected besides the movie's inherent low-frequency sound effects, and these features do not match the movie's pre-defined sound quality features, these additional low-frequency features are identified as contrast anomalies. Finally, based on the contrast anomalies (i.e., the air conditioner hum) and the environmental low-frequency features extracted from the ambient sound information (i.e., the actually detected air conditioner hum), the environmental interference features are accurately determined. In this way, the low-frequency noise from the air conditioner can be accurately identified as environmental interference, without misjudging the low-frequency sound effects of explosion scenes in movies as noise and removing them. This ensures that users can fully experience the immersive sound effects of the movie, while effectively suppressing the interference of environmental noise on user feedback.
[0041] This embodiment uses original media information to guide the sound quality assessment process, ensuring that the analysis of the low-frequency characteristics of the current audio stream is based on the context of the media content. By determining the sound quality characteristics of the media content, a benchmark for the low-frequency characteristics of the media content itself can be established. Subsequently, comparing the sound quality characteristics of the media content with preset media content characteristics can identify anomalous features that do not match the expected sound quality of the media content. These anomalous features are more likely to originate from external environmental interference rather than the media content itself. Finally, by combining these anomalous features with environmental low-frequency features extracted from ambient sound information, environmental interference characteristics can be more accurately determined. This ensures the accuracy of subsequent noise reduction processing of the original user feedback information and avoids the erroneous removal of valid media information.
[0042] In some embodiments of this application described above, the step of determining environmental interference characteristics based on the contrasting abnormal characteristics and the low-frequency environmental characteristics includes: The comparative anomaly features and the environmental low-frequency features are analyzed for similarities to obtain comparative identical features. This step refers to detecting and quantifying the similarity between the two in the frequency domain, time domain, or energy distribution. For example, methods such as correlation analysis, pattern matching, or feature vector distance calculation can be used to identify common or highly similar low-frequency components. The purpose is to identify common interference components that exist both in media content anomalies and in environmental low-frequency noise, as these components may be key factors leading to misjudgments.
[0043] The contrast anomaly features and the low-frequency environmental features are superimposed to obtain the original environmental interference features. This step involves combining the two linearly or nonlinearly to form a preliminary characterization of the environmental interference. For example, the energy spectra of the two can be simply weighted and summed, or the feature vectors can be concatenated to obtain the original environmental interference features. The purpose is to gather all possible interference information to form a comprehensive preliminary estimate of the interference signal.
[0044] Using the aforementioned comparative identical features, the original environmental interference features are corrected to obtain the environmental interference features. This step refers to modifying the superimposed original environmental interference features based on the results of the identical feature analysis. For example, if the identical feature analysis shows that a low-frequency component of a certain frequency band is significantly present in both the comparative anomaly features and the environmental low-frequency features, then this component can be considered a relatively certain interference, and its weight in the original environmental interference features can be adjusted or enhanced. Conversely, if a component appears only in one of them, it may be necessary to reduce its weight in the original environmental interference features or suppress it. The purpose is to eliminate or weaken any misjudged components that may exist in the original environmental interference features, highlight the true environmental interference, and thus obtain more accurate environmental interference features.
[0045] This embodiment effectively solves the aforementioned problem of insufficient accuracy in determining environmental interference features by introducing identical feature analysis and correction processing. Specifically, firstly, by performing identical feature analysis on the contrasting anomaly features and low-frequency environmental features, the common low-frequency components between the two can be identified. These common components are often key areas where media content anomalies and environmental noise are confused. Secondly, by superimposing the contrasting anomaly features and low-frequency environmental features, all potential interference information can be comprehensively captured, forming the original environmental interference features. Using the previously obtained identical features, the original environmental interference features are corrected. This effectively filters out interference components that are not purely caused by the environment or that highly overlap with the characteristics of the media content, resulting in purer and more accurate environmental interference features. This allows for more accurate identification of genuine environmental interference, avoiding misjudging the characteristics of the media content itself as interference, thus providing a more reliable basis for subsequent noise reduction processing of user feedback information.
[0046] In some embodiments of this application described above, the step of obtaining the original user feedback information, current audio stream information, and ambient sound information for a period of time after the user adjusts the sound effect parameters includes the following steps: Acquire spatial information about the user's location. This step involves identifying and recording the user's physical spatial range using sensors (e.g., microphone arrays, cameras, infrared sensors, etc.) or pre-configured settings. Its purpose is to provide a foundation for subsequent spatial division and information collection.
[0047] The spatial information of the area is divided into sub-regions. This step involves logically or physically dividing the entire spatial area where the user is located into several smaller, independently processable sub-regions. For example, a room can be divided into multiple listening zones or activity zones. This division can be based on factors such as spatial geometry, acoustic characteristics, and user activity patterns, with the aim of achieving refined management and localized acquisition of spatial information.
[0048] Based on each defined sub-region, user feedback information and current audio stream information within that sub-region are acquired. This step involves independently collecting user feedback information (such as visual information, physiological activity information, and auditory information) and current audio stream information for each defined sub-region. For example, specific sensors can be deployed within a particular sub-region to capture user facial expressions, heart rate changes, ambient sounds, and media audio within that region. The aim is to ensure that the collected information has a clear spatial assignment, avoiding confusion between information from different regions.
[0049] Based on the sub-regional information of user feedback and the sub-regional information of the current audio stream within each sub-region, the original user feedback information and the current audio stream information are obtained. This step involves integrating or summarizing the local information obtained from each sub-region to form the global original user feedback information and current audio stream information. Integration can be a simple data splicing or a fusion based on some weight or priority, with the aim of constructing a holistic and more comprehensive original information from precise local information.
[0050] Specifically, in a living room environment, a user is experiencing immersive sound effects through an audio output device. First, the spatial information of the living room is acquired, such as by defining its boundaries through a room layout diagram or acoustic modeling. Then, the living room space is divided into three sub-zones: a sofa area, a dining area, and a window-side lounge area. In the sofa area, microphone arrays and cameras deployed near the sofa capture the user's vocal feedback, facial expressions, and posture information, as well as the audio stream of that area. Similarly, corresponding sensors are deployed in the dining area and the window-side lounge area to acquire user feedback sub-zone information and current audio stream sub-zone information for their respective areas. For example, the dining area might have the sound of clinking cutlery, while the window-side lounge area might be affected by traffic noise outside the window. Finally, the user feedback sub-zone information and current audio stream sub-zone information collected from the three sub-zones are integrated, for example, through weighted averaging or dynamic weights based on the user's current location, to obtain the raw user feedback information and current audio stream information for the entire living room. This ensures that the acquired raw information fully considers spatial differences, allowing subsequent sound effect adjustments to better adapt to the user's actual listening environment and state.
[0051] This embodiment introduces the concept of regional spatial information and meticulously divides it into multiple sub-regions, thereby achieving spatialized and refined acquisition of original user feedback information and current audio stream information. Decomposing the entire space into manageable sub-regions allows for independent and accurate acquisition of sub-regional information from user feedback and the current audio stream within each sub-region. Ultimately, by integrating the sub-regional information, more spatially representative and accurate original user feedback information and current audio stream information can be obtained, providing high-quality input data for subsequent environmental interference feature determination and user feedback information denoising.
[0052] In some embodiments of this application described above, the step of obtaining the original user feedback information and the current audio stream information based on the sub-region information of user feedback within each divided sub-region and the sub-region information of the current audio stream includes: The user feedback sub-region information and the current audio stream sub-region information within each segmented sub-region are denoised separately to obtain denoised user feedback sub-region information and denoised current audio stream sub-region information within each segmented sub-region. This step refers to applying appropriate noise suppression techniques to the user feedback data (e.g., user visual information, user physiological activity information, and user vocal information) and current audio stream data obtained from each segmented sub-region to eliminate or reduce interference components such as environmental noise, equipment noise, and data acquisition errors. For example, algorithms such as digital filtering, spectral subtraction, and wavelet denoising can be used to process audio information, image denoising can be performed on visual information, and signal smoothing can be performed on physiological activity information. The aim is to improve the purity and accuracy of information within each sub-region, laying the foundation for subsequent information fusion.
[0053] Information fusion processing is performed on the denoised user feedback sub-region information and the denoised current audio stream sub-region information within each sub-region to obtain the original user feedback information and the current audio stream information. This step integrates similar information from different sub-regions after denoising to form unified, complete, and high-quality original user feedback information and current audio stream information. For example, techniques such as weighted averaging, Kalman filtering, sensor fusion algorithms, and multimodal data fusion can be used to effectively integrate the denoised data from multiple sub-regions to compensate for the limitations of single sub-region information and enhance the robustness of the information. The aim is to extract more representative and accurate overall information from information obtained from multiple perspectives.
[0054] This embodiment first performs denoising processing on the user feedback sub-region information and the current audio stream sub-region information within each divided sub-region, effectively removing potential local noise and interference from the sub-region data and ensuring the purity of the input information. Subsequently, by performing information fusion processing on the denoised sub-region information, the purified data from different sub-regions is intelligently integrated, overcoming the potential bias or incompleteness of information from a single sub-region. Due to this step-by-step and refined processing, the final user feedback original information and current audio stream information have higher accuracy, completeness, and reliability, providing a high-quality data foundation for subsequent sound effect parameter adjustments and immersive sound effect interaction.
[0055] In some embodiments of this application described above, the step of performing information fusion processing on the denoised user feedback sub-region information and the denoised current audio stream sub-region information within each divided sub-region to obtain the original user feedback information and the current audio stream information includes: Information fusion processing is performed on the denoised user feedback sub-region information and the denoised current audio stream sub-region information within each sub-region to obtain initial user feedback information and initial current audio stream information. This step involves integrating the denoised user feedback sub-region information from different sub-regions and integrating the sub-region information of the current audio stream to form preliminary initial user feedback information and initial current audio stream information. Information fusion processing can employ various techniques, such as weighted averaging, feature concatenation, or machine learning models, to aggregate scattered sub-region information into unified initial information. The aim is to integrate scattered, denoised local information into globally usable preliminary information.
[0056] The initial user feedback information and the initial current audio stream information are each subjected to integrity checks to obtain verified initial user feedback information and verified initial current audio stream information. This step refers to performing data integrity checks on the initially fused initial user feedback information and initial current audio stream information. Specifically, it checks for missing, corrupted, or outlier data. For example, for the initial user feedback information, it checks whether it contains all expected visual, physiological, and auditory information, and whether the format of this information meets the requirements; for the initial current audio stream information, it checks whether the audio data stream is continuous and whether there are packet losses or encoding errors. The purpose is to ensure that the acquired initial information is complete and reliable, avoiding the impact of incomplete or corrupted data on the accuracy of subsequent processing.
[0057] The verified initial user feedback information and the verified initial current audio stream information are preprocessed to obtain the original user feedback information and the current audio stream information. This step refers to further standardization, normalization, feature extraction, or format conversion operations performed on the data after it has passed integrity verification. For example, user feedback information from different sources or in different formats can be unified into a standard format, and spectral analysis or feature extraction can be performed on the audio stream information to facilitate subsequent sound effect adjustment algorithms. The purpose is to transform the verified initial information into a data format suitable for the input of subsequent algorithm models, thereby improving the efficiency and accuracy of data processing.
[0058] This embodiment achieves preliminary information fusion by combining the denoised user feedback sub-region information and the denoised current audio stream sub-region information within each divided sub-region, thus converging scattered local information. Simultaneously, integrity verification is performed on the obtained initial user feedback information and current audio stream information to promptly identify and eliminate missing, corrupted, or abnormal data, ensuring data reliability. This integrity verification ensures that subsequent data processing is based on high-quality, defect-free data. Subsequently, the verified initial user feedback information and current audio stream information are preprocessed to further transform the data into a standardized format suitable for algorithm model input. This not only improves data processing efficiency but also provides more accurate and consistent input for subsequent sound effect parameter adjustments and optimizations.
[0059] In some embodiments of this application described above, the step of determining the sound effect adjustment coefficient based on the adjusted sound effect parameters and user feedback information includes: Based on the user feedback information, sound effect response parameters are determined. This step involves extracting and quantifying the user's actual perception and reaction to sound effect adjustments from the user feedback information. User feedback information may include the user's visual information, physiological activity information, and vocal information. For example, visual information may include facial expressions (such as smiling, frowning) and changes in eye contact; physiological activity information may include changes in physiological indicators such as heart rate, skin conductance, and electroencephalogram (EEG); and vocal information may include voice comments, sighs, or other nonverbal sounds. The raw feedback data is processed and analyzed to generate a series of sound effect response parameters that can quantify the user's emotions, satisfaction, or discomfort.
[0060] The adjusted sound effect parameters and sound effect response parameters are analyzed and processed to obtain adjustment effect parameters. This step refers to performing correlation analysis and evaluation on the sound effect adjustment parameters and the sound effect response parameters extracted from user feedback. This analysis aims to determine whether the actual effect of the sound effect adjustment is consistent with the expected goal, and the degree of user response to the adjustment. For example, if the sound effect parameter is adjusted to enhance bass, and the sound effect response parameters show that the user exhibits positive emotions and physiological excitement, then the adjustment can be considered to have produced a positive effect. This analysis can use statistical methods, machine learning algorithms, or pre-set rule models to quantify this correlation, thereby obtaining comprehensive adjustment effect parameters.
[0061] The adjustment effect parameters are compared with preset effect parameter thresholds to obtain the sound effect adjustment effect coefficient. This step involves comparing the obtained adjustment effect parameters with one or more pre-set effect parameter thresholds. The thresholds serve as a standard for evaluating the sound effect adjustment effect, determining whether the adjustment is successful, unsuccessful, or neutral. For example, positive and negative thresholds can be set. If the adjustment effect parameter exceeds the positive threshold, the sound effect adjustment effect coefficient is positive, indicating a good adjustment effect; if it is below the negative threshold, it is negative, indicating a poor adjustment effect; and if it falls between the two, it may be neutral. Through comparison, a quantified sound effect adjustment effect coefficient can be obtained, which directly reflects the effectiveness of the sound effect adjustment and the degree of improvement in user experience.
[0062] Specifically, when a user plays music using an audio output device, the bass effect of that music is enhanced. For a period of time after the adjustment, user feedback is continuously acquired, such as facial expressions (e.g., smiling) captured by a camera, a slight increase in the user's heart rate detected by a wearable device, and positive feedback (e.g., "This bass effect is great!") received through a microphone. Based on user feedback, sound effect response parameters are determined. For example, the degree of smile is quantified as an expression positivity score, heart rate changes are quantified as a physiological arousal index, and voice feedback is converted into an emotional tendency score using natural language processing. Subsequently, the sound effect response parameters are analyzed and processed against the previously adjusted bass enhancement parameters. For example, if the magnitude of bass enhancement is positively correlated with the user's expression positivity, physiological arousal, and emotional tendency in their voice, a higher adjustment effect parameter will be obtained. Finally, this adjustment effect parameter is compared with a preset effect parameter threshold. If the adjustment effect parameter exceeds the preset positive threshold, the sound effect adjustment coefficient is determined to be positive (e.g., 1), indicating that the bass enhancement adjustment has achieved a good immersive sound effect interaction effect; conversely, if it is below the negative threshold, it may be determined to be negative, indicating that the adjustment effect is not good.
[0063] This embodiment first extracts sound effect response parameters from user feedback information, quantifying the user's actual perception and reaction. Then, it comprehensively analyzes the quantified user response with the actual sound effect adjustment parameters to objectively evaluate the actual effect of the adjustment and obtain adjustment effect parameters. Finally, by comparing these adjustment effect parameters with preset evaluation criteria (i.e., preset effect parameter thresholds), the sound effect adjustment effect coefficient can be systematically and accurately determined. This ensures higher accuracy and reliability in determining the sound effect adjustment effect coefficient, avoids subjective judgment or fuzzy evaluation, and provides a solid data foundation for subsequent parameter optimization.
[0064] In some embodiments of this application described above, the step of evaluating the correlation between the adjusted sound effect parameters and user feedback information based on the sound effect adjustment coefficient includes: Based on the sound effect adjustment effect coefficients, a correlation evaluation model is determined. This step involves constructing or selecting a suitable mathematical model or algorithm framework based on the determined sound effect adjustment effect coefficients to quantify the relationship between adjusted sound effect parameters and user feedback information. The sound effect adjustment effect coefficients are a quantitative indicator of the initial effect of sound effect adjustments, reflecting the initial trend in user experience after the adjustments. The correlation evaluation model can be a statistical model, such as a regression model or a correlation analysis model, or a machine learning model, such as a support vector machine, neural network, or decision tree model. Its purpose is to provide a structured and quantifiable evaluation tool for subsequent correlation evaluation.
[0065] Using the aforementioned correlation evaluation model, the adjusted sound effect parameters and user feedback information are evaluated to determine the degree of correlation between them. This step involves taking the adjusted sound effect parameters and user feedback information as input, processing and calculating them through a pre-defined correlation evaluation model, and outputting a quantified degree of correlation. The adjusted sound effect parameters refer to the sound effect configuration after initial adjustments, while the user feedback information includes the user's visual information, physiological activity information, and vocal information, comprehensively reflecting the user's real-time perception and reaction to the sound effect adjustments. The correlation evaluation model analyzes the inherent relationships between these input data, for example, analyzing the correlation between changes in specific sound effect parameters and the user's heart rate, pupil changes, or vocal emotion, thereby determining the degree of correlation between the adjusted sound effect parameters and the user feedback information. This degree of correlation can be a numerical value, such as a floating-point number between 0 and 1, representing the strength of the correlation, or it can be a classification result, such as strong correlation, medium correlation, or weak correlation. Its purpose is to objectively and accurately measure the impact of sound effect adjustments on the user's immersive experience.
[0066] This embodiment introduces a correlation evaluation model, making the assessment of the correlation between sound effect adjustment parameters and user feedback information more scientific and precise. First, the correlation evaluation model is determined based on the sound effect adjustment effect coefficient. This model is constructed or selected based on the preliminary effect of the sound effect adjustment, thus ensuring the model's fit with the actual application scenario. For example, if the sound effect adjustment effect coefficient indicates that the user's response to low-frequency enhancement is relatively positive, the correlation evaluation model may focus on analyzing the relationship between low-frequency parameters and the user's physiological pleasure. Second, using this model to conduct a correlation evaluation of adjusted sound effect parameters and user feedback information allows for the systematic matching and analysis of complex, multi-dimensional user feedback information (such as visual, physiological, and auditory information) with specific changes in sound effect parameters. Due to the structured evaluation method, the evaluation results can more accurately reflect the real impact of sound effect adjustments on the user's immersive experience, avoiding the bias of subjective judgment and providing a solid data foundation for subsequent parameter optimization.
[0067] For an immersive sound effect interaction method for an audio output device based on any of the above embodiments, please refer to [link to relevant documentation]. Figure 2 The present invention also provides an immersive sound effect interaction system for an audio output device, the system comprising an information acquisition module 210, a parameter adjustment module 220, a coefficient determination module 230, a correlation evaluation module 240, an information optimization module 250, and a parameter optimization module 260.
[0068] The information acquisition module 210 is used to acquire the current sound effect parameters of the audio output device, the original parameter adjustment information, and the original media information of the current media content output by the audio output device.
[0069] The parameter adjustment module 220 is used to adjust the current sound effect parameters using the parameter adjustment original information and media original information, to obtain the adjusted sound effect parameters and user feedback information on the sound effect parameters within a certain period of time after adjustment. The user feedback information includes the user's visual information, the user's physiological activity information and the sound information emitted by the user.
[0070] The coefficient determination module 230 is used to determine the sound effect adjustment coefficient based on the adjusted sound effect parameters and user feedback information.
[0071] The correlation evaluation module 240 is used to evaluate the correlation between the adjusted sound effect parameters and user feedback information based on the sound effect adjustment coefficient.
[0072] The information optimization module 250 is used to optimize the original information of parameter adjustment according to the correlation degree and the correlation degree threshold to obtain parameter adjustment optimization information.
[0073] The parameter optimization module 260 is used to optimize and adjust the current sound effect parameters using the parameter adjustment and optimization information, so as to complete the interactive optimization and adjustment of the current sound effect parameters.
[0074] In this embodiment, the information acquisition module 210 comprehensively collects sound effects, parameters, and media information. The parameter adjustment module 220 performs initial sound effect adjustments. Subsequently, the coefficient determination module 230 and the correlation evaluation module 240 quantify and evaluate the adjustment effect. Finally, the information optimization module 250 and the parameter optimization module 260 form a closed-loop feedback to continuously optimize and adjust the sound effect parameters. Through a systematic interactive optimization mechanism, the accuracy of sound effect adjustments and the immersive user experience are ensured, effectively avoiding damage to the user experience caused by improper adjustments.
[0075] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the present invention specification.
Claims
1. A method for immersive sound effect interaction in an audio output device, characterized in that, include: Obtain the current sound effect parameters, original parameter adjustment information, and original media information of the current media content output by the audio output device; Using the parameters to adjust the original information and media original information, the current sound effect parameters are adjusted to obtain the adjusted sound effect parameters and user feedback information on the sound effect parameters within a certain period of time after the adjustment. The user feedback information includes the user's visual information, the user's physiological activity information, and the sound information emitted by the user. Based on the adjusted sound effect parameters and user feedback, the sound effect adjustment coefficient is determined; Based on the sound effect adjustment coefficient, assess the correlation between the adjusted sound effect parameters and user feedback information; Based on the correlation degree and the correlation degree threshold, the original parameter adjustment information is optimized to obtain parameter adjustment optimization information; Using the parameter adjustment and optimization information, the current sound effect parameters are optimized and adjusted to complete the interactive optimization and adjustment of the current sound effect parameters.
2. The immersive sound effect interaction method for an audio output device according to claim 1, characterized in that, In the step of adjusting the original information and media original information using the parameters, adjusting the current sound effect parameters, and obtaining the adjusted sound effect parameters and user feedback information on the sound effect parameters within a certain period of time after the adjustment, the step of obtaining user feedback information includes: Acquire raw user feedback information, current audio stream information, and ambient sound information regarding the user's adjustments to sound effect parameters over a period of time after the adjustments are made. Low-frequency features are extracted from the environmental sound information to obtain environmental low-frequency features; Based on the original media information, the low-frequency characteristics of the current audio stream information, and the low-frequency characteristics of the environment, the environmental interference characteristics are determined; By utilizing the environmental interference characteristics, the original user feedback information is denoised to obtain the user feedback information.
3. The immersive sound effect interaction method for an audio output device according to claim 2, characterized in that, The steps for determining environmental interference characteristics based on the original media information, the low-frequency characteristics of the current audio stream information, and the low-frequency characteristics of the environment include: Using the original media information, the low-frequency features of the current audio stream information are used to determine the sound quality characteristics of the media content in the current audio stream information. The audio quality features of the media content are compared with preset media content features to obtain comparison anomaly features; Based on the aforementioned anomaly characteristics and the aforementioned low-frequency environmental characteristics, environmental interference characteristics are determined.
4. The immersive sound effect interaction method for an audio output device according to claim 3, characterized in that, The steps for determining environmental interference characteristics based on the aforementioned anomaly characteristics and the aforementioned low-frequency environmental characteristics include: By performing a similarity analysis on the aforementioned anomaly features and the aforementioned low-frequency environmental features, similarity features are obtained; The contrasting anomaly features and the low-frequency environmental features are superimposed to obtain the original environmental interference features; By using the aforementioned comparative identical features, the original features of the environmental interference are corrected to obtain the environmental interference features.
5. The immersive sound effect interaction method for an audio output device according to claim 2, characterized in that, In the step of obtaining the original user feedback information, current audio stream information, and ambient sound information regarding the sound effect parameters after a period of time following adjustment, the step of obtaining the original user feedback information and current audio stream information includes: Obtain the user's location spatial information; The spatial information of the region is divided to obtain each sub-region; Based on each of the sub-regions, obtain the sub-region information of user feedback and the sub-region information of the current audio stream within each sub-region; Based on the sub-region information of user feedback within each divided sub-region and the sub-region information of the current audio stream, the original user feedback information and the current audio stream information are obtained.
6. The immersive sound effect interaction method for an audio output device according to claim 5, characterized in that, The step of obtaining the original user feedback information and the current audio stream information based on the sub-region information of user feedback within each divided sub-region and the sub-region information of the current audio stream includes: Denoising is performed on the sub-region information of user feedback and the sub-region information of the current audio stream within each sub-region, resulting in the denoised sub-region information of user feedback and the denoised sub-region information of the current audio stream within each sub-region. Information fusion processing is performed on the denoised user feedback sub-region information and the denoised current audio stream sub-region information within each divided sub-region to obtain the original user feedback information and the current audio stream information.
7. The immersive sound effect interaction method for an audio output device according to claim 6, characterized in that, The steps for fusing the denoised user feedback sub-region information and the denoised current audio stream sub-region information within each divided sub-region to obtain the original user feedback information and the current audio stream information include: Information fusion processing is performed on the denoised user feedback sub-region information and the denoised current audio stream sub-region information in each divided sub-region to obtain the initial user feedback information and the initial current audio stream information. The integrity of the initial user feedback information and the initial current audio stream information are checked separately to obtain the initial user feedback information and the initial current audio stream information that have passed the checks. The verified initial information of user feedback and the verified initial information of current audio stream are preprocessed to obtain the original information of user feedback and the current audio stream information.
8. The immersive sound effect interaction method for an audio output device according to claim 1, characterized in that, The steps for determining the sound effect adjustment coefficient based on the adjusted sound effect parameters and user feedback include: Based on the user feedback information, determine the sound effect response parameters; The sound effect parameters and sound effect response parameters are analyzed and processed to obtain the adjustment effect parameters; The sound effect adjustment coefficient is obtained by comparing the adjustment effect parameters with the preset effect parameter thresholds.
9. The immersive sound effect interaction method for an audio output device according to claim 1, characterized in that, The step of evaluating the correlation between the adjusted sound effect parameters and user feedback information based on the sound effect adjustment coefficient includes: Based on the sound effect adjustment coefficients, determine the correlation evaluation model; Using the aforementioned correlation evaluation model, the correlation evaluation is performed on the adjusted sound effect parameters and user feedback information to obtain the degree of correlation between the adjusted sound effect parameters and user feedback information.
10. An immersive audio effect interactive system for an audio output device, characterized in that, The system includes: The information acquisition module is used to acquire the current sound effect parameters of the audio output device, the original parameter adjustment information, and the original media information of the current media content output by the audio output device; The parameter adjustment module is used to adjust the current sound effect parameters using the parameter adjustment original information and media original information, and to obtain the adjusted sound effect parameters and user feedback information on the sound effect parameters within a certain period of time after the adjustment. The user feedback information includes the user's visual information, the user's physiological activity information and the sound information emitted by the user. The coefficient determination module is used to determine the sound effect adjustment coefficient based on the adjusted sound effect parameters and user feedback information; The correlation evaluation module is used to evaluate the correlation between the adjusted sound effect parameters and user feedback information based on the sound effect adjustment coefficient. The information optimization module is used to optimize the original information of parameter adjustment based on the correlation degree and the correlation degree threshold to obtain parameter adjustment optimization information; The parameter optimization module is used to optimize and adjust the current sound effect parameters using the parameter adjustment and optimization information, so as to complete the interactive optimization and adjustment of the current sound effect parameters.