Intelligent sound equipment sound field adaptive tuning method fusing multi-mode perception

By integrating multimodal perception and deep learning, combined with reinforcement learning and dynamic feedback, high-precision, real-time adaptive sound field tuning of smart audio devices in complex environments has been achieved. This solves the problem of sound-image mismatch in traditional audio systems under environmental changes, and has significant real-time performance and optimization accuracy.

CN121865160APending Publication Date: 2026-04-14GUANGZHOU LANJUE AUDIO EQUIP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU LANJUE AUDIO EQUIP CO LTD
Filing Date
2025-10-30
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing smart speaker devices cannot accurately acquire spatial geometric information, the listener's position, and dynamic changes in the environment, resulting in a mismatch between the sound image center and the sound field. Furthermore, they lack a real-time compensation mechanism, making it difficult to achieve high-precision adaptive sound field tuning in complex environments.

Method used

Acoustic, visual, and spatial location data are collected using a microphone array, camera, and spatial sensors. Multimodal feature fusion is achieved through deep neural networks, adaptive tuning parameters are generated by combining reinforcement learning algorithms, and real-time sound field correction is achieved through dynamic feedback.

Benefits of technology

It achieves real-time sound field adaptation with millisecond-level response time, surpassing the optimization precision of professional sound engineers, and maintains stable sound field estimation capability in complex noise or low-light environments, significantly improving the intelligence level of the audio system.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention relates to the technical field of tuning methods, and discloses an intelligent sound equipment sound field adaptive tuning method fusing multi-mode perception, and the method comprises the steps: respectively collecting acoustic, visual and spatial position data through a microphone array, a camera and a space sensor; multi-modal feature fusion is realized by using a deep neural network; constructing a sound field model according to the fusion features and identifying a scene type; self-adaptive tuning parameters are generated through a reinforcement learning algorithm; and real-time sound field correction is realized according to dynamic feedback. Compared with a traditional tuning method based on acoustic single-channel feedback, the method has the advantages that the response time is shortened to the millisecond level through multi-modal fusion, real real-time sound field self-adaption is achieved, acoustic mapping rules of different rooms are learned through the deep neural network, and the effect close to or beyond the manual tuning effect can be achieved without manual intervention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of tuning methods, specifically to an intelligent speaker sound field adaptive tuning method that integrates multimodal perception. Background Technology

[0002] Most existing smart speaker devices achieve simple automatic equalization (Auto EQ) functions through acoustic feedback, such as collecting echo characteristics through microphone arrays and then adjusting the frequency response curve to adapt to the room's reflection characteristics. However, this method relies solely on single acoustic data and cannot accurately capture spatial geometry, the listener's position, and dynamic environmental changes. For example, when there are large areas of glass, furniture, or people in the room, sound wave reflections change over time, and traditional tuning algorithms cannot adapt in a timely manner. Equalization methods based solely on acoustic models struggle to distinguish between "obstacle reflections" and "effective direct sound delays," leading to distortion problems such as low-frequency booming and phase misalignment. When the listener moves or changes posture, the sound image center and sound field equalization no longer match, and current technologies lack real-time compensation mechanisms for the "spatial distribution of the human ear." For a long time, enabling speaker systems to understand the spatial environment "like a human" and automatically optimize the sound field has been a technological bottleneck in the field of audio engineering. This invention addresses this pain point by proposing, for the first time, a multimodal perception tuning method that integrates visual, acoustic, and spatial sensing technologies. It utilizes deep neural networks to achieve cross-modal feature association learning, thereby enabling high-precision, low-latency adaptive sound field optimization. Therefore, there is an urgent need for an intelligent speaker sound field adaptive tuning method that integrates multimodal perception to solve the aforementioned problems. Summary of the Invention

[0003] The purpose of this invention is to provide an intelligent speaker sound field adaptive tuning method that integrates multimodal perception to solve the problems mentioned in the background art.

[0004] To achieve the above objectives, the present invention provides the following technical solution:

[0005] A method for adaptive sound field tuning of an intelligent speaker integrating multimodal perception includes: collecting acoustic, visual, and spatial location data through a microphone array, camera, and spatial sensor; fusing multimodal features using a deep neural network; constructing a sound field model based on the fused features and identifying scene types; generating adaptive tuning parameters through a reinforcement learning algorithm; and achieving real-time sound field correction based on dynamic feedback. Preferably, the microphone array includes at least eight MEMS microphones distributed around the speaker housing to form a spatial sound pressure distribution matrix. Preferably, the visual information is collected by a wide-angle camera, and a convolutional neural network is used to identify room boundaries and obstacle distribution. Preferably, the spatial sensor includes an infrared or ultrasonic module for real-time detection of the listener's relative position and posture. Preferably, the multimodal feature fusion uses a three-branch deep network structure to process acoustic, visual, and spatial features respectively, outputting a unified semantic vector. Preferably, the sound field model uses a hybrid structure of convolutional neural networks and graph convolutional networks to model the room's acoustic response function. Preferably, the adaptive tuning parameters are generated through a reinforcement learning algorithm, with the training objective function being to minimize auditory error and optimize sound image stability. Preferably, the system has a dynamic feedback channel, which enables self-calibration of tuning parameters through real-time sound wave monitoring and visual re-recognition. Preferably, the method can reload new equalization parameters within 50 milliseconds after detecting environmental changes, achieving imperceptible sound field updates. Preferably, the method can still operate stably in complex noise or low-light environments, overcoming the problem of traditional acoustic feedback failing in high-noise scenes through visual-assisted sound field estimation.

[0006] Compared with the prior art, the present invention has the following beneficial effects:

[0007] 1. Significantly improved real-time performance: Compared with traditional tuning methods based on acoustic single-channel feedback, this invention shortens the response time to the millisecond level through multimodal fusion, achieving true "real-time sound field adaptation".

[0008] 2. Optimization precision surpassing that of professional sound engineers: Deep neural networks learn the acoustic mapping patterns of different rooms, achieving near-or even better sound mixing results without human intervention.

[0009] 3. Unexpected technical effects: Experiments have shown that the multimodal model can still maintain stable sound field estimation capabilities in low light or high noise environments, which is significantly better than traditional equalization algorithms based on microphone feedback, overcoming the previous technical bias that "visual features cannot assist in sound field control". Detailed Implementation

[0010] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0011] A method for adaptive sound field tuning of an intelligent speaker that integrates multimodal perception includes: collecting acoustic, visual, and spatial location data through a microphone array, a camera, and a spatial sensor, respectively; using a deep neural network to achieve multimodal feature fusion; constructing a sound field model based on the fused features and identifying scene types; generating adaptive tuning parameters through a reinforcement learning algorithm; and achieving real-time sound field correction based on dynamic feedback.

[0012] The microphone array includes at least eight MEMS microphones distributed around the speaker housing to form a spatial sound pressure distribution matrix.

[0013] The visual information is collected by a wide-angle camera and identified by a convolutional neural network to identify room boundaries and obstacle distribution.

[0014] The spatial sensor includes an infrared or ultrasonic module, used to detect the relative position and posture of the listener in real time.

[0015] The multimodal feature fusion adopts a three-branch deep network structure to process acoustic features, visual features and spatial features respectively, and output a unified semantic vector.

[0016] The sound field model employs a hybrid structure of convolutional neural networks and graph convolutional networks to model the room's acoustic response function.

[0017] The adaptive tuning parameters are generated through a reinforcement learning algorithm, and the training objective function is to minimize auditory error and optimize sound image stability.

[0018] The system is equipped with a dynamic feedback channel, which enables self-calibration of tuning parameters through real-time sound wave monitoring and visual re-recognition.

[0019] The method can reload new equalization parameters within 50 milliseconds after detecting environmental changes, achieving a seamless sound field update.

[0020] This method can still operate stably in complex noise or low light environments, and overcomes the problem of traditional acoustic feedback failing in high-noise scenes by using visual-assisted sound field estimation.

[0021] This invention provides an intelligent speaker sound field adaptive tuning method that integrates multimodal perception. Its working principle is based on multimodal information fusion, deep learning feature mapping, reinforcement learning parameter optimization, and a real-time dynamic feedback mechanism. The core of this method lies in perceiving the auditory scene through multiple sensing modes, fusing and reasoning acoustic information, visual spatial information, and human-computer interaction state information to achieve a closed loop similar to the human auditory system: "environmental understanding—judgment—tuning—feedback." The entire method can be divided into the following five main stages:

[0022] (1) Multimodal data acquisition stage;

[0023] (2) Multimodal feature fusion and correlation modeling stage;

[0024] (3) Sound field modeling and scene recognition stage;

[0025] (4) Adaptive tuning parameter generation stage;

[0026] (5) Dynamic feedback and self-calibration stage.

[0027] The principles and implementation mechanisms are explained below:

[0028] I. Multimodal Data Acquisition Stage: Traditional acoustic tuning methods rely solely on microphones to collect echo signals, resulting in a single sampling dimension and difficulty in describing the sound wave propagation patterns in complex spatial environments. This invention constructs a multimodal sensing system, integrating a microphone array, a visual camera, and spatial sensors (including infrared or ultrasonic modules) into the audio system to achieve multi-dimensional environmental perception. 1. Microphone Array Acquisition of Acoustic Information: The microphone array typically consists of 8-16 MEMS microphones, distributed in a ring or matrix on the surface of the speaker enclosure. The sound source location and reflection path are calculated using the phase difference method and the time difference of arrival method (TDOA). The signals acquired by the array are processed by short-time Fourier transform (STFT) to obtain a spectrum diagram, which is used to describe the reflection characteristics and energy distribution of different frequency bands. 2. Camera Acquisition of Visual Spatial Information: A front-facing wide-angle camera captures the listening space in real time, generating a three-dimensional spatial point cloud using a depth estimation network (DepthNet). This image data is used to identify room boundaries, wall materials, furniture positions, etc., thereby determining the reflection coefficient and diffusion characteristics. For example, when the camera detects a large area of ​​glass or a smooth wall, the system automatically increases mid-to-high frequency attenuation compensation. 3. Spatial sensors acquire the relative position and dynamic posture of the human ear. Infrared or ultrasonic sensing modules calculate the three-dimensional coordinates of the listener relative to the sound source through reflection delay, and combine this with accelerometer readings or a human posture estimation model to determine whether the listener is within the main acoustic axis. When the listener moves or turns their head, the system can capture the posture changes in real time, providing a basis for subsequent sound image relocalization. Through a time synchronization module, all three types of sensor data are accompanied by a unified timestamp, forming a complete time sequence input, ensuring the alignment and correlation of multimodal data in subsequent processing stages.

[0029] II. In the multimodal feature fusion and correlation modeling stage, the core challenge of multimodal data lies in how to map heterogeneous signals to the same feature space. This invention adopts a multi-branch fusion structure based on deep neural networks. After extracting acoustic features, visual features, and spatial features separately, semantic alignment and tensor concatenation are performed in the fusion layer to form a unified multimodal joint feature vector. 1. The acoustic feature extraction branch uses a convolutional neural network (CNN) to extract reflection patterns, reverberation delay, and energy attenuation features from the spectrogram; the network training objective is to reconstruct the room impulse response (RIR) curve to reflect the temporal characteristics of sound wave propagation in space. 2. The visual feature extraction branch uses a ResNet network to extract features from the spatial image, identify the room geometry, reflective surface location, and material type, and output a high-dimensional feature tensor containing spatial geometric information. 3. The spatial location feature branch processes the dynamic pose data of the listener through a long short-term memory network (LSTM) to capture the temporal dependence of position changes. 4. Feature Fusion Mechanism: The outputs of the three branches are concatenated in the fusion layer, and the weights of each modality feature are calculated through an attention mechanism to highlight the most critical perceptual signals in the current scene. After dimensionality reduction by a fully connected layer, the fused output forms a low-dimensional "environment description vector," which serves as the input to the sound field modeling and scene recognition modules.

[0030] III. Sound Field Modeling and Scene Recognition Stage: The core task of this stage is to establish a mapping model from multimodal features to acoustic spatial response. This invention proposes a sound field modeling algorithm based on a hybrid architecture of graph convolutional neural networks (GCN) and convolutional networks to infer the room acoustic response function (RIR) and scene category. 1. Sound Field Topology Modeling: The room is divided into multiple nodes, each representing a spatial unit, whose features include sound pressure, reflection coefficient, and absorption rate. The weights of the connection edges between nodes are determined based on spatial distance and obstacle obstruction. The graph convolutional network propagates features from this graph structure, establishing a mapping relationship between the local sound field and the overall sound field. 2. Scene Recognition Model: The model simultaneously outputs room type labels, such as "living room," "bedroom," "meeting room," and "vehicle space." Different types of scenes correspond to different acoustic objective functions. For example, in a meeting scene, speech intelligibility needs to be enhanced, while in a music mode, the naturalness of spatial reverberation needs to be maintained. 3. The room response prediction model is trained by minimizing the difference between the actual echo signal and the predicted RIR. It can generate room impulse response estimates in real time after inputting arbitrary multimodal data, thus providing data support for subsequent parameter optimization. Through this stage, the system obtains a complete description of the environmental acoustic state and has the ability to autonomously judge the difference between the current sound field structure and the desired auditory effect.

[0031] IV. Adaptive Tuning Parameter Generation Stage: After completing scene recognition and sound field modeling, the system needs to determine the optimal tuning parameters for that environment. This invention introduces an adaptive parameter optimization mechanism based on Reinforcement Learning (RL), obtaining the optimal tuning strategy through continuous trial and error and feedback learning. 1. State and Action Definition: State: Consists of the environment description vector, the current sound field response, and the listener's position; Action: Includes adjusting equalizer gain, delay compensation, sound image center, reverberation parameters, etc. 2. Reward Function Design: The reward function comprehensively considers three indicators: subjective listening satisfaction (simulated by a human ear model), minimization of acoustic error, and uniformity of energy distribution. A higher reward value indicates that the tuning effect is closer to the ideal state. 3. Policy Training and Inference: The Deep Deterministic Policy Gradient (DDPG) algorithm is used to continuously optimize the tuning parameters through online interaction. The model can ultimately output the optimal parameter set in real time under different environments, achieving precise acoustic adaptive control. 4. Parameter Execution and Transition Smoothing: When the system generates new tuning parameters, a smooth transition algorithm is used to prevent abrupt changes in tone or delay, ensuring the continuity and naturalness of the listening experience.

[0032] Experimental results show that the reinforcement learning model of the present invention can converge after about 20 iterations, and the final tuning accuracy is improved by about 35% and the response speed is improved by about 60%.

[0033] V. Dynamic Feedback and Self-Calibration Stage: The sound field is dynamically changing; for example, opening and closing doors and windows, movement of people, and noise interference can all cause abrupt changes in acoustic characteristics. This invention sets up a continuous feedback path, enabling the system to automatically correct itself after detecting changes, forming a true closed-loop control. 1. Real-time Acoustic Feedback Detection: A microphone array continuously monitors the spatial response curve after the sound is emitted. When the reverberation time (RT60) or frequency response curve differs from the prediction model by more than a set threshold (e.g., ±1.5dB), feedback correction is immediately triggered. 2. Visual Dynamic Re-identification: A camera continuously captures scene changes. When furniture movement or personnel displacement is detected, the spatial layout is automatically recalculated and the acoustic model topology is updated. 3. Self-Calibration Mechanism: By rapidly updating the corresponding feature weights and RIR parameters in the fusion model, low-latency self-repair is achieved. The entire process is completed within 50 milliseconds, making the sound change almost imperceptible to the human ear. 4. The incremental learning and memory retention system can accumulate historical scene data during long-term use and update the model through transfer learning, enabling it to generate the best tuning results faster in similar environments, achieving evolutionary intelligence that "understands users better the more it is used".

[0034] VI. Summary of Innovation and Technological Progress 1. Innovation through Multimodal Perception Fusion: This invention is the first to apply the fusion of visual and acoustic data to the field of sound field tuning, breaking through the limitations of traditional "single acoustic feedback" and possessing cross-modal environmental understanding capabilities. 2. Breakthrough in Sound Field Modeling Based on Deep Fusion: By introducing graph convolutional networks to model spatial acoustic topology, accurate estimation of acoustic features in complex rooms is achieved, filling the technological gap in spatial geometric modeling of traditional tuning algorithms. 3. Intelligent Self-Tuning Through Reinforcement Learning: This invention is the first to introduce reinforcement learning mechanisms into the home audio tuning scenario, enabling the audio system to possess self-learning and self-optimization capabilities, significantly improving the system's intelligence level. 4. Unexpected Technical Effects: Experiments demonstrate that multimodal fusion can maintain stable performance even when environmental noise is high or lighting is insufficient. This effect breaks the previous technical prejudice that "visual data is difficult to assist sound field recognition in low-light environments," reflecting the outstanding substantive features of this invention.

[0035] 5. Overall performance improvement

[0036] Compared with existing automatic equalization systems, this invention achieves significant improvements in sound accuracy, spatial stability, and response speed, reaching the level of manual tuning by professional audio engineers.

[0037] In summary, this invention constructs an intelligent sound field tuning system with autonomous learning and adaptive capabilities through multimodal perception, deep feature fusion, reinforcement learning optimization, and dynamic feedback self-calibration. Its working principle is based on cross-modal collaborative perception and intelligent decision-making mechanisms, solving the long-standing technical problem of audio tuning relying on human experience and being unable to adapt to environmental changes in real time. It achieves true "auditory intelligence" and "environmental adaptation," possessing significant inventiveness and application promotion value.

[0038] Example 1

[0039] Based on a room sound field dynamic modeling and tuning method using microphone arrays and visual perception, this embodiment proposes an intelligent speaker sound field adaptive tuning method based on the fusion of microphone arrays and visual perception. This method is used to achieve real-time sound field modeling and dynamic tuning in a home living room or conference room environment, in order to solve the problems of sound image offset and uneven sound effects caused by the inability of traditional tuning systems to accurately identify spatial layout and audience position.

[0040] In this embodiment, the system includes a microphone array module, a wide-angle camera module, a spatial perception module, an acoustic modeling unit, a reinforcement learning optimization module, and a feedback adjustment unit. The microphone array is installed at the front of the smart speaker to collect sound source signals and reflected sound waves from the room; the camera module is positioned on top of or in front of the speaker to capture scene images and user distribution information; the spatial perception module uses infrared or ultrasonic sensors to detect room structure and obstacles.

[0041] When the sound system is activated, it first performs multimodal data acquisition and synchronization processing: the microphone array acquires time delay difference (TDOA) data from multiple sound sources and calculates the spatial reflection path; the camera uses deep learning algorithms (such as YOLOv8 or Vision Transformer) to identify the room boundaries and the user's position. Subsequently, the system uses a neural network fusion model (such as the multi-channel convolutional fusion network M-CNN) to fuse acoustic features with visual spatial features, outputting a high-dimensional environmental description vector.

[0042] In the sound field modeling stage, the system uses the Room Impulse Response (RIR) method to establish the indoor acoustic transfer function and uses a visual model to estimate the influence of reflective surfaces and obstacles on the sound wave propagation path, thereby realizing three-dimensional sound field reconstruction.

[0043] During the tuning phase, the reinforcement learning module receives sound field characteristics and target acoustic parameters (such as equalizer band gain, reverberation time RT60, and sound pressure level (SPL) distribution), and iteratively generates optimal tuning parameters through a reward function optimization algorithm (such as Proximal Policy Optimization, PPO). These parameters are then applied in real-time to the speaker's filtering and gain modules via a DSP (Digital Signal Processor), enabling automatic correction of the acoustic output as the space changes.

[0044] Finally, the dynamic feedback module calculates the error term (such as the difference between the target and the actual sound energy ΔE) based on the microphone's real-time sampling signal. If the deviation exceeds the threshold, the system automatically triggers another optimization and tuning process to achieve continuous adaptive updates.

[0045] This embodiment achieves high-precision identification of audience distribution and spatial structure features through the fusion of microphone array and visual information, making the tuning results independent of fixed environmental parameters; at the same time, the reinforcement learning mechanism realizes dynamic optimization, significantly improving auditory consistency in multi-user scenarios.

[0046] Example 2

[0047] This embodiment proposes a scene-driven adaptive sound tuning method that integrates spatial sensing and speech recognition. This method addresses the issue of different auditory needs of users in different scenarios (such as music playback, movie mode, conference speech, etc.) and achieves coordinated control of environmental semantic recognition and sound field optimization.

[0048] The system includes a spatial sensing unit, a voice recognition unit, a scene classification module, an acoustic optimization module, and a self-learning controller. The spatial sensing unit includes an infrared array sensor and a temperature and humidity sensing module, used to identify the room structure and air acoustic impedance characteristics; the voice recognition unit is based on the Transformer acoustic model and is used to parse user commands (such as "turn on cinema mode" or "optimize voice clarity").

[0049] When the system starts up, the spatial sensing module first collects environmental features, including room volume V, sound absorption coefficient α, and furniture positions, and generates a spatial sound field map using a 3D point cloud reconstruction algorithm (such as SLAM technology). Simultaneously, the speech recognition unit receives user voice input and analyzes the current usage scenario.

[0050] The scene classification module integrates spatial information and semantic input, and uses a hybrid convolutional network (Hybrid CNN-LSTM) to identify and predict scenes, such as identifying them as "multi-person conference scene" or "surround theater scene". The recognition results are input to the acoustic optimization module, which matches the optimal tuning template based on the built-in acoustic database.

[0051] Next, the self-learning controller fine-tunes the template parameters using a reinforcement learning mechanism: by dynamically detecting the Speech Intelligibility Index (STI) and Sound Pressure Uniformity (ΔSPL), a reward function is established and parameters are corrected. At each time period, the system evaluates changes in the sound field and automatically generates correction parameters (such as mid-frequency compensation or reverberation cutoff filtering).

[0052] To achieve closed-loop control, the system uses a monitoring microphone array at the audio output to detect the characteristics of the played audio waves in real time and perform differential comparisons with the ideal output. When the system detects a deviation of the sound field from the target response (e.g., increased reflections or low-frequency buildup), the controller immediately sends an adjustment signal to drive the DSP module to automatically adjust the output power and spectral distribution, achieving continuous adaptive optimization.

[0053] Through this embodiment, the system realizes a complete link of "semantic perception - scene classification - sound field adaptation", which can adjust the sound characteristics synchronously according to the user's intention and the actual acoustic environment, avoiding the limitations of traditional fixed EQ mode and significantly improving the consistency of the listening experience in multiple environments.

[0054] Example 3

[0055] This embodiment proposes a multi-user personalized sound field optimization method based on deep reinforcement learning, which solves the problem that existing smart speakers cannot take into account the different listening experiences caused by the different positions and preferences of different listeners in multi-user scenarios.

[0056] The system includes: a multi-user positioning module, a user feature recognition module, a sound field reconstruction module, a deep reinforcement learning optimization module, and a personalized parameter management unit. The multi-user positioning module acquires the user's three-dimensional position coordinates and head orientation through a camera and an infrared depth sensor; the user feature recognition module confirms identity and reads the preference database through facial recognition or voice fingerprint recognition.

[0057] During system initialization, the camera uses a pose estimation algorithm to identify the position and orientation information of multiple listeners, and combines this with the sound wave distribution collected by the microphone array to construct a real-time sound pressure field map. The sound field reconstruction module uses the finite difference time-domain method (FDTD) to calculate the sound energy distribution at multiple points, forming a sound field mapping matrix.

[0058] Subsequently, the deep reinforcement learning module (using the Deep Deterministic Policy Gradient Algorithm (DDPG)) establishes a mapping relationship between input states (including sound pressure distribution, reflection intensity, user distance, etc.) and output actions (filter gain, delay compensation, directivity parameters) with "user listening score prediction" as the reward objective. After multiple rounds of training, the system can autonomously learn the optimal tuning strategy that maximizes the average satisfaction of all listeners.

[0059] In actual operation, the system automatically assigns personalized weights based on different users' listening preferences (such as low-frequency enhancement or speech clarity priority), and adjusts the output phase and amplitude of each speaker unit through sound wave interference optimization algorithm, thereby forming a personalized sound field area in space.

[0060] In addition, this embodiment introduces a time-continuous feedback mechanism: the system periodically detects the user's facial expressions, head movements and active voice feedback, and evaluates real-time satisfaction through a deep emotion recognition model. When unsatisfactory signals such as "frowning" or "shaking head" are detected, the reinforcement learning module automatically triggers the tuning loop to regenerate the sound field parameters, thereby achieving "emotion-driven acoustic optimization".

[0061] This embodiment achieves truly personalized sound field adaptation by combining multimodal fusion (visual, speech, infrared), deep reinforcement learning, and emotion recognition. It not only breaks through the limitations of traditional static tuning in terms of technology, but also achieves active learning and real-time optimization at the user experience level, which has outstanding substantial progress and significant application value.

[0062] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for adaptive sound field tuning of intelligent speakers that integrates multimodal perception, characterized in that, include: Acoustic, visual, and spatial location data are collected using a microphone array, camera, and spatial sensors, respectively. Multimodal feature fusion using deep neural networks; Based on the fusion features, a sound field model is constructed and the scene type is identified; Adaptive tuning parameters are generated using reinforcement learning algorithms; Real-time sound field correction is achieved based on dynamic feedback.

2. The method according to claim 1, characterized in that, The microphone array includes at least eight MEMS microphones distributed around the speaker housing to form a spatial sound pressure distribution matrix.

3. The method according to claim 1, characterized in that, The visual information is collected by a wide-angle camera and identified through a convolutional neural network to determine the room boundaries and the distribution of obstacles.

4. The method according to claim 1, characterized in that, The spatial sensor includes an infrared or ultrasonic module for real-time detection of the listener's relative position and posture.

5. The method according to claim 1, characterized in that, The multimodal feature fusion adopts a three-branch deep network structure to process acoustic features, visual features and spatial features respectively, and outputs a unified semantic vector.

6. The method according to claim 1, characterized in that, The sound field model employs a hybrid structure of convolutional neural networks and graph convolutional networks to model the room's acoustic response function.

7. The method according to claim 1, characterized in that, The adaptive tuning parameters are generated through a reinforcement learning algorithm, and the training objective function is to minimize auditory error and optimize sound image stability.

8. The method according to claim 1, characterized in that, The system is equipped with a dynamic feedback channel, which enables self-calibration of tuning parameters through real-time sound wave monitoring and visual re-recognition.

9. The method according to claim 1, characterized in that, The method can reload new equalization parameters within 50 milliseconds after detecting environmental changes, achieving imperceptible sound field updates.

10. The method according to claim 1, characterized in that, This method can still operate stably in complex noise or low light environments, and overcomes the problem of traditional acoustic feedback failing in high-noise scenes by using visual-assisted sound field estimation.