Multi-medium sound field control system based on AI algorithm
By combining the finite-difference time domain method and graph neural network for sound field modeling, and combining it with proximal strategy optimization and differential evolution algorithm, and introducing laser Doppler vibrometer and microphone array for feedback evaluation, the problem of insufficient generalization ability of multi-media sound field control system in complex environments is solved, and higher control accuracy and stability are achieved.
Patent Information
- Application Number
- CN202510798950.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-06-16
AI Technical Summary
Existing multi-media sound field control systems based on AI algorithms have poor generalization capabilities in complex or unknown media environments and cannot effectively respond to changes in the sound field, resulting in unsatisfactory control effects.
The finite difference time domain method combined with graph neural network is used for sound field modeling, and the proximal strategy optimization and differential evolution algorithm are combined for control. Laser Doppler vibrometer and microphone array are introduced for feedback evaluation to optimize the sound field control parameters.
The accuracy of sound field modeling and the adaptability of control strategies are improved, the generalization ability and control accuracy of the system in different media environments are enhanced, and the robustness and stability of sound field control are improved.
Smart Images

Figure CN120640202A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligence, and more specifically, to a multi-media sound field control system based on AI algorithm. Background Art
[0002] Multi-media sound field manipulation systems utilize the propagation characteristics of sound waves in different media and are widely used in fields such as acoustics, communications, medicine, and environmental monitoring. With the rapid development of artificial intelligence (AI) algorithms, traditional sound field control methods have gradually evolved towards intelligent and automated approaches. AI-based multi-media sound field manipulation systems utilize advanced technologies such as deep learning and machine learning to precisely control the sound field by monitoring, analyzing, and optimizing the propagation paths and effects of sound waves in multiple media in real time. The development of this system's background technology stems from in-depth research on the propagation characteristics of sound waves in the acoustic field, and has gradually been applied to the manipulation of sound waves in various media (such as air, water, and solid materials). Traditional sound field manipulation relies heavily on physical models and empirical design, which have limitations. The introduction of AI algorithms overcomes these limitations. Through big data analysis and model training, sound field manipulation systems can automatically adapt to complex environments and precisely adjust parameters such as sound wave frequency, intensity, and direction.
[0003] In existing technologies, many AI algorithms rely on large amounts of experimental data for training, but this data is often limited by the experimental environment and lacks sufficient diversity. As a result, the AI models have poor generalization capabilities and may not be able to effectively respond to changes in the sound field in complex or unknown media environments, resulting in suboptimal control effects. Although AI algorithms can perform well under certain fixed conditions, in actual applications, due to changes in the media environment (such as changes in temperature, pressure, humidity, and other factors), the system's adaptability is insufficient. Existing technologies have not yet achieved sufficiently intelligent dynamic adjustment capabilities, and may not be able to maintain optimal sound field control effects under changing environmental conditions. Summary of the Invention
[0004] The purpose of the present invention is to provide a multi-media sound field control system based on an AI algorithm to solve the problems raised in the above-mentioned background technology: In the prior art, multiple AI algorithms rely on a large amount of experimental data for training, and these data are often limited by the experimental environment and lack sufficient diversity. Therefore, the generalization ability of the AI model is poor, and in a complex or unknown medium environment, it may not be able to effectively respond to changes in the sound field, resulting in unsatisfactory control effects. Although the AI algorithm can perform well under certain fixed conditions, in actual applications, due to changes in the medium environment (such as changes in temperature, pressure, humidity and other factors), the system's adaptability is insufficient. The existing technology has not yet achieved sufficiently intelligent dynamic adjustment capabilities, and may not be able to maintain the best sound field control effect under changing environmental conditions.
[0005] Technical solution: The multi-media sound field control system based on AI algorithm includes a data acquisition module, a sound field modeling module, an AI optimization control module, an adaptive adjustment module, a hardware execution module and a feedback evaluation module; Among them, the data acquisition module acquires acoustic signals in air media, water media, and solid media based on multi-channel synchronous sampling; the sound field modeling module uses the finite difference time domain method combined with the graph neural network to calculate the sound wave propagation characteristics; the AI optimization control module combines the proximal strategy optimization algorithm and the differential evolution algorithm to calculate the sound field control parameters; the adaptive adjustment module uses the model migration method to optimize the weight parameters of the AI optimization control module; the hardware execution module excites the target sound field based on the piezoelectric transducer array; the feedback evaluation module uses a laser Doppler vibrometer and a microphone array to obtain sound field measurement data and correct the control parameters.
[0006] Preferably, the sound field modeling module includes a physical modeling unit and a data-driven modeling unit; The physical modeling unit calculates the sound field distribution in the inhomogeneous medium based on the finite difference time domain method; The data-driven modeling unit uses a graph neural network to learn the impact of medium characteristics on sound wave propagation; The physical modeling unit adopts an absorbing boundary condition to weaken the boundary reflection of the calculation area; The data-driven modeling unit uses a multi-layer graph convolutional network to extract medium topological features.
[0007] Preferably, in the graph convolutional network of the data-driven modeling unit, the input layer uses an adjacency matrix to characterize the spatial topological relationship of the multi-media sound field; the hidden layer uses a self-attention mechanism to adjust the propagation weights between media; the output layer calculates the sound field distribution and compares it with the finite difference time domain calculation results, and optimizes the network parameters by gradient descent.
[0008] Preferably, the self-attention mechanism of the hidden layer adopts a multi-head attention structure; The multi-head attention structure optimizes feature extraction by computing different weight matrices in parallel; The weight matrix adopts a dynamic weight update strategy based on position encoding; The dynamic weight update strategy adaptively adjusts attention allocation based on the similarity between neighboring nodes.
[0009] Preferably, the AI optimization control module includes a reinforcement learning unit and an evolutionary algorithm unit; The reinforcement learning unit uses a proximal strategy optimization algorithm to train the intelligent agent to calculate the optimal sound field control strategy; The evolutionary algorithm unit uses a differential evolution algorithm to optimize the hyperparameters of the reinforcement learning unit; The optimization objectives of the reinforcement learning unit include sound pressure level deviation, energy utilization and computational overhead; The evolutionary algorithm unit adjusts the search step size based on an adaptive mutation factor.
[0010] Preferably, the reinforcement learning unit adopts a hierarchical reward mechanism to optimize the sound field control strategy; The main reward is calculated based on the sound pressure level deviation of the target area to optimize the gradient direction; The auxiliary reward is normalized and adjusted based on the sound field energy utilization and computational complexity; The reward function adopts a dynamic weight distribution strategy; The dynamic weight allocation strategy calculates the magnitude of change between the current strategy and the historical strategy based on KL divergence.
[0011] Preferably, the reinforcement learning unit uses self-supervised learning to enhance generalization ability; Wherein, the self-supervised learning adopts contrastive learning method to construct positive and negative sample pairs; The positive sample pairs are composed of optimal sound field control parameters under similar environments; The negative sample pairs are composed of misadjusted parameters under different medium conditions; By maximizing the similarity of positive samples and minimizing the similarity of negative samples.
[0012] Preferably, the evolutionary algorithm unit uses a hierarchical differential evolution strategy to optimize the hyperparameters of the reinforcement learning unit; The hierarchical differential evolution strategy includes individual variation, population screening, and hierarchical crossover; The individual variation adaptively adjusts the variation amplitude based on the gradient information of the optimal individual; The population screening adopts an elite retention strategy to inherit the optimal hyperparameters; The hierarchical crossover searches for a Pareto optimal solution based on a multi-objective optimization strategy.
[0013] Preferably, the hierarchical differential evolution strategy is combined with a neural architecture search method to optimize the hyperparameter search space; The neural architecture search method uses a gradient-based search strategy to optimize the network structure; The search strategy optimizes the architecture via second-order gradient updates; The neural network structure of the reinforcement learning unit selects the optimal architecture among the fully connected network, convolutional network, and transformer structure based on a search strategy.
[0014] Compared with the prior art, the advantages of the present invention are: (1) This paper combines the finite difference time domain method (FDTD) and graph neural network (GNN) to optimize the accuracy of sound field modeling in multi-media environments, overcoming the problem of insufficient simulation accuracy of traditional methods in complex media.
[0015] (2) The present invention improves the adaptability of the control strategy through proximal policy optimization (PPO) and differential evolution algorithm (DE), and adopts self-supervised learning to enhance the generalization ability in different media environments, thereby improving the training efficiency and accuracy.
[0016] (3) The present invention improves the closed-loop stability and accuracy of the sound field control and enhances the robustness of the system in practical environments by introducing multi-sensor feedback of a laser Doppler vibrometer and a microphone array and combining it with error correction. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a module schematic diagram of a multi-media sound field control system based on an AI algorithm of the present invention; DETAILED DESCRIPTION
[0018] For examples, see Figure 1 ,A multi-media sound field control system based on AI algorithm includes a data acquisition module, a sound field modeling module, an AI optimization control module, an adaptive adjustment module, a hardware execution module, and a feedback evaluation module; Among them, the data acquisition module acquires acoustic signals in air media, water media, and solid media based on multi-channel synchronous sampling; the sound field modeling module uses the finite difference time domain method combined with graph neural network to calculate the sound wave propagation characteristics; the AI optimization control module combines the proximal strategy optimization algorithm and the differential evolution algorithm to calculate the sound field control parameters; the adaptive adjustment module uses the model migration method to optimize the weight parameters of the AI optimization control module; the hardware execution module excites the target sound field based on the piezoelectric transducer array; the feedback evaluation module uses a laser Doppler vibrometer and a microphone array to obtain sound field measurement data and correct the control parameters.
[0019] Specifically, the data acquisition module uses multi-channel synchronous sampling technology to collect acoustic signals from air, water, and solid media. It uses a high-precision analog-to-digital converter to digitize the tiny sound pressure signals and combines them with environmental perception sensors (temperature, humidity, and density sensors) to synchronously obtain the physical properties of the media, providing basic data for modeling. The acoustic field modeling module includes a physical modeling unit and a data-driven modeling unit. The former calculates the sound wave propagation path, reflection, and refraction effects in heterogeneous media based on the finite-difference time-domain (FDTD) method, and uses absorbing boundary conditions (PML) to weaken boundary reflections. The latter learns the topological characteristics of the medium through a graph neural network (GNN). The input layer expresses spatial connectivity using an adjacency matrix, and the hidden layer introduces a multi-head self-attention mechanism to adjust the propagation weights between media. The output layer compares the modeling results with the FDTD results and jointly optimizes the weight parameters using a gradient descent algorithm. The AI optimization control module includes a reinforcement learning unit and an evolutionary algorithm unit. The reinforcement learning unit uses the proximal policy optimization (PPO) algorithm to train the control agent. Its state space is the current sound field state, and its action space is the combination of excitation parameters (such as transducer frequency, phase, amplitude, etc.). The reward function combines the main reward (sound pressure deviation in the target area) and auxiliary rewards (energy utilization and computational overhead) and uses a dynamic weight allocation strategy controlled by KL divergence. The evolutionary algorithm unit uses the differential evolution (DE) algorithm to optimize the reinforcement learning hyperparameters and introduces a hierarchical differential evolution strategy, including individual mutation (adaptive adjustment of mutation amplitude based on gradient information), population screening (using an elite retention strategy to inherit excellent parameters), and hierarchical crossover (multi-objective optimization to obtain the Pareto optimal solution). The adaptive adjustment module optimizes AI module weights through a model transfer mechanism. When the acoustic field medium, target area, or transducer array configuration changes, the system quickly converges to a new optimal control strategy through historical model transfer learning. It also enhances generalization capabilities by combining self-supervised contrastive learning. Positive samples are derived from optimal control parameters in similar media environments, while negative samples are derived from misadjusted parameters. Transfer learning is achieved by maximizing the similarity of positive samples and minimizing the similarity of negative samples. The hardware execution module uses a piezoelectric transducer array to generate adjustable sound waves. The array layout adaptively optimizes the density and direction based on the modeling results. The drive signal is modulated in real time by a high-precision waveform controller and fine-tuned based on real-time feedback to ensure that the generated sound field covers the target area. The feedback evaluation module obtains sound field information in real time through a laser Doppler vibrometer and a three-dimensional microphone array, constructs an error evaluation function, and transmits a correction signal back to the AI control module to achieve closed-loop control. The measurement results are used to evaluate parameters such as sound pressure level, phase offset, and energy distribution, thereby further improving control accuracy.
[0020] The sound field modeling module includes a physical modeling unit and a data-driven modeling unit; The physical modeling unit calculates the acoustic field distribution in inhomogeneous media based on the finite difference time domain method; The data-driven modeling unit uses graph neural networks to learn the impact of medium characteristics on sound wave propagation; The physical modeling unit uses absorbing boundary conditions to weaken the boundary reflection of the calculation area; The data-driven modeling unit uses a multi-layer graph convolutional network to extract medium topological features.
[0021] Specifically, the sound field modeling module adopts the following steps: A1. Use the FDTD method to divide the mesh model and set the medium parameters (density, sound velocity, etc.) and boundary absorption conditions; A2. Input the initial excitation source parameters and iteratively calculate the sound field changes at each time step; A3. Build a topology model and establish the connection relationships between media nodes to form an adjacency matrix; A4. A graph neural network extracts the topological structure of the medium and calculates the sound propagation path adjustment coefficient through multi-layer graph convolution (GCN) and attention mechanism; A5. Cross-validate with FDTD calculation results and use gradient descent to optimize model parameters to improve modeling accuracy.
[0022] In the graph convolutional network of the data-driven modeling unit, the input layer uses the adjacency matrix to represent the spatial topological relationship of the multi-media sound field; the hidden layer uses the self-attention mechanism to adjust the propagation weights between media; the output layer calculates the sound field distribution and compares it with the finite difference time domain calculation results, and optimizes the network parameters through gradient descent.
[0023] Specifically, the self-attention mechanism of the graph neural network uses the following formula to adjust the weights:
[0024] in: is the node feature vector, is the weight matrix of query and key, is the dimension normalization factor, is the position encoding function.
[0025] The self-attention mechanism of the hidden layer adopts a multi-head attention structure; The multi-head attention structure optimizes feature extraction by computing different weight matrices in parallel; The weight matrix adopts a dynamic weight update strategy based on position encoding; The dynamic weight update strategy adaptively adjusts attention allocation based on the similarity between neighboring nodes.
[0026] Specifically, the reward function of the reinforcement learning unit adopts the following structure:
[0027] in: is the target sound pressure, is the actual sound pressure, is the energy utilization rate, To calculate resource consumption, is the weight factor, which is adjusted dynamically; A small constant to prevent the denominator from being zero.
[0028] The AI optimization control module includes a reinforcement learning unit and an evolutionary algorithm unit; The reinforcement learning unit uses a proximal strategy optimization algorithm to train the intelligent agent to calculate the optimal sound field control strategy; The evolutionary algorithm unit uses the differential evolution algorithm to optimize the hyperparameters of the reinforcement learning unit; The optimization objectives of the reinforcement learning unit include sound pressure level deviation, energy utilization, and computational overhead; The evolutionary algorithm unit adjusts the search step size based on the adaptive mutation factor.
[0029] Specifically, the hierarchical evolution process in the differential evolution algorithm is as follows: B1. Initialize the population and set the individual parameter vector; B2. Adaptively adjust the coefficient of variation and crossover probability according to the gradient change of the optimal individual in each iteration; B3. Use the Pareto optimal solution to score and screen candidate parameters; B4. Introduce neural architecture search and combine candidate architecture space (fully connected, convolutional, transformer) to optimize the AI control network structure and improve the strategy convergence speed.
[0030] The reinforcement learning unit uses a hierarchical reward mechanism to optimize the sound field control strategy; The main reward is calculated based on the deviation of the sound pressure level in the target area to optimize the gradient direction; The auxiliary reward is normalized and adjusted based on the sound field energy utilization and computational complexity; The reward function adopts a dynamic weight distribution strategy; The dynamic weight allocation strategy calculates the change between the current strategy and the historical strategy based on KL divergence.
[0031] Specifically, it supports multi-modal input and remote linkage operation. After the task objectives are input through the remote control platform, the system can adaptively configure the sound field control plan under different media structures, and support real-time three-dimensional sound image rendering and multi-task area sound focus control.
[0032] The reinforcement learning unit uses self-supervised learning to enhance generalization capabilities; Among them, self-supervised learning uses contrastive learning methods to construct positive and negative sample pairs; The positive sample pairs are composed of the optimal sound field control parameters in similar environments; Negative sample pairs are composed of misadjusted parameters under different medium conditions; By maximizing the similarity of positive samples and minimizing the similarity of negative samples.
[0033] Specifically, a joint loss function formula for sound field modeling and optimization control in the present invention is as follows:
[0034] in: The sound pressure values are predicted by the physical model and GNN respectively. KL represents the KL divergence of the control strategy change, is the gradient of the reward function with respect to the network parameters, is the weighting factor of the loss term.
[0035] The evolutionary algorithm unit uses a hierarchical differential evolution strategy to optimize the hyperparameters of the reinforcement learning unit; Hierarchical differential evolution strategies include individual mutation, population screening, and hierarchical crossover; Individual variation adaptively adjusts the variation amplitude based on the gradient information of the optimal individual; Population screening uses an elite retention strategy to inherit the optimal hyperparameters; Hierarchical cross-searching searches for Pareto optimal solutions based on multi-objective optimization strategies.
[0036] Specifically, a multi-medium dynamic control sound focusing index formula is:
[0037] in, is the sound pressure at the i-th sampling point; is the maximum sound pressure value; is the standard deviation of the sound pressure, is the adjustment coefficient for regulating uniformity.
[0038] Hierarchical differential evolution strategy combined with neural architecture search method to optimize hyperparameter search space; Among them, the neural architecture search method uses a gradient-based search strategy to optimize the network structure; The search strategy optimizes the architecture via second-order gradient updates; The neural network structure of the reinforcement learning unit selects the optimal architecture among the fully connected network, convolutional network, and transformer structure based on the search strategy.
[0039] The above shows and describes the basic principles, main features and advantages of the present invention; those skilled in the art should understand that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only preferred examples of the present invention and are not intended to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements all fall within the scope of the present invention to be protected; the scope of protection claimed in the present invention is defined by the attached claims and their equivalents.
Claims
1. A multi-media sound field control system based on AI algorithm, characterized in that: The multi-media sound field control system based on AI algorithm includes a data acquisition module, a sound field modeling module, an AI optimization control module, an adaptive adjustment module, a hardware execution module and a feedback evaluation module; Among them, the data acquisition module acquires acoustic signals in air media, water media, and solid media based on multi-channel synchronous sampling; the sound field modeling module uses the finite difference time domain method combined with the graph neural network to calculate the sound wave propagation characteristics; the AI optimization control module combines the proximal strategy optimization algorithm and the differential evolution algorithm to calculate the sound field control parameters; the adaptive adjustment module uses the model migration method to optimize the weight parameters of the AI optimization control module; the hardware execution module excites the target sound field based on the piezoelectric transducer array; the feedback evaluation module uses a laser Doppler vibrometer and a microphone array to obtain sound field measurement data and correct the control parameters.
2. The multi-media sound field manipulation system according to claim 1, characterized in that: The sound field modeling module includes a physical modeling unit and a data-driven modeling unit; The physical modeling unit calculates the sound field distribution in the inhomogeneous medium based on the finite difference time domain method; The data-driven modeling unit uses a graph neural network to learn the impact of medium characteristics on sound wave propagation; The physical modeling unit adopts an absorbing boundary condition to weaken the boundary reflection of the calculation area; The data-driven modeling unit uses a multi-layer graph convolutional network to extract medium topological features.
3. The multi-media sound field manipulation system according to claim 2, characterized in that: In the graph convolutional network of the data-driven modeling unit, the input layer uses an adjacency matrix to characterize the spatial topological relationship of the multi-media sound field; the hidden layer uses a self-attention mechanism to adjust the propagation weights between media; the output layer calculates the sound field distribution and compares it with the finite difference time domain calculation results, and optimizes the network parameters through gradient descent.
4. The multi-media sound field manipulation system according to claim 3, characterized in that: The self-attention mechanism of the hidden layer adopts a multi-head attention structure; The multi-head attention structure optimizes feature extraction by computing different weight matrices in parallel; The weight matrix adopts a dynamic weight update strategy based on position encoding; The dynamic weight update strategy adaptively adjusts attention allocation based on the similarity between neighboring nodes.
5. The multi-media sound field manipulation system according to claim 1, characterized in that: The AI optimization control module includes a reinforcement learning unit and an evolutionary algorithm unit; The reinforcement learning unit uses a proximal strategy optimization algorithm to train the intelligent agent to calculate the optimal sound field control strategy; The evolutionary algorithm unit uses a differential evolution algorithm to optimize the hyperparameters of the reinforcement learning unit; The optimization objectives of the reinforcement learning unit include sound pressure level deviation, energy utilization and computational overhead; The evolutionary algorithm unit adjusts the search step size based on an adaptive mutation factor.
6. The multi-media sound field manipulation system according to claim 5, characterized in that: The reinforcement learning unit uses a hierarchical reward mechanism to optimize the sound field control strategy; The main reward is calculated based on the sound pressure level deviation of the target area to optimize the gradient direction; The auxiliary reward is normalized and adjusted based on the sound field energy utilization and computational complexity; The reward function adopts a dynamic weight distribution strategy; The dynamic weight allocation strategy calculates the magnitude of change between the current strategy and the historical strategy based on KL divergence.
7. The multi-media sound field manipulation system according to claim 6, characterized in that: The reinforcement learning unit uses self-supervised learning to enhance generalization ability; The self-supervised learning adopts contrastive learning method to construct positive and negative sample pairs; The positive sample pairs are composed of optimal sound field control parameters under similar environments; The negative sample pairs are composed of misadjusted parameters under different medium conditions; By maximizing the similarity of positive samples and minimizing the similarity of negative samples.
8. The multi-media sound field manipulation system according to claim 5, characterized in that: The evolutionary algorithm unit uses a hierarchical differential evolution strategy to optimize the hyperparameters of the reinforcement learning unit; The hierarchical differential evolution strategy includes individual variation, population screening, and hierarchical crossover; The individual variation adaptively adjusts the variation amplitude based on the gradient information of the optimal individual; The population screening adopts an elite retention strategy to inherit the optimal hyperparameters; The hierarchical crossover searches for a Pareto optimal solution based on a multi-objective optimization strategy.
9. The multi-media sound field manipulation system according to claim 8, characterized in that: The hierarchical differential evolution strategy is combined with a neural architecture search method to optimize the hyperparameter search space; The neural architecture search method uses a gradient-based search strategy to optimize the network structure; The search strategy optimizes the architecture via second-order gradient updates; The neural network structure of the reinforcement learning unit selects the optimal architecture among the fully connected network, convolutional network, and transformer structure based on a search strategy.
Citation Information
Patent Citations
Self-adaptive sound field regulation and control method and device
CN113099356A
Method for optimizing parameters of audio output device, electronic device and storage medium
CN117880702A
AI scene construction omnibearing system and method for psychological adjuvant therapy
CN118236601A
Natural environment bird monitoring method based on multi-modal fusion deep learning and computer device
CN119027775A
Echo suppression method, device and system of spherical screen display screen system
CN119323952A
Cited By
Microphone with adaptive pickup module and pickup control method
CN120343477A
Intelligent sound field adaptive system of digital professional sound equipment
CN121711603A