Self-adaptive sound field control and privacy protection method and system based on semantic driving

By using real-time RIR estimation and dynamic adjustment of MPC weights, combined with beamforming and noise masking, the problem of sound field control and privacy protection in public address systems in complex public spaces is solved, achieving high-precision voice transmission and privacy protection, and improving the system's environmental adaptability and security.

CN121968007APending Publication Date: 2026-05-01LINKER
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LINKER
Filing Date
2025-12-30
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing public address systems suffer from a lack of environmental adaptability in complex public spaces, rigid control strategies, severe multi-area interference, and difficulty in achieving privacy protection. In particular, the sound field coverage is uneven under the influence of changes in crowd density and temperature, failing to meet the needs of semantic understanding and privacy protection.

Method used

By estimating the room impulse response (RIR) in real time using a distributed microphone array, and combining user identity and scene semantic features, the model predictive control (MPC) weights are dynamically adjusted. A combination of beamforming and noise masking is used to achieve adaptive sound field control and privacy protection.

Benefits of technology

It achieves high-precision sound field control in complex environments, ensuring voice clarity and privacy protection, reducing crosstalk interference, improving system flexibility and security, and meeting the needs of safe broadcasting in emergency situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121968007A_ABST
    Figure CN121968007A_ABST
Patent Text Reader

Abstract

The invention discloses a self-adaptive sound field control and privacy protection method and system based on semantic driving. The method comprises the following steps: firstly, estimating environment impulse response and inverting equivalent sound velocity in real time through a distributed microphone array, and generating scene semantic features in combination with user identity and position; thirdly, mapping the semantic features into semantic state tags, dynamically adjusting weights and constraint conditions in a model predictive control cost function according to the semantic state tags, and realizing self-adaptive switching of a control strategy among a privacy protection mode, an emergency priority mode and a conventional equilibrium mode; and finally, driving a loudspeaker array based on the solved parameters, performing beam forming on the target area, forming beam null for the adjacent potential eavesdropping area, and directionally playing masking noise based on residual speech spectrum synthesis. Through the closed-loop environment perception and psychoacoustic masking technology, the problems that the open space sound field control strategy is rigid and privacy protection is difficult are solved, and the environment adaptability and safety of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of acoustic signal processing and intelligent control technology, and in particular to a semantically driven adaptive sound field control and privacy protection method and system for use in complex public spaces such as airports, train stations, and large stadiums. Background Technology

[0002] Public address systems are an indispensable infrastructure in public places such as airports, train stations, and large conference centers. As people's demands for auditory experience and privacy protection increase, traditional unified broadcasting across the entire area or simple physical zone broadcasting can no longer meet the increasingly complex needs of various scenarios.

[0003] The existing public address system suffers from the following four major pain points: 1. Coarse sound field coverage and lack of environmental adaptability: Traditional systems are typically based on ideal acoustic models (such as assuming a constant sound speed of 343 m / s and ignoring sound absorption by crowds) for parameter settings. However, in real-world scenarios, changes in crowd density drastically alter the sound absorption coefficient of the space (humans are strong sound absorbers), and temperature changes affect the speed of sound wave propagation. This open-loop control method results in severe high-frequency attenuation and unclear intelligibility when there are many people, and echoing and howling when there are few people, making it impossible to guarantee speech intelligibility.

[0004] 2. Lack of semantic understanding and rigid control strategies: The control logic of existing systems is usually static or based solely on simple timed tasks. The system cannot perceive the current business semantics (such as the difference between "VIPs are having a private conversation" and "emergency fire alarm"), and often uses fixed control weights, failing to make optimal decisions for the current scenario regarding volume comfort, energy consumption, and interference resistance.

[0005] 3. Severe multi-region interference: When different content is played simultaneously in multiple adjacent regions, sound wave superposition interference is likely to occur. Although existing beamforming technology can target the target region, it often ignores the suppression of non-target regions, resulting in severe crosstalk.

[0006] 4. Physical Limitations of Privacy Protection Technologies: Achieving voice privacy protection in open spaces is an industry-wide challenge. Some existing technologies attempt to create silent zones by playing inverse sound waves (destructive interference) using active noise cancellation principles. However, physical acoustics principles show that in open, large-scale public spaces, due to the complex reflection paths, achieving stable broadband destructive interference at the boundaries using only a few speakers is almost impossible, and may even result in greater noise at certain locations due to comb filtering effects.

[0007] Therefore, there is an urgent need for an intelligent sound field control system that can sense environmental parameters (such as RIR) and business semantics in real time, dynamically adjust control strategies based on semantics, and achieve effective privacy protection through psychoacoustic masking rather than simple physical cancellation. Summary of the Invention

[0008] This invention primarily addresses the technical problems of existing technologies, such as the lack of environmental adaptability in sound field control, rigid multi-objective optimization strategies, and difficulty in achieving privacy protection in open spaces. It provides a semantically driven adaptive sound field control and privacy protection method and system.

[0009] The present invention addresses the aforementioned technical problems primarily through the following technical solution: a semantically driven adaptive sound field control and privacy protection method, applied to a public space equipped with several distributed loudspeakers and several distributed microphone arrays, comprising the following steps: S1: Acquire environmental and user status data: Acoustic signals are collected through the distributed microphone arrays to estimate the room impulse response and environmental acoustic parameters of the public space in real time. The room impulse response is RIR. At the same time, user identification information and location coordinates are acquired. Based on the identification information, users are divided into a first category of users who require privacy protection and a second category of users who do not require privacy protection. Scene semantic feature vectors are generated by combining crowd density and the urgency of broadcast content. S2: Generate semantic-driven control strategy: Map the scene semantic feature vector to the semantic state label of the current system operation, the label includes at least privacy protection state and normal state; Based on the semantic state label, dynamically adjust the weight coefficients and constraints in the model predictive control (MPC) cost function; The cost function includes at least a weighting term for zone volume deviation, speaker power consumption, and cross-zone interference energy; when in privacy protection mode, the weighting coefficient of the cross-zone interference energy is increased compared to the normal mode. S3: Calculate adaptive sound field optimization parameters: Based on the adjusted cost function and constraints, solve for the gain, time delay and phase control parameters of each speaker; and calculate the inverse filter according to the frequency response characteristics of the RIR inversion to perform pre-distortion compensation on the audio signal to be played. S4: Perform sound field reconstruction and privacy protection: drive the distributed broadcast loudspeakers to form a main lobe using a beamforming algorithm for the area where the first type of user is located, and form beam nulls for adjacent areas; at the same time, synthesize masking noise and play the masking noise directionally to the adjacent areas through loudspeakers located at the edge.

[0010] Traditional sound field control is often open-loop, meaning that preset parameters remain unchanged. This invention, however, introduces a closed-loop feedback mechanism. The room impulse response (RIR) is a core parameter describing the acoustic characteristics of the sound field. By estimating the RIR in real time, the system can understand the reverberation time (RT60) and frequency response characteristics of the current space. The scene semantic feature vector (Ft) is key to abstracting physical world parameters (such as the number of people, noise level) into computer-understandable "semantics," typically containing dimensions such as: [crowd density D, background noise N, user identity V, urgency level T]. This transformation from physical quantities to semantic quantities is the foundation for subsequent intelligent control.

[0011] Existing MPC algorithms typically use fixed weights, which cannot adapt to changing needs. This invention establishes a "semantic-weight" mapping mechanism. For example, when the semantics are identified as a privacy protection state, the system considers preventing sound from reaching the next room more important than the uniformity of the current area's volume. Therefore, the control algorithm automatically increases the weight ω3 of the cross-area interference term (e.g., to 10 times that of the normal state), forcing the optimization result to form an extremely steep sound pressure gradient at the boundary, even if this sacrifices some energy consumption or smoothness.

[0012] To address the issue of high-frequency loss caused by sound absorption in crowds, this invention does not simply increase the volume, but rather performs pre-distortion. If RIR analysis shows that the ambient frequency band above 4kHz attenuates too rapidly, the system will pre-emphasize this frequency band using an inverse filter Hinv(f) before sending the audio. Furthermore, preferably, the solution process also incorporates trajectory prediction. By predicting the user's position at several future moments (e.g., 5 seconds in the future), a predictive field-of-view constraint is introduced into the MPC, allowing the sound field focus to smoothly transition with the user's movement, avoiding sudden changes in volume.

[0013] Unlike existing technologies that attempt to eliminate sound, this invention employs a combination of beam nulling and acoustic masking.

[0014] First, using the MVDR algorithm, it is necessary not only to make the target area audible (main lobe), but also to force the interference area (potential eavesdropping area) to be inaudible (zero trap). Preferably, the zero trap depth is set to reduce the leakage by at least 15dB.

[0015] Secondly, for the faint speech remaining after the zero trap, the system no longer attempts to cancel it out, but instead plays a specially designed masking noise to cover it up.

[0016] Preferably, in step S1, the step of acquiring environmental parameters and estimating RIR in real time adopts a dual-mode excitation mechanism: Based on the signal transmission characteristics between the microphone array and the preset anchor point speaker, the room impulse response (RIR) of the public space is estimated in real time using the currently playing audio signal or the generated test signal as the excitation source; wherein, the system broadcast status is determined, and if it is in a silent state for more than a preset duration threshold, the speaker is driven to emit a test signal that is inaudible to the human ear, and the RIR is estimated using the echo of the test signal; The equivalent sound velocity of the current sound propagation medium is inverted from the RIR and used to correct the calculation of the user's position coordinates.

[0017] When a broadcast system is not playing content (quiet period), traditional adaptive algorithms will stop updating due to the lack of excitation signals, potentially causing parameter mismatch in the first broadcast after the quiet period ends. This invention creatively introduces a pilot mechanism in a non-audible frequency band. The inaudible test signal is preferably a low-energy sweep signal or white noise in the 18kHz-20kHz range. Because this frequency band is outside the range of human hearing but within the microphone's response range, the system can continuously detect the environment without disturbing the user. Simultaneously, by measuring the transmission time from a speaker to the microphone at a fixed distance, the current equivalent speed of sound, v, can be inferred. eff (For example, increased temperature leads to a faster speed of sound), thus correcting the TDOA positioning formula Δt=Δd / v eff This significantly improves positioning accuracy.

[0018] Preferably, in step S1, the method for obtaining the user's identity information includes: Identity and permissions can be obtained by reading the user's portable location tag; or by extracting the user's voice signal to establish a temporary acoustic feature mapping, and deregistering the feature after the user leaves the specific area.

[0019] To balance personalized service and privacy protection, this invention abandons the approach of establishing a permanent voiceprint database. Portable positioning tags (such as RFID or UWB tags) are preferred, as this is the most secure method of physical identification. If voiceprint recognition must be used, this invention emphasizes temporariness and regionality; that is, voiceprint features are extracted only when a user enters a specific service area (such as a VIP room) for current beam tracking, and the data is destroyed immediately once the user leaves, thus achieving strict privacy protection.

[0020] Preferably, in step S2, the MPC cost function J(k) is expressed as: J(k)=Σ M i=1 [ω1(S)·(L i (k)-L target ) 2 +ω2(S)·P i(k)+ω2(S)·I total (k)]; Where M is the total number of partitions, Li(k) is the actual volume of the i-th partition at time k, Ltarget is the target volume, Pi(k) is the power consumption of the speaker group serving the i-th partition at time k, Itotal(k) represents the total sound energy leaked from the interference source area (such as the waiting area) to the sensitive area (such as the VIP area) at time k, and ω1(S), ω2(S) and ω3(S) are the weight coefficients that change with the semantic state label S.

[0021] This variable structure control allows a system to exhibit multiple characteristics.

[0022] Preferably, step S2 further includes priority arbitration logic: The semantic state label also includes an emergency priority state; When an emergency priority state is determined, the optimization process of the MPC cost function is bypassed, the privacy protection strategy is lifted, and the gain of the speaker is directly adjusted to the maximum output value.

[0023] In extreme situations such as fires and earthquakes, complex algorithm calculations may introduce unnecessary delays, and privacy protection becomes less important. This invention employs a high-priority circuit breaker mechanism. Upon detecting an emergency semantic (such as a fire alarm signal), the system immediately switches from intelligent mode to direct-through mode, ensuring that the alarm sound can cover the entire area at maximum sound pressure level, complying with public safety regulations.

[0024] Preferably, in step S3, the process of solving for each speaker parameter further includes: Predict the user's movement trajectory sequence at several future moments; During MPC optimization, the cumulative cost within the predicted field of view corresponding to the motion trajectory sequence is calculated to achieve a smooth transition of the sound field focus.

[0025] Preferably, in step S4, the synthesis and playback of the masking noise specifically includes: Analyze the spectral envelope of the residual speech signal reaching the adjacent region; Generate a masking noise signal whose power spectral density matches the spectral envelope; The Speech Transmission Index (STI) is calculated based on the signal-to-noise ratio of the masking noise signal and the residual speech signal. The emission intensity of the masking noise is then adjusted so that the STI value remains below a preset intelligibility threshold.

[0026] Masking noise is not simple white noise, but rather correlated noise dynamically synthesized based on the spectral envelope of the residual speech signal (such as Babble Noise, or the noise of multiple people). This noise perfectly covers the energy concentration points of the leaked speech in the frequency domain, yet sounds like natural background noise. The preset intelligibility threshold is STI less than 0.3. This ensures privacy (unable to hear the content clearly) while preventing excessive masking noise from interfering with the environment.

[0027] Preferably, the beamforming algorithm is the minimum variance distortionless response (MVDR) algorithm, and the depth of the beam null is set such that the signal energy attenuation in the adjacent region is reduced by at least 15 dB relative to the main lobe.

[0028] Preferably, the plurality of distributed broadcast loudspeakers are logically and dynamically divided into target sound field playback units and edge interference units, and the same physical loudspeaker dynamically switches between the roles of the playback unit and the interference unit based on its relative position to the user.

[0029] This system does not require the deployment of dedicated jamming speakers. Each speaker in the system is generic, and its role is defined by software. When a user moves, a speaker that was originally in the center playing the broadcast may become a speaker in a peripheral position playing masking noise after the user moves away. This software-defined hardware design greatly reduces deployment costs and improves system flexibility.

[0030] A semantically driven adaptive sound field control and privacy protection system is applied to a public space equipped with several distributed loudspeakers and several distributed microphone arrays, comprising: The environment and user perception module is configured to acquire acoustic signals through the microphone array to estimate the room impulse response (RIR) and environmental parameters, and to obtain the user's identity information and location coordinates to generate a scene semantic feature vector; The semantic-driven strategy generation module is configured to map the scene semantic feature vector to a semantic state label, and dynamically adjust the weight coefficients and constraints in the model predictive control (MPC) cost function based on the label. The adaptive calculation module is configured to solve for the speaker control parameters based on the adjusted cost function and to calculate the inverse filter for audio predistortion compensation. The sound field execution module is configured to drive the loudspeaker to perform beamforming on the first type of user area, and to create beam nulls and directional playback masking noise on adjacent areas.

[0031] The substantial effects of this invention are: 1. Solved the problem of rigid control strategies: By introducing scene semantic features and variable structure MPC, this invention can dynamically adjust the weights and constraints of the algorithm according to business needs (such as privacy, urgency, energy saving), realizing the leap from passive playback to active cognitive control.

[0032] 2. Breakthrough in the physical bottleneck of privacy protection in open spaces: This invention innovatively proposes a psychoacoustic scheme of beam nulling + dynamic masking. By introducing the STI index into the control closed loop, it achieves content-incomprehensible privacy protection with minimal noise cost, and has extremely high practical value in large open public spaces.

[0033] 3. Improved environmental adaptability and robustness: By utilizing pilot frequencies during quiet periods and echo feedback to update RIR and sound velocity in real time, this invention can automatically compensate for the sound absorption effect caused by changes in crowd density and the positioning error caused by temperature changes, ensuring high-precision sound field control even in complex dynamic environments.

[0034] 4. Public safety is guaranteed: The highest priority emergency arbitration logic is set up to ensure that the system can switch to pass-through mode instantly in emergency situations, which complies with security standards. Attached Figure Description

[0035] Figure 1 This is a flowchart of a semantically driven adaptive sound field control and privacy protection method according to the present invention. Detailed Implementation

[0036] The technical solution of the present invention will be further described in detail below through embodiments and in conjunction with the accompanying drawings.

[0037] Example: This invention discloses a semantically driven adaptive sound field control and privacy protection system, applied to a large public space (such as the waiting hall of Terminal 3 at an airport). The system mainly includes a perception layer, a computation layer, and an execution layer.

[0038] Perception layer: Distributed microphone array: Several omnidirectional microphone arrays (one set deployed per 500m²) are evenly distributed on the ceiling. The model is Yamaha ADECIA DM3, with a sampling rate of 48kHz and a bit depth of 24bit.

[0039] Location tag readers: Deploy UWB or RFID readers in key areas to read portable location tags carried by users.

[0040] Computation layer: The AI ​​control server, model Huawei Atlas 900 PoD, is used to run the sound field control algorithm of this invention.

[0041] Execution layer: Speaker array: Composed of several independently controllable smart speakers, model Bosch LBC 3480. It is worth noting that the speakers in this system are physically interchangeable, and their logical roles (broadcast units or interference units) are dynamically assigned by an algorithm based on the user's location.

[0042] A semantically driven adaptive sound field control and privacy protection method, such as Figure 1 As shown, it includes the following steps: S1: Obtain environmental and user status data This step aims to build the system's full-dimensional perception capability.

[0043] 1. Closed-loop environmental parameter inversion The real-time RIR estimation step is a continuous adaptive system identification process. To ensure accurate acquisition of environmental parameters at any time (regardless of whether there is a broadcast task), this system employs a dual-mode excitation mechanism: Service signal excitation mode (default): When the system is in normal broadcast state, the currently playing audio stream is used as a reference signal, and the signal collected by the microphone is subjected to NLMS adaptive filtering processing to calculate the RIR in real time.

[0044] Dedicated pilot excitation mode (silence compensation): When the system detects that it has been in a broadcast silence state for more than a preset time (such as 5 minutes), in order to prevent the environmental parameters from expiring, the system automatically drives the speaker to emit a test signal that is inaudible to the human ear (preferably a sweep signal of 18kHz-20kHz or band-limited white noise) as a reference signal, and uses the echo of the test signal to perform the RIR estimation.

[0045] In this embodiment, the RIR can be calculated using an adaptive system identification method well known to those skilled in the art.

[0046] Preferably, the frequency domain block normalized least mean square (FDAF-NLMS) algorithm is used to reduce computational complexity and improve convergence speed. Its core iterative formula is: ; W represents the filter coefficients, i.e., the RIR estimate; E represents the error signal; and X represents the reference signal.

[0047] It should be noted that this invention is not limited to a specific adaptive algorithm; algorithms such as RLS and APA can also be applied. The core of this invention lies in constructing the following dual-mode stimulus mechanism: When there is a broadcast task, the broadcast audio stream is used directly as the reference signal X. k ; During the silent period, an 18kHz pilot signal is automatically generated as the reference signal X.k .

[0048] Based on the estimated RIR, the system further extracts: Equivalent speed of sound v eff By measuring the arrival time of the direct wave peak, the current sound velocity in the medium is calculated, which is used to correct the TDOA positioning formula Δt=Δd / v eff .

[0049] Frequency response decay curve H env (f): Analyze the frequency domain characteristics of RIR through FFT transformation. If the high-frequency attenuation is significant, it is determined that the sound absorption effect of the current population is enhanced.

[0050] 2. Privacy-compliant user identification To mitigate biometric privacy risks, this invention employs portable location tags (such as RFID tags integrated into electronic boarding passes or UWB tags for VIP cards) to identify users. The system reads the tag ID, queries the permission list, and categorizes users into two types: Category 1 users (requiring privacy protection, such as VIPs) and Category 2 users (not requiring privacy protection).

[0051] If voiceprint recognition is used, the system will only establish a temporary acoustic feature mapping when the user enters a specific area, and will immediately cancel and delete the feature after the user leaves the area (based on location determination), thus achieving self-destructing messages.

[0052] 3. Semantic Feature Vector Generation The system collects the above data and generates a scene semantic feature vector F. t =[D,N,V,T]. Where D is the crowd density, N is the background noise sound pressure level, V is the user type flag, and T is the urgency of the broadcast content.

[0053] S2: Generative semantic-driven control strategy 1. Semantic State Classification Model The system uses a pre-built multi-layer decision tree model or logical rule base to process vector F. t Mapped to semantic state label S: First-level judgment (safety layer): Detects the urgency level (T) in the feature vector. If T > T thresh (e.g., 0.9), for example, T=1 (fire alarm), directly output S=S2 (emergency priority state), without further judgment. At this time, the priority arbitration logic is triggered, bypassing subsequent MPC optimization, directly locking all speaker gains to the maximum value, and disabling privacy function to ensure safe broadcast coverage.

[0054] Second-level judgment (privacy layer): If the first level fails, check V (user). If V=1 (including the first type of user) and D (crowd density)<0.8 (not extremely crowded), output S=S1 (privacy protection status).

[0055] The third level of judgment (normal layer or equalization layer): If none of the above are hit, the judgment is based on N (background noise). If N>70dB, the output S is in noise reduction mode (preferring sharpness); otherwise, the output S=S3 (normal equalization state).

[0056] The advantage of using the above explicit rules is that the computation latency is extremely low (<1ms) and it is interpretable, avoiding the unpredictability of black box models.

[0057] 2. Variable Structure MPC Cost Function The system dynamically adjusts the weights of the MPC cost function J(k) based on the state S. The formula is as follows: J(k)=Σ M i=1 [ω1(S)·(L i (k)-L target ) 2 +ω2(S)·P i (k)+ω2(S·Σ j≠i I i,j (k)]; Parameter explanation: i and j are the indices of the partitions; M is the total number of partitions. For example, if the public space is divided into three logical areas: a waiting area, a VIP room, and a dining area, then M=3. The MPC algorithm will simultaneously perform joint optimization of the sound field in these three areas; k represents the discrete time step. Since MPC is a computer-controlled algorithm, it operates on a time-slice basis. k represents the current moment, and k+1 represents the next control cycle; L i (k) represents the predicted volume (sound pressure level SPL) of the i-th zone, i.e., the sound pressure level predicted based on the current speaker gain to reach the center of that zone. i (k)-L target ) 2 This represents the volume deviation term for the i-th partition. The mean squared error form is used here to more severely penalize larger volume deviations, ensuring stable volume; P i (k) represents the speaker power consumption (energy consumption term) of the i-th partition at time k. This is the control input cost of the system, which is proportional to the square of the speaker gain. Reducing this term helps save energy and prevents the system from outputting excessive current in pursuit of minute volume accuracy; Σ j≠i I i,j (k) represents the total energy term of cross-regional interference, where I i,j(k) specifically refers to the acoustic energy leaked from the i-th partition (as the source) to the j-th partition (as the object of disturbance). This term is calculated by using the sound field transfer matrix to sum the energy generated by the non-target loudspeaker in the target area. Summing this term means that the system should minimize global crosstalk interference. ω1(S), ω2(S), and ω3(S) are weighting coefficients that vary with the semantic state label S.

[0058] Σ j≠i I i,j (k) (i.e., I) total (k) is not directly obtained from microphone measurements, but is a prediction based on a sound field transfer model. The system pre-stores or updates a transfer matrix M in real time, where the elements M x,y Σ represents the sound attenuation coefficient from the x-th loudspeaker to the center of the y-th zone. j≠i I i,j (k) is calculated as: the sum of the remaining energy reaching the center of the target zone after the current gain of all speakers in the non-target zone is attenuated by the transfer matrix M.

[0059] By minimizing this term in the cost function and combining it with semantically driven weights ω3(S), the algorithm can automatically search for a set of optimal beamforming coefficients, enabling the loudspeaker to form energy nulls in adjacent sensitive areas while covering the local area, thereby achieving mathematical anti-interference optimization.

[0060] It should be noted that the above cost function formula is a preferred embodiment of the present invention. In practical applications, those skilled in the art can adjust the specific form of the cost function according to specific needs, and this should also fall within the protection scope of the present invention. For example: Error term form: The volume deviation term is not limited to the square form (L−L) target ) 2 (L2 norm), can also be expressed in absolute value form |L−L target | (L1 norm) to reduce sensitivity to outliers.

[0061] Interference term form: Cross-regional interference terms are not limited to the algebraic sum of all interferences; they can also be expressed in the form max(I) i,j This means focusing on eliminating the impact of the area with the most severe interference.

[0062] Constraint extension: Other penalty terms can be added to the cost function as needed, such as a penalty term ΔL for the rate of change of volume. 2 This is to limit the drastic adjustment of volume.

[0063] Weighting adjustment strategy: In the normal state (S3), the weight configuration is balanced (e.g., 0.6:0.2:0.2), and the update step size is widened to 200ms to save energy.

[0064] In the privacy-preserving state (S1), the cross-regional interference weight ω3 is significantly increased (e.g., increased to 2.0). This means that the algorithm aims to reduce the interference term I. i,j It will sacrifice energy consumption P i The partial volume flatness forces beamforming to produce extremely strong directivity.

[0065] S3: Calculate adaptive sound field optimization parameters The system uses gradient descent or sequential quadratic programming (SQP) to solve the above cost function and obtain the gain G of each loudspeaker. i Delay τ i and phase ϕ i The specific processing procedure is as follows: (1) Initialization: The control parameters [G] at the previous moment k−1 ,ϕ k−1 [WarmStart] is used as the initial value for the current optimization to reduce the number of iterations.

[0066] (2) Iteration: Calculate the gradient ∇J of J(k) with respect to each control variable, update the parameters along the negative gradient direction until the cost function decreases by less than the preset threshold or reaches the maximum number of iterations (e.g., 10 times).

[0067] (3) Constraint handling: If the updated parameters exceed the hardware constraints (e.g., gain > 12dB), the projection method is used to truncate them to the boundary value.

[0068] At the same time, using H obtained from step S1 env (f) Calculate the inverse filter H inv (f)=1 / H env (f). Before playing the audio signal, perform pre-distortion compensation to amplify the high-frequency components that are easily absorbed by the environment in advance to ensure clear listening.

[0069] In addition, the solution process also incorporates user trajectories predicted by Kalman filtering, introducing a prediction field of view (such as the next 5 seconds) into MPC to achieve a smooth transition of the sound field focus.

[0070] The user trajectory prediction module establishes the following motion state equation: State vector X k =[x,y,v x ,v y ] T Where (x,y) is the position, (v x ,v y () represents speed.

[0071] Assuming the user moves at a constant speed for a short period of time, the state transition matrix A is set as follows: X k+1 =AX k +w k ; The system uses the location calculated by TDOA as the observation value Z. k The standard prediction-update step of Kalman filtering is used to smooth measurement noise and extrapolate the future N. p The position coordinates at each moment.

[0072] S4: Sound Field Execution and Privacy Protection For the areas where the first type of users are located, the system adopts a dual protection mechanism of beam nulling and dynamic masking.

[0073] 1. Beam null (physical layer suppression) The system drives the playback unit in the target area to perform MVDR beamforming, with constraints including the formation of deep nulls in the direction of adjacent potential eavesdropping areas. In this embodiment, the null depth is set such that the leakage signal energy is attenuated by at least 15 dB.

[0074] 2. Dynamically related masking (psychoacoustic inhibition) To address the faint leaked voice that remains after the zero trap, the system drives the interference unit located at the edge to play masking noise.

[0075] The dynamic synthesis process of masking noise specifically includes: (1) Acquisition of basic sources: The system has multiple noisy sound substrates of different languages ​​and genders, and these substrate signals have flat long-term average spectrum.

[0076] (2) Feature extraction: extracting the leaked speech X from adjacent areas as monitored in real time. leak (t) Perform a short-time Fourier transform (STFT) to extract its spectral envelope E. leak (f) (reflects the position of the resonance peak).

[0077] (3) Shaping filter: using the extracted E leak (f) Design a time-varying filter H shape (f).

[0078] (4) Synthesis: The selected multi-person noisy sound substrate is passed through the filter so that the energy of the output noise is concentrated in the frequency band where the leakage speech energy is strongest (such as 500Hz-3kHz).

[0079] STI closed-loop control: The system calculates the Speech Transmission Index (STI) at the boundary and dynamically adjusts the emission intensity of the masking noise, aiming to suppress the STI value to below 0.3 (the incomprehensible threshold). Due to the high noise spectrum matching degree, only a small masking volume is needed to achieve the incomprehensible effect, avoiding noise pollution.

[0080] The system's performance in real-world scenarios has shown that: Positioning accuracy: Combined with sound velocity calibration during the silent period, the positioning error is stable at ±0.4m.

[0081] Interference control: Interference attenuation between adjacent zones reaches 32dB.

[0082] Privacy protection: The accuracy of voice content recognition at 1m outside the privacy zone drops to 8%, and the STI value remains stable below 0.3.

[0083] The specific embodiments described herein are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains may make various modifications or additions to the described specific embodiments or use similar methods to substitute them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.

[0084] Although this document uses various terms extensively, the possibility of using other terms is not excluded. These terms are used merely for the convenience of describing and explaining the essence of the invention; interpreting them as any additional limitation would contradict the spirit of the invention.

Claims

1. A semantically driven adaptive sound field control and privacy protection method, applied to a public space equipped with several distributed loudspeakers and several distributed microphone arrays, characterized in that, Includes the following steps: S1: Acquire environmental and user status data: Acoustic signals are collected through the distributed microphone arrays to estimate the room impulse response and environmental acoustic parameters of the public space in real time. The room impulse response is RIR. At the same time, user identification information and location coordinates are acquired. Based on the identification information, users are divided into a first category of users who require privacy protection and a second category of users who do not require privacy protection. Scene semantic feature vectors are generated by combining crowd density and the urgency of broadcast content. S2: Generate semantic-driven control strategy: Map the scene semantic feature vector to the semantic state label of the current system operation, the label includes at least privacy protection state and normal state; Based on the semantic state label, dynamically adjust the weight coefficients and constraints in the model prediction control, i.e., MPC cost function; The cost function includes at least a weighting term for zone volume deviation, speaker power consumption, and cross-zone interference energy; when in privacy protection mode, the weighting coefficient of the cross-zone interference energy is increased compared to the normal mode. S3: Calculate adaptive sound field optimization parameters: Based on the adjusted cost function and constraints, solve for the gain, time delay and phase control parameters of each speaker; and calculate the inverse filter according to the frequency response characteristics of the RIR inversion to perform pre-distortion compensation on the audio signal to be played. S4: Perform sound field reconstruction and privacy protection: drive the distributed broadcast loudspeakers to form a main lobe using a beamforming algorithm for the area where the first type of user is located, and form beam nulls for adjacent areas; at the same time, synthesize masking noise and play the masking noise directionally to the adjacent areas through loudspeakers located at the edge.

2. The semantically driven adaptive sound field control and privacy protection method according to claim 1, characterized in that, In step S1, the step of real-time RIR estimation specifically includes: Based on the signal transmission characteristics between the microphone array and the preset anchor point speaker, the room impulse response of the public space is estimated in real time using the currently playing audio signal or the generated test signal as the excitation source; wherein, the system broadcast status is determined, and if it is in a silent state for more than a preset duration threshold, the speaker is driven to emit a test signal that is inaudible to the human ear, and the echo of the test signal is used for RIR estimation. The equivalent sound velocity of the current sound propagation medium is inverted from the RIR and used to correct the calculation of the user's position coordinates.

3. The semantically driven adaptive sound field control and privacy protection method according to claim 1, characterized in that, In step S1, the methods for obtaining the user's identity information include: Identity and permissions can be obtained by reading the user's portable location tag; or by extracting the user's voice signal to establish a temporary acoustic feature mapping, and deregistering the feature after the user leaves the specific area.

4. The semantically driven adaptive sound field control and privacy protection method according to claim 1, characterized in that, In step S2, the MPC cost function J(k) is expressed as: J(k)=Σ M i=1 [ω1(S)·(L i (k)-L target ) 2 +ω2(S) P i (k)+ω2(S)·I total (k)]; Where M is the total number of partitions, L i (k) represents the actual volume of the i-th partition at time k, L target For the target volume, P i (k) represents the power consumption of the speaker group serving the i-th zone at time k. total (k) represents the total acoustic energy leaked from the interference source area (such as the waiting area) to the sensitive area (such as the VIP area) at time k, and ω1(S), ω2(S) and ω3(S) are the weight coefficients that change with the semantic state label S.

5. The semantically driven adaptive sound field control and privacy protection method according to claim 1, characterized in that, The S2 step also includes priority arbitration logic: The semantic state label also includes an emergency priority state; When an emergency priority state is determined, the optimization process of the MPC cost function is bypassed, the privacy protection strategy is lifted, and the gain of the speaker is directly adjusted to the maximum output value.

6. The semantically driven adaptive sound field control and privacy protection method according to claim 1, characterized in that, In step S3, the process of solving for each loudspeaker parameter also includes: Predict the user's movement trajectory sequence at several future moments; During MPC optimization, the cumulative cost within the predicted field of view corresponding to the motion trajectory sequence is calculated to achieve a smooth transition of the sound field focus.

7. The semantically driven adaptive sound field control and privacy protection method according to claim 1, characterized in that, In step S4, the synthesis and playback of the masking noise specifically includes: Analyze the spectral envelope of the residual speech signal reaching the adjacent region; Generate a masking noise signal whose power spectral density matches the spectral envelope; The Speech Transmission Index (STI) is calculated based on the signal-to-noise ratio (SNR) of the masking noise signal and the residual speech signal. The transmission intensity of the masking noise is then adjusted so that the STI value remains below a preset intelligibility threshold.

8. The semantically driven adaptive sound field control and privacy protection method according to claim 1, characterized in that, In step S4, the beamforming algorithm is a minimum variance distortionless response algorithm, and the depth of the beam null is set such that the signal energy attenuation in the adjacent region is reduced by at least 15 dB relative to the main lobe.

9. The semantically driven adaptive sound field control and privacy protection method according to claim 1, characterized in that, In step S4, the plurality of distributed broadcast loudspeakers are logically and dynamically divided into target sound field playback units and edge interference units. The same physical loudspeaker dynamically switches between the roles of the playback unit and the interference unit based on its relative position to the user.

10. A semantically driven adaptive sound field control and privacy protection system, applied to a public space equipped with several distributed loudspeakers and several distributed microphone arrays, characterized in that... include: The environment and user perception module is configured to acquire acoustic signals through the microphone array to estimate the room impulse response (RIR) and environmental parameters, and to obtain the user's identity information and location coordinates to generate a scene semantic feature vector; The semantic-driven strategy generation module is configured to map the scene semantic feature vector to a semantic state label, and dynamically adjust the weight coefficients and constraints in the model prediction control cost function based on the label. The adaptive calculation module is configured to solve for the speaker control parameters based on the adjusted cost function and to calculate the inverse filter for audio predistortion compensation. The sound field execution module is configured to drive the loudspeaker to perform beamforming on the first type of user area, and to create beam nulls and directional playback masking noise on adjacent areas.