A deep reinforcement learning-based sound masking strategy adaptive optimization method

CN122551816APending Publication Date: 2026-08-11BEIJING DEQI LIDA INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-18
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

目标空间内语音事件具有突发性和方位变化特征,传统声掩蔽方法多依据声压级或简单频谱特征触发掩蔽,难以及时提取语音共振峰、空间到达方向和瞬态语义区域,导致掩蔽响应滞后或掩蔽区域不准确;现有深度强化学习方法在声掩蔽控制中通常直接采用通用状态、动作和奖励设计,未充分结合语音可懂度、声压泄漏、用户舒适度、能耗和动作平滑等声学约束,容易出现策略震荡、过度掩蔽或听觉烦躁;传统滤波式掩蔽频谱存在频带边界突变,难以在语音共振峰附近形成自然渐变的保护带宽;传统波束赋形依赖人工标定声场传递函数,对房间混响、扬声器布设和目标区域变化的适应能力不足;不同空间之间声学差异较大,直接迁移声掩蔽模型容易产生冷启动时间长和策略负迁移问题

Benefits of technology

本发明通过音频Mel频谱与多通道时延特征的流式时频注意力融合,生成带方位标签的瞬态语义掩码,解决传统声掩蔽仅依据声压级或单一频谱触发而难以定位突发语音区域的问题,使声掩蔽策略能够针对语音共振峰和空间方位进行及时响应;通过心理声学对比编码和信息瓶颈压缩生成感知潜变量,解决高维声学特征冗余、边缘端推理负担大以及人工心理声学指标表达不足的问题,提高状态表征的紧凑性和感知相关性;通过改进型TD-MPC2声掩蔽策略网络引入KalMamba声场隐状态估计、潜空间轨迹预测和模型预测控制动作选择,并结合语音可懂度安全阈值、声压泄漏约束、用户舒适度约束、能耗约束和二阶差分平滑惩罚,解决通用强化学习策略震荡、过度掩蔽和听觉烦躁的问题,实现更稳定的目标声掩蔽动作参数输出;通过平流扩散频谱重塑形成渐变保护带宽,解决传统滤波频带边界突变的问题;通过隐式声场表示模型和反馈校正机制,减少人工声场标定依赖,提高目标掩蔽区投射精度;通过云侧策略潜空间类别原型聚合,为新目标空间生成初始声掩蔽策略,降低跨空间部署冷启动成本,对研发办公、会议保密和智能声环境管理具有工程应用意义。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551816A_ABST
    Figure CN122551816A_ABST
Patent Text Reader

Abstract

This invention discloses an adaptive optimization method for acoustic masking strategies based on deep reinforcement learning, comprising the following steps: acquiring ambient sound streams from a multi-channel microphone array within the target space and generating transient semantic masks with directional labels; extracting target audio segments and generating perceptual latent variables; constructing a deep reinforcement learning state space, action space, and acoustic masking strategy reward value; generating target acoustic masking action parameters through an improved TD-MPC2 acoustic masking strategy network; constructing a masking energy concentration field and generating an acoustic masking spectrum with a gradually changing protection bandwidth; generating phase-weighted coefficients and amplitude-weighted coefficients using an implicit sound field representation model, projecting the acoustic masking signal onto the target masking area; and aggregating strategy latent space category prototypes on the cloud side to generate an initial acoustic masking strategy. This invention improves the adaptability, privacy protection, and cross-space deployment efficiency of acoustic masking strategies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of acoustic privacy protection and intelligent sound field control technology, and in particular to an adaptive optimization method for sound masking strategies based on deep reinforcement learning. Background Technology

[0002] With the increasing prevalence of open-plan offices, collaborative R&D, meetings, and remote voice interaction scenarios, sound masking and spatial sound field control technologies to address the risk of voice content leakage have received widespread attention. Existing sound masking systems primarily rely on white noise, pink noise, preset natural sounds, or masking signals based on fixed filters, and play the masking sound in the target area using speaker arrays. However, these systems commonly suffer from the following problems in practical applications: Speech events in the target space are characterized by suddenness and directional changes. Traditional sound masking methods often rely on sound pressure level or simple spectral features to trigger masking, making it difficult to extract speech formants, spatial directions of arrival, and transient semantic regions in a timely manner, resulting in delayed masking responses or inaccurate masking areas. Existing deep reinforcement learning methods in sound masking control typically use general state, action, and reward designs directly, without fully considering acoustic constraints such as speech intelligibility, sound pressure leakage, user comfort, energy consumption, and action smoothness, which can easily lead to policy oscillations, over-masking, or auditory annoyance. Traditional filter-based masking spectra have abrupt changes in frequency band boundaries, making it difficult to form a naturally gradual protection bandwidth near speech formants. Traditional beamforming relies on manually calibrated sound field transfer functions, which are insufficiently adaptable to changes in room reverberation, speaker placement, and target area. There are significant acoustic differences between different spaces, and directly transferring the sound masking model can easily lead to long cold start times and negative policy transfer problems.

[0003] Therefore, how to provide an adaptive optimization method for acoustic masking strategies based on deep reinforcement learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0004] One objective of this invention is to propose an adaptive optimization method for sound masking strategies based on deep reinforcement learning. This invention utilizes multi-channel microphone array perception, psychoacoustic contrastive coding, an improved TD-MPC2 sound masking strategy network, advection diffusion spectrum reshaping, and an implicit sound field representation model. It details the adaptive optimization process from environmental sound stream recognition, perception latent variable generation, sound masking action decision-making to multi-channel loudspeaker projection and cloud-side strategy prototype transfer. It has the advantages of good privacy protection, high sound field control accuracy, high user comfort, and high cross-space deployment efficiency.

[0005] An adaptive optimization method for acoustic masking strategy based on deep reinforcement learning according to an embodiment of the present invention includes the following steps: Step 1: Acquire the ambient sound stream of a multi-channel microphone array in the target space, extract the audio Mel spectrum and multi-channel time delay features, and perform streaming time-frequency attention fusion processing to generate a transient semantic mask with orientation labels; Step 2: Extract the target audio segment based on the transient semantic mask, perform psychoacoustic contrastive coding and information bottleneck compression on the target audio segment, and generate perceptual latent variables; Step 3: Construct a deep reinforcement learning state space based on transient semantic mask and perceptual latent variables, construct the acoustic masking parameter set as the action space, and generate acoustic masking policy reward values ​​based on privacy and security constraints and acoustic operation constraints; Step 4: Input the deep reinforcement learning state space and action space into the improved TD-MPC2 acoustic masking policy network, and generate target acoustic masking action parameters by KalMamba acoustic field hidden state estimation, latent space trajectory prediction, candidate action sequence evaluation and model prediction to control action selection. Step 5: Construct a masking energy concentration field based on the target acoustic masking action parameters and the speech formant frequency band, and perform advection diffusion spectrum reshaping on the masking energy concentration field to generate an acoustic masking spectrum with a gradually changing protection bandwidth; Step 6: Input the sound masking spectrum and target spatial coordinate parameters into the implicit sound field representation model to generate the phase weighting coefficients and amplitude weighting coefficients of the multi-channel loudspeaker, and project the sound masking signal onto the target masking area; Step 7: Based on the sound field feedback data and the sound masking strategy reward value, form a state transition sample, write back to update the improved TD-MPC2 sound masking strategy network, aggregate the strategy latent space category prototype on the cloud side, generate an initial sound masking strategy based on the acoustic characteristics of the new target space, and complete the adaptive optimization of the sound masking strategy.

[0006] Optionally, the ambient sound stream is acquired by a multi-channel microphone array under the same sampling clock, and carries the channel number, acquisition timestamp, and sound pressure sampling value; The ambient sound stream is streamed and framed according to a preset frame length and preset frame shift to generate a multi-channel audio frame sequence. The center channel of the array is selected as the reference channel, and DC component removal, pre-emphasis, windowing, short-time spectrum transformation and Mel filter bank mapping are performed on the reference channel audio frames to generate the audio Mel spectrum. The reference channel audio frame is paired with the other channel audio frames according to the same frame number, and generalized cross-correlation-phase transformation processing is performed to obtain the sound arrival time difference. The phase difference at the same frequency band position is extracted, and the sound arrival time difference, phase difference, cross-correlation peak intensity and channel pair number are concatenated to generate multi-channel time delay features. Align the audio Mel spectrum and multi-channel delay features according to frame number and acquisition timestamp to generate a time-frequency spatial joint frame sequence; perform streaming time-frequency attention fusion processing on the time-frequency spatial joint frame sequence, encapsulate the audio Mel spectrum into a time-frequency token, encapsulate the multi-channel delay features into a spatial token, and establish a causal buffer window composed of the current frame and its preceding buffer frames; Within the causal cache window, time-frequency tokens are used as attention query terms, and spatial tokens are used as attention key and value terms. Time-frequency spatial attention weights are generated based on the consistency of frequency band energy changes, the stability of sound arrival time difference, and the stability of phase difference. The spatial tokens are then weighted and aggregated to generate a fused token sequence. The threshold for sudden speech change is determined based on the statistical results of frequency band energy change in the background audio frame. The fusion tokens that reach the threshold for sudden speech change and whose sound arrival time difference falls within the same preset directional interval are marked as sudden speech tokens. A transient semantic mask is generated based on the frame number, Mel frequency band number and preset directional interval corresponding to the sudden speech token, and a directional label is written into it to generate a transient semantic mask with a directional label.

[0007] Optionally, step two specifically includes: Based on the effective speech regions, frame numbers, Mel frequency band numbers, and location labels in the transient semantic mask with location labels, continuous effective speech regions are merged, and target audio segments are extracted from the ambient sound stream based on the merged start and end frames. Amplitude normalization and endpoint smoothing are performed on the target audio segment. Based on short-time energy, A-weighted sound pressure level, Mel band energy distribution, spectral centroid, low-frequency fluctuations of the audio envelope, and linear predictive coding, a psychoacoustic reference feature set is generated, including energy trajectory, loudness approximate trajectory, high-frequency energy proportion, spectral centroid trajectory, modulation spectrum features, and formant positions. Based on psychoacoustic reference feature groups and location labels, comparative sample pairs are constructed. Two target audio segments with psychoacoustic reference features under the same location label that are within the same preset classification range are marked as positive sample pairs, and two target audio segments with any psychoacoustic reference feature that are within different preset classification ranges are marked as negative sample pairs. The Mel frequency band energy distribution, Mel frequency cepstral coefficients, psychoacoustic reference feature set and orientation label of the target audio segment are concatenated into an encoded input sequence, which is then input into a one-dimensional convolutional coding layer and a gated recurrent coding layer to generate an initial perceptual embedding vector. Based on positive and negative sample pairs, a contrastive learning constraint is performed on the initial perceptual embedding vector to generate a perceptual embedding vector. The perceptual embedding vector is then input into a linear bottleneck layer and compressed into a low-dimensional vector of a preset dimension through a fully connected mapping to generate perceptual latent variables.

[0008] Optionally, step three specifically includes: Based on the effective speech region and perceptual latent variables in the transient semantic mask, the time frame position, Mel frequency band number, orientation label and perceptual latent variables corresponding to the effective speech region are concatenated into the current acoustic state vector and written into the deep reinforcement learning state space. The sound masking frequency band center, bandwidth, bandwidth energy, time envelope, spatial projection intensity, and loudspeaker channel weights are used as a set of sound masking parameters. The action value boundaries, action change step size, and channel weight normalization range are set respectively to generate the action space. Calculate the speech transmission index based on the target audio segment and compare it with the preset speech intelligibility security threshold to generate privacy security reward items or privacy security penalty items; The sound pressure leakage is calculated based on the A-weighted equivalent sound pressure level collected by microphones in the target concealment area and non-target area, and compared with the preset sound pressure leakage threshold to generate a sound pressure leakage constraint term. Based on the A-weighted equivalent sound pressure level, spectral centroid change, short-time energy fluctuation, and loudspeaker output power of the acoustic masking signal, user comfort constraints and energy consumption constraints are generated. The second-order difference smoothing penalty term is calculated based on the set of acoustic masking parameters at three consecutive time points; Using a preset baseline reward value as the initial value, the reward value for sound masking strategy is generated by adding or subtracting items based on privacy and security reward items, privacy and security penalty items, sound pressure leakage constraint items, user comfort constraint items, energy consumption constraint items, and second-order difference smoothing penalty items. The result after the addition or subtraction is limited to the preset reward value range.

[0009] Optionally, the improved TD-MPC2 acoustic masking strategy network includes a KalMamba acoustic field hidden state estimation layer, a latent space trajectory prediction layer based on a latent space dynamics model, a candidate action sequence evaluation layer based on a value prediction network, and a model prediction control action selection layer. The KalMamba acoustic field hidden state estimation layer performs field normalization, sliding time window splicing, and one-dimensional convolution projection on the current acoustic state vector based on the current acoustic state vector and historical acoustic field feedback data to generate an acoustic state input sequence. The acoustic state input sequence is written into a selective state-space scanning structure, and feedback residuals are generated based on historical sound field feedback data. Sound field confidence hidden states are generated through the prediction and correction steps of Kalman filtering. The latent space trajectory prediction layer is based on the sound field confidence latent state and candidate sound masking action parameters. The candidate sound masking action parameters are generated into a latent space action vector through a multilayer perceptron action embedding layer and input into the latent space dynamics model to predict the sound field confidence latent state, speech transmission index prediction value, sound pressure leakage prediction value and user comfort prediction value at the next moment. The candidate action sequence evaluation layer generates candidate sound masking action sequences based on the action space. The candidate sound masking action sequences are then progressively input into the latent space dynamics model for rolling prediction. The single-step prediction reward is calculated based on the predicted values ​​of the speech transmission index, sound pressure leakage, user comfort, and changes in sound masking action parameters at each prediction step. Finally, the value prediction network generates the candidate action sequence evaluation results. The model predicts the control action selection layer to remove candidate sound masking action sequences that exceed the preset speech intelligibility safety threshold, preset sound pressure leakage threshold, preset smoothing threshold, or preset energy consumption threshold, and selects the sound masking action parameter with the highest long-term masking benefit from the remaining candidate sound masking action sequences as the target sound masking action parameter. A temporal difference objective is constructed based on state transition samples and value prediction network output, and the parameters of the KalMamba acoustic field hidden state estimation layer, latent space trajectory prediction layer, and candidate action sequence evaluation layer are updated using mini-batch gradient descent.

[0010] Optionally, step five specifically includes: Based on the acoustic masking frequency band center, bandwidth, frequency band energy and time envelope in the target acoustic masking action parameters, as well as the speech formant frequency band in the transient semantic mask, the speech formant frequency band is converted into a set of frequency sampling points in the short-time Fourier spectrum through Mel frequency band inverse mapping to generate the formant target frequency band; Based on the center of the sound masking frequency band, bandwidth, band energy, and time envelope, a masking energy concentration field is established on the frequency axis of the short-time Fourier spectrum; When performing advection processing on the masking energy concentration field, the advection direction is generated based on the frequency offset direction between the initial energy peak position and the target frequency band of the resonance peak. The masking energy is moved along the frequency axis using first-order upwind finite difference, and the spectral energy is kept continuous by linear interpolation. When performing diffusion processing on the masking energy concentration field after advection processing, the target frequency band of the resonance peak is used as the diffusion center. The second-order central finite difference is used to extend the masking energy to the adjacent frequency sampling points on both sides, and the energy amplitude is reduced according to the distance to generate a gradually changing protection bandwidth that covers the target frequency band of the resonance peak and whose edge energy decreases step by step. The masking energy concentration field after diffusion processing is subjected to Hanning window spectrum smoothing, energy limiting and inter-frame crossfading processing. The processed masking energy concentration field is converted into a short-time Fourier amplitude spectrum and written into the background sound phase or random phase according to the availability of the background sound phase to generate a sound masking spectrum with a gradient protection bandwidth.

[0011] Optionally, the target spatial coordinate parameters include the three-dimensional coordinates of the multi-channel loudspeaker, the three-dimensional coordinates of the sampling points in the target masking area, and the three-dimensional coordinates of the monitoring points in the non-target area. The implicit sound field representation model includes a Fourier position coding layer, a multilayer perceptron sound field transmission prediction layer, a complex weight solving layer, and a feedback correction layer. Before training the implicit sound field representation model, logarithmic sweep signals are played sequentially by each channel loudspeaker to collect response signals from the target masking area and non-target areas. Amplitude response samples and phase response samples are generated through short-time Fourier transform and impulse response estimation. The Fourier position coding layer generates a position coding vector based on the three-dimensional coordinates of the loudspeaker, the three-dimensional coordinates of the sampling points, and the frequency sampling points in the sound masking spectrum; the multilayer perceptron sound field transmission prediction layer predicts the amplitude transmission value and the phase transmission value based on the position coding vector; When training the implicit sound field representation model, amplitude response samples and phase response samples are used as sound field transfer supervision data, and the training constraints are: the sound pressure level of the target masking area reaches the target masking sound pressure range, the sound pressure level of the non-target area does not exceed the preset sound pressure leakage threshold, the phase continuity of adjacent frequency sampling points, and the amplitude smoothness of adjacent speaker channels. The complex weighting solution layer uses constrained least squares to solve the complex driving weights of the multi-channel loudspeaker based on the sound masking spectrum, amplitude transfer value, and phase transfer value, and decomposes them into phase weighting coefficients and amplitude weighting coefficients. The acoustic masking signal for each speaker channel is generated based on the phase weighting coefficient and the amplitude weighting coefficient and projected onto the target masking area; The feedback correction layer calculates the residuals based on the measured A-weighted equivalent sound pressure level and speech transmission index after projection, and uses recursive least squares to correct the phase weighting coefficients and amplitude weighting coefficients.

[0012] Optionally, step seven specifically includes: Based on the sound field feedback data, the sound masking strategy reward value, the current acoustic state vector, the target sound masking action parameters, and the acoustic state vector at the next moment, the data is packaged into state transition samples according to the acquisition timestamp and written back to the experience cache of the improved TD-MPC2 sound masking strategy network. The sound field confidence hidden state, target sound masking action parameter distribution and long-term masking benefit are extracted from the updated and improved TD-MPC2 sound masking strategy network to generate a local sound masking strategy representation. Local acoustic features are generated based on reverberation time, background noise A-weighted equivalent sound pressure level, target cover area, distance from target cover area to non-target area, number of loudspeaker channels and loudspeaker spacing. The local acoustic features and local acoustic masking strategies are encoded into strategy latent space vectors, and local strategy latent space category prototypes are generated according to reverberation time. The cloud side performs a sample-weighted average of the prototypes of the latent space categories of the same local strategy to generate the prototypes of the latent space categories of the cloud side strategy and the corresponding acoustic feature centers. When a new target space is accessed, the interpolation weights are determined based on the normalized weighted Euclidean distance between the acoustic features of the new target space and the centers of each acoustic feature, and the cloud-side strategy latent space category prototypes are weighted and summed to generate the initialization vector of the new target space strategy latent space. Based on the new target space strategy, the initial sound field confidence hidden state, initial candidate action sampling center, initial action value range and initial parameters of the value prediction network are generated from the latent space initialization vector, thus forming the initial sound masking strategy.

[0013] The beneficial effects of this invention are: This invention generates a transient semantic mask with orientation labels by fusing audio Mel spectrum with streaming time-frequency attention based on multi-channel delay features. This solves the problem that traditional sound masking, which relies solely on sound pressure level or a single spectrum trigger, struggles to locate sudden speech regions, enabling sound masking strategies to respond promptly to speech formants and spatial orientation. Furthermore, it generates perceptual latent variables through psychoacoustic contrastive coding and information bottleneck compression, addressing issues of redundant high-dimensional acoustic features, heavy inference burden at the edge, and insufficient expression of artificial psychoacoustic indicators, thus improving the compactness of state representation and perceptual relevance. Finally, it introduces KalMamba latent state estimation, latent space trajectory prediction, and model predictive control through an improved TD-MPC2 sound masking strategy network. Action selection, combined with speech intelligibility safety thresholds, sound pressure leakage constraints, user comfort constraints, energy consumption constraints, and second-order difference smoothing penalties, addresses the issues of oscillation, over-masking, and auditory annoyance in general reinforcement learning strategies, achieving more stable output of target acoustic masking action parameters. It resolves the problem of abrupt changes in traditional filter band boundaries by reshaping the advection diffusion spectrum to form a gradually changing protection bandwidth. Through implicit sound field representation models and feedback correction mechanisms, it reduces dependence on manual sound field calibration and improves the projection accuracy of the target masking area. Furthermore, by aggregating cloud-side strategy latent space category prototypes, it generates initial acoustic masking strategies for new target spaces, reducing the cold-start cost of cross-space deployment. This has engineering application significance for R&D offices, meeting confidentiality, and intelligent sound environment management. Attached Figure Description

[0014] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of an adaptive optimization method for acoustic masking strategies based on deep reinforcement learning proposed in this invention; Figure 2 This is a schematic diagram of an adaptive optimization method for acoustic masking strategies based on deep reinforcement learning proposed in this invention. Figure 3 This is a framework diagram of the improved TD-MPC2 acoustic masking strategy network in the adaptive optimization method of acoustic masking strategy based on deep reinforcement learning proposed in this invention. Detailed Implementation

[0015] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0016] refer to Figures 1-3 An adaptive optimization method for acoustic masking strategies based on deep reinforcement learning includes the following steps: Step 1: Acquire the ambient sound stream of a multi-channel microphone array in the target space, extract the audio Mel spectrum and multi-channel time delay features, and perform streaming time-frequency attention fusion processing to generate a transient semantic mask with orientation labels; Step 2: Extract the target audio segment based on the transient semantic mask, perform psychoacoustic contrastive coding and information bottleneck compression on the target audio segment, and generate perceptual latent variables; Step 3: Construct a deep reinforcement learning state space based on transient semantic mask and perceptual latent variables, construct the acoustic masking parameter set as the action space, and generate acoustic masking policy reward values ​​based on privacy and security constraints and acoustic operation constraints; Step 4: Input the deep reinforcement learning state space and action space into the improved TD-MPC2 acoustic masking policy network, and generate target acoustic masking action parameters by KalMamba acoustic field hidden state estimation, latent space trajectory prediction, candidate action sequence evaluation and model prediction to control action selection. Step 5: Construct a masking energy concentration field based on the target acoustic masking action parameters and the speech formant frequency band, and perform advection diffusion spectrum reshaping on the masking energy concentration field to generate an acoustic masking spectrum with a gradually changing protection bandwidth; Step 6: Input the sound masking spectrum and target spatial coordinate parameters into the implicit sound field representation model to generate the phase weighting coefficients and amplitude weighting coefficients of the multi-channel loudspeaker, and project the sound masking signal onto the target masking area; Step 7: Based on the sound field feedback data and the sound masking strategy reward value, form a state transition sample, write back to update the improved TD-MPC2 sound masking strategy network, aggregate the strategy latent space category prototype on the cloud side, generate an initial sound masking strategy based on the acoustic characteristics of the new target space, and complete the adaptive optimization of the sound masking strategy.

[0017] In this embodiment, the target space is an indoor acoustic area equipped with a multi-channel microphone array, and the ambient sound stream is continuous audio data collected by each microphone channel under the same sampling clock and carrying the channel number, collection timestamp, and sound pressure sampling value. The ambient sound stream is streamed and framed according to the preset frame length and preset frame shift to generate a multi-channel audio frame sequence. The center channel of the array is selected as the reference channel. DC component removal, pre-emphasis, windowing, short-time spectrum transformation and Mel filter bank mapping are performed on the reference channel audio frames. The energy values ​​of each frame in the Mel auditory frequency band are arranged according to the frame number to generate the audio Mel spectrum. The reference channel audio frame is paired with the other channel audio frames according to the same frame number. Generalized cross-correlation-phase transformation processing is performed. The sampling point offset corresponding to the cross-correlation peak is read and converted into the sound arrival time difference according to the sampling rate. At the same time, the phase difference at the same frequency band position is extracted. The sound arrival time difference, phase difference, cross-correlation peak intensity and channel pair number are concatenated to generate multi-channel time delay features. Align the audio Mel spectrum and multi-channel delay features according to frame number and acquisition timestamp to generate a time-frequency spatial joint frame sequence; perform streaming time-frequency attention fusion processing on the time-frequency spatial joint frame sequence. The streaming time-frequency attention fusion processing includes: encapsulating the audio Mel spectrum into a time-frequency token that records the frame number, Mel frequency band number and frequency band energy value; encapsulating the multi-channel delay features into a spatial token that records the channel pair number, sound arrival time difference and phase difference; and establishing a causal buffer window composed of the current frame and its preceding buffer frames. Within the causal buffer window, time-frequency tokens are used as attention queries, and spatial tokens are used as attention keys and values. Time-frequency spatial attention weights are generated based on the consistency of frequency band energy changes, the stability of sound arrival time difference, and the stability of phase difference under the same frame number. Stability means that the inter-frame variation of the corresponding feature within the causal buffer window does not exceed a preset stability threshold. Spatial tokens are weighted and aggregated according to the time-frequency spatial attention weights, and the weighted aggregation result is written into the corresponding time-frequency token to generate a fused token sequence. Based on the statistical results of frequency band energy changes in the background audio frames, a sudden speech change threshold is determined. Fusion tokens whose frequency band energy changes are not less than the sudden speech change threshold and whose sound arrival time differences fall within the same preset directional interval are marked as sudden speech tokens, and the remaining fusion tokens are marked as background tokens. A transient semantic mask is generated based on the frame number, Mel frequency band number, and preset directional interval corresponding to the sudden speech token, and the directional label corresponding to the preset directional interval is written into the transient semantic mask to generate a transient semantic mask with directional labels. The directional label is the sound source direction number that represents the arrival direction of the target audio segment relative to the multi-channel microphone array.

[0018] In this embodiment, step two specifically includes: Read the effective speech region, frame number, Mel frequency band number and directional label from the transient semantic mask with directional label, merge the continuous effective speech regions, and extract the target audio segment from the ambient sound stream according to the start and end frames of the merged segment. The target audio segment is short audio data containing burst speech components and carrying directional labels. Amplitude normalization and endpoint smoothing are performed on the target audio segment. The energy trajectory is calculated based on short-time energy. The loudness approximate trajectory is calculated based on A-weighted sound pressure level. The high-frequency energy ratio is calculated based on Mel frequency band energy distribution. The spectral centroid trajectory is calculated based on spectral centroid. The modulation spectrum features are calculated based on low-frequency fluctuations of the audio envelope. The positions of the first formant, second formant, and third formant are extracted based on linear predictive coding to generate a psychoacoustic reference feature set. Comparative sample pairs are constructed based on psychoacoustic reference feature sets and azimuth labels. The loudness approximation trajectory, high-frequency energy ratio, spectral centroid trajectory, modulation spectrum features, and formant positions in the psychoacoustic reference feature sets are classified according to the feature distribution of historical target audio segments. Two target audio segments with all psychoacoustic reference features under the same azimuth label falling within the same preset classification range are marked as positive sample pairs, and two target audio segments with any psychoacoustic reference feature falling within different preset classification ranges are marked as negative sample pairs. The target audio segment is subjected to psychoacoustic contrast coding. The Mel frequency band energy distribution, Mel frequency cepstral coefficients, psychoacoustic reference feature set and orientation label of the target audio segment are concatenated into the coding input sequence. The Mel frequency cepstral coefficients are speech spectrum envelope features obtained by logarithmic compression and cepstral transformation of Mel frequency band energy. The encoded input sequence is fed into a one-dimensional convolutional coding layer and a gated recurrent coding layer. First, the one-dimensional convolutional coding layer extracts the local energy changes between adjacent frequency bands, and then the gated recurrent coding layer extracts the temporal changes between consecutive frames to generate an initial perceptual embedding vector. Based on positive and negative sample pairs, a contrastive learning constraint is performed on the initial perceptual embedding vector to reduce the distance between the initial perceptual embedding vectors corresponding to positive sample pairs in the latent space and increase the distance between the initial perceptual embedding vectors corresponding to negative sample pairs in the latent space, thus generating a perceptual embedding vector. The perceptual embedding vector is input into a linear bottleneck layer, which is a fully connected mapping layer with an output dimension smaller than the input dimension. After being compressed into a low-dimensional vector of a preset dimension by the fully connected mapping, the low-dimensional vector retains the difference information corresponding to the speech formant position, modulation spectrum features and orientation labels through contrastive learning constraints, thereby generating perceptual latent variables.

[0019] In this embodiment, step three specifically includes: Read the effective speech region and perceptual latent variables from the transient semantic mask, concatenate the time frame position, Mel frequency band number, orientation label and perceptual latent variables corresponding to the effective speech region into the current acoustic state vector, and write it into the deep reinforcement learning state space; The sound masking frequency band center, bandwidth, bandwidth energy, time envelope, spatial projection intensity, and loudspeaker channel weights are used as a set of sound masking parameters. The action value boundaries, action change step size, and channel weight normalization range are set respectively to generate the action space. Calculate the speech transmission index based on the target audio segment corresponding to the effective speech region, compare the speech transmission index with the preset speech intelligibility security threshold, and generate privacy security reward items or privacy security penalty items. The sound pressure leakage is calculated based on the A-weighted equivalent sound pressure level collected by microphones in the target concealment area and non-target area, and compared with the preset sound pressure leakage threshold to generate a sound pressure leakage constraint term. Based on the A-weighted equivalent sound pressure level, spectral centroid change, short-time energy fluctuation, and loudspeaker output power of the sound masking signal, a user comfort evaluation value and a sound pressure energy consumption constraint are generated. These values ​​are then compared with preset comfort thresholds and preset energy consumption thresholds to generate user comfort constraint items and energy consumption constraint items. Read the sound masking parameter set for three consecutive time moments, calculate the parameter change at the current time relative to the previous time moment, and then calculate the magnitude of the change in the parameter relative to the previous parameter change to generate a second-order difference smoothing penalty term. When generating the sound masking strategy reward value, a preset baseline reward value is used as the initial value. A privacy and security reward item is added when the voice transmission index is not greater than a preset voice intelligibility safety threshold. When the voice transmission index is greater than the preset voice intelligibility safety threshold, a privacy and security penalty item is deducted according to the extent of exceeding the threshold. When the sound pressure leakage is greater than a preset sound pressure leakage threshold, a sound pressure leakage constraint item is deducted according to the extent of exceeding the leakage limit. When the user comfort evaluation value is less than a preset comfort threshold, a user comfort constraint item is deducted according to the extent of insufficient comfort. When the speaker output power is greater than a preset energy consumption threshold, an energy consumption constraint item is deducted according to the extent of exceeding the energy consumption limit. When the second-order differential smoothing penalty item is greater than a preset smoothing threshold, a smoothing penalty is deducted according to the extent of exceeding the limit. The result after these additions and subtractions is limited to the preset reward value range to generate the sound masking strategy reward value.

[0020] In this embodiment, the improved TD-MPC2 acoustic masking strategy network includes a KalMamba acoustic field hidden state estimation layer, a latent space trajectory prediction layer based on a latent space dynamics model, a candidate action sequence evaluation layer based on a value prediction network, and a model prediction control action selection layer. The KalMamba acoustic field latent state estimation layer reads the current acoustic state vector and historical acoustic field feedback data from the deep reinforcement learning state space. The historical acoustic field feedback data consists of speech transmission index, sound pressure leakage, and user comfort evaluation values, which are collected and calculated by microphones in the target masking area and non-target areas after the acoustic masking signal is projected onto the target masking area. Field normalization, sliding time window splicing, and one-dimensional convolutional projection are performed on the time frame position, Mel frequency band number, azimuth label, and perceptual latent variables in the current acoustic state vector to generate an acoustic state input sequence. The acoustic state input sequence is written into a selective state space scanning structure to extract the temporal dependent states of speech formant changes, short-time energy changes, and azimuth changes. Feedback residuals are generated based on the historical acoustic field feedback data, and the temporal dependent states are corrected using the prediction and correction steps of Kalman filtering to generate the acoustic field confidence latent state. The latent space trajectory prediction layer based on the latent space dynamics model reads the sound field confidence latent state and candidate sound masking action parameters. It inputs the sound masking frequency band center, bandwidth, frequency band energy, time envelope, spatial projection intensity and speaker channel weight into the multilayer perceptron action embedding layer to generate a latent space action vector. The sound field confidence latent state and the latent space action vector are concatenated and then input into the latent space dynamics model to predict the sound field confidence latent state, speech transmission index prediction value, sound pressure leakage prediction value and user comfort prediction value at the next moment. The candidate action sequence evaluation layer based on the value prediction network generates multiple candidate sound masking action sequences by using cross-entropy sampling based on the action value boundaries, action change step size, and channel weight normalization range in the action space. The candidate sound masking action sequences are then progressively input into the latent space dynamics model for rolling prediction. Based on the predicted values ​​of speech transmission index, sound pressure leakage, user comfort, and sound masking action parameters obtained at each prediction step, the single-step prediction reward is calculated according to the sound masking strategy reward value generation method in claim 4. The value prediction network then estimates the single-step prediction reward for consecutive prediction steps to generate the candidate action sequence evaluation result. The model predictive control action selection layer eliminates candidate sound masking action sequences based on the evaluation results of candidate action sequences. These sequences include those with predicted speech transmission index values ​​exceeding a preset speech intelligibility safety threshold, predicted sound pressure leakage values ​​exceeding a preset sound pressure leakage threshold, second-order change amplitude exceeding a preset smoothing threshold, or loudspeaker output power exceeding a preset energy consumption threshold. The layer then selects the candidate sound masking action sequence with the highest long-term masking benefit from the remaining candidate sound masking action sequences and determines the sound masking action parameters corresponding to the current moment in this candidate sound masking action sequence as the target sound masking action parameters. When updating the policy based on the state transition samples, the current acoustic state vector, target acoustic masking action parameters, acoustic masking policy reward value, and next-time acoustic state vector are read from the experience cache. The next-time value prediction result output by the value prediction network is read. The cumulative discount of the acoustic masking policy reward value and the next-time value prediction result is used to construct a temporal difference objective. The next-time acoustic state vector is used as the state supervision data of the latent space dynamics model. The temporal difference objective is used as the reward supervision data of the value prediction network. The parameters of the KalMamba acoustic field hidden state estimation layer, latent space trajectory prediction layer, and candidate action sequence evaluation layer are updated using mini-batch gradient descent. The target acoustic masking action parameters that have been redefined after the policy update are output. The improved TD-MPC2 acoustic masking strategy network retains the original TD-MPC2 latent space dynamics model, candidate action sequence evaluation, model prediction control action selection, and temporal difference update structure based on state transition samples. It is used to predict sound field changes in the latent space and select the target acoustic masking action parameters at the current moment. The difference lies in the addition of a KalMamba acoustic field latent state estimation layer at the state input end, which generates acoustic field confidence latent states through one-dimensional convolutional projection, selective state space scanning structure, and Kalman filter correction; and the introduction of speech transmission index prediction, sound pressure leakage prediction, user comfort prediction, and the change amplitude of sound masking action parameters at the action evaluation end.

[0021] The above improvements are designed to address the issues of observable sound field in the target space, strong reverberation disturbances, and the tendency of ordinary reinforcement learning strategies to oscillate and over-mask. The improvement ensures that candidate sound masking action sequences are selected only after satisfying constraints on privacy, sound pressure leakage, comfort, energy consumption, and smoothness, thereby improving sound masking stability, reducing auditory annoyance, and reducing sound pressure leakage in non-target areas.

[0022] In this embodiment, step five specifically includes: Read the sound masking frequency band center, bandwidth, frequency band energy and time envelope from the target sound masking action parameters, and read the speech formant frequency band from the transient semantic mask. Convert the speech formant frequency band into a set of frequency sampling points in the short-time Fourier spectrum through Mel frequency band inverse mapping to generate the formant target frequency band. The initial energy peak position is determined based on the center of the sound masking frequency band, the initial energy coverage range is determined based on the bandwidth, the initial energy amplitude of each frequency sampling point is determined based on the frequency band energy, the energy change boundary between adjacent audio frames is determined based on the time envelope, and a masking energy concentration field is established on the frequency axis of the short-time Fourier spectrum. The masking energy concentration field is a frequency domain energy matrix that records the masking energy amplitude of different frequency sampling points in each audio frame. When performing advection processing on the masking energy concentration field, the advection direction is generated based on the frequency offset direction between the initial energy peak position and the resonant peak target frequency band. The masking energy is moved along the frequency axis using a first-order windward finite difference method, so that the masking energy is migrated from the initial energy coverage area to the resonant peak target frequency band frame by frame. During the migration process, linear interpolation is performed on adjacent frequency sampling points to maintain the continuity of spectral energy. When performing diffusion processing on the masking energy concentration field after advection processing, the target frequency band of the resonance peak is used as the diffusion center. The second-order central finite difference is used to extend the masking energy in the target frequency band of the resonance peak to the adjacent frequency sampling points on both sides. The energy amplitude is reduced in order from near to far from the diffusion center to generate a gradient protection bandwidth that covers the target frequency band of the resonance peak and whose edge energy decreases step by step. The masking energy concentration field after diffusion processing is subjected to spectral smoothing, energy limiting and inter-frame crossfade processing. Spectral smoothing uses the Hanning window to smooth the energy amplitude of adjacent frequency sampling points along the frequency axis. Energy limiting restricts the energy amplitude of each frequency sampling point to the range of frequency band energy values ​​corresponding to the target sound masking action parameters. Inter-frame crossfade connects the masking energy changes of adjacent audio frames according to the time envelope. The processed masking energy concentration field is converted into a short-time Fourier amplitude spectrum. When background sound phase information exists, the background sound phase is written in; when background sound phase information does not exist, a random phase is written in, generating an acoustic masking spectrum with a gradually changing protection bandwidth.

[0023] In this embodiment, step six specifically includes: The target spatial coordinate parameters include the three-dimensional coordinates of the multi-channel loudspeaker, the three-dimensional coordinates of the sampling points in the target masking area, and the three-dimensional coordinates of the monitoring points in the non-target area. The implicit sound field representation model includes a Fourier position coding layer, a multilayer perceptron sound field transmission prediction layer, a complex weight solving layer, and a feedback correction layer. Before training the implicit sound field representation model, the speakers of each channel sequentially play logarithmic sweep signals, and the response signals are collected by microphones in the target masking area and non-target area. Short-time Fourier transform and impulse response estimation are performed on the response signals to generate amplitude response samples and phase response samples from each speaker to each sampling point. The Fourier position coding layer reads the three-dimensional coordinates of the loudspeaker, the three-dimensional coordinates of the sampling point, and the frequency sampling points in the sound masking spectrum. It maps the coordinates and frequency sampling points into position coding vectors. The multilayer perceptron sound field transfer prediction layer predicts the amplitude transfer value and phase transfer value from the corresponding loudspeaker to the corresponding sampling point based on the position coding vector. When training the implicit sound field representation model, amplitude response samples and phase response samples are used as sound field transmission supervision data. The training constraints are that the sound pressure level in the target masking area reaches the target masking sound pressure range, the sound pressure level in the non-target area does not exceed the preset sound pressure leakage threshold, the phase change of adjacent frequency sampling points is continuous, and the amplitude change of adjacent speaker channels is smooth. Mini-batch gradient descent is used to update the parameters of the Fourier position coding layer and the multilayer perceptron sound field transmission prediction layer. The complex weighting solution layer reads the sound masking spectrum and the amplitude and phase transfer values ​​output by the multilayer perceptron sound field transfer prediction layer. It uses constrained least squares to solve the complex driving weights of the multi-channel loudspeaker and decomposes the complex driving weights into phase weighting coefficients and amplitude weighting coefficients. The acoustic masking spectrum is weighted in multiple channels based on phase weighting coefficients and amplitude weighting coefficients. The acoustic masking signal for each speaker channel is generated by short-time Fourier inverse transform and then projected onto the target masking area by the multi-channel loudspeaker. The feedback correction layer reads the measured A-weighted equivalent sound pressure level of the target masking area, the measured A-weighted equivalent sound pressure level of the non-target area, and the speech transmission index after the sound masking signal is projected. The measured results are compared with the prediction results of the implicit sound field representation model to calculate the residuals. The recursive least squares correction phase weighting coefficient and amplitude weighting coefficient are used to generate the feedback-corrected multi-channel loudspeaker driving parameters.

[0024] In this embodiment, step seven specifically includes: Read the sound field feedback data and the sound masking strategy reward value. The sound field feedback data includes the voice transmission index of the target masking area, the A-weighted equivalent sound pressure level of the target masking area, the A-weighted equivalent sound pressure level of the non-target area, and the user comfort evaluation value. Encapsulate the sound field feedback data, the sound masking strategy reward value, the current acoustic state vector, the target sound masking action parameters, and the acoustic state vector at the next moment into a state transition sample according to the acquisition timestamp, and write it back to the experience cache of the improved TD-MPC2 sound masking strategy network. The sound field confidence hidden state, target sound masking action parameter distribution and long-term masking benefit are read from the updated and improved TD-MPC2 sound masking strategy network and spliced ​​to generate a local sound masking strategy characterization; the reverberation time of the target space, the background noise A-weighted equivalent sound pressure level, the area of ​​the target masking area, the distance from the target masking area to the non-target area, the number of loudspeaker channels and the spacing of loudspeaker placement are collected to generate local acoustic features; The local acoustic features and local acoustic masking strategy representations are input into the fully connected encoder to generate a strategy latent space vector. The center vector of the strategy latent space vector is obtained according to the reverberation time, and the local strategy latent space category prototype is generated. The cloud side receives the local strategy latent space category prototype, the mean of local acoustic features and the number of samples, and performs a weighted average of the prototypes of the same type according to the number of samples to generate the cloud side strategy latent space category prototype and the corresponding acoustic feature center. When a new target space is accessed, the acoustic features of the new target space are collected. The acoustic features of the new target space and the acoustic feature centers of each acoustic feature are calculated using normalized weighted Euclidean distance. A preset number of cloud-side strategy latent space category prototypes are selected in ascending order of distance. The reciprocal of each normalized weighted Euclidean distance is taken and normalized to generate interpolation weights. The selected cloud-side strategy latent space category prototypes are weighted and summed according to the interpolation weights to generate the initialization vector of the new target space strategy latent space. The initial sound field confidence latent state, initial candidate action sampling center, initial action value range and initial parameters of the value prediction network are generated based on the initialization vector of the new target space strategy latent space. These are written into the improved TD-MPC2 sound masking strategy network of the new target space to generate the initial sound masking strategy, and are locally updated according to the state transition samples during the operation of the new target space.

[0025] Example 1: To verify the feasibility of this invention in practice, it was applied to a sound masking scenario in a research and development office area. This office area, approximately 180m², includes an open workstation area, two small meeting rooms, and a secure communication area. An 8-channel microphone array and a 12-channel speaker array are deployed. The sampling rate is set to 16kHz. The ambient sound stream is streamed in frames with a 32ms frame length and a 16ms frame shift. The number of Mel filter banks is set to 64, the causal buffer window is set to 8 frames, and the sudden speech change threshold is determined based on the average frequency band energy change of the previous 60s background audio frames plus 2.5 times the standard deviation. The preset stability threshold is set to ensure that the time difference between adjacent buffer frames does not exceed 0.18ms and the phase difference does not exceed 0.22rad. When the system is running, a multi-channel microphone array simultaneously collects conversations, keyboard typing, air conditioning noise, and conference voices from R&D personnel. First, it generates audio Mel spectrum and multi-channel time delay features. Then, it generates transient semantic masks with location labels through streaming time-frequency attention fusion processing. This marks voice segments such as "sudden voices near the conference room door" and "short sentences of communication on the right side of the open workstation" and hands them over to subsequent sound masking strategies for processing.

[0026] In this embodiment, the target audio segment is extracted based on the transient semantic mask, and short-time energy, A-weighted sound pressure level, spectral centroid, modulation spectrum features, and the positions of the first to third formants are extracted to construct a psychoacoustic reference feature set. A preset grading range is divided into five grading ranges according to the feature distribution of historical target audio segments. Segments with the same psychoacoustic reference features within the same grading range under the same directional label are marked as positive sample pairs, while segments with any feature in different grading ranges are marked as negative sample pairs. The encoded input sequence is input into a one-dimensional convolutional coding layer and a gated recurrent coding layer, and then compressed into 16-dimensional perceptual latent variables through a linear bottleneck layer. The deep reinforcement learning state space consists of the effective speech region and perceptual latent variables in the transient semantic mask, while the action space includes the sound masking frequency band center, bandwidth, band energy, temporal envelope, spatial projection intensity, and speaker channel weights. The preset speech intelligibility safety threshold is set to STI≤0.32, the preset sound pressure leakage threshold is set to no more than -8dB relative to the target masking area in the non-target area, the preset comfort threshold corresponds to an A-weighted equivalent sound pressure level of no more than 48dB(A) and a short-term energy fluctuation of no more than 3.5dB, the preset energy consumption threshold is set to an average output power of no more than 6W for a single-channel speaker, and the preset smoothing threshold is set to a second-order change amplitude of the sound masking parameter set at three consecutive moments of no more than 0.12.

[0027] The improved TD-MPC2 acoustic masking strategy network runs on a local edge server. The KalMamba acoustic field latent state estimation layer generates acoustic field confidence latent states based on the current acoustic state vector and historical acoustic field feedback data. The latent space trajectory prediction layer predicts the speech transmission index, sound pressure leakage, and user comfort prediction values ​​for the next moment. The candidate action sequence evaluation layer performs rolling predictions on 16 sets of candidate acoustic masking action sequences. The model prediction control action selection layer eliminates candidate sequences that exceed preset speech intelligibility safety thresholds, preset sound pressure leakage thresholds, preset smoothing thresholds, or preset energy consumption thresholds, and outputs the target acoustic masking action parameters. Subsequently, a masking energy concentration field is constructed based on the speech formant frequency band. First-order upwind finite difference is used for advection processing, and second-order center finite difference is used for diffusion processing. Finally, a masking spectrum with a gradually changing protection bandwidth is generated through Hanning window spectral smoothing and inter-frame crossfading. The implicit sound field representation model is initially trained using a logarithmic sweep frequency signal. Inputs include the three-dimensional coordinates of multi-channel loudspeakers, the three-dimensional coordinates of sampling points in the target masking area, and the three-dimensional coordinates of monitoring points in the non-target area. Outputs phase-weighted coefficients and amplitude-weighted coefficients, and performs recursive least-squares feedback correction based on measured A-weighted equivalent sound pressure level and speech transmission index. The cloud side only receives the local policy latent space category prototype and the local acoustic feature mean, without uploading ambient sound streams or target audio clips. When a new conference room is launched, the normalized weighted Euclidean distance is calculated based on reverberation time, background noise A-weighted equivalent sound pressure level, target masking area, distance from the target masking area to the non-target area, number of loudspeaker channels, and loudspeaker spacing, and an initial sound masking strategy is generated.

[0028] This embodiment compares four methods. Method A is a fixed pink noise masking, played at a constant sound pressure level; Method B is a traditional filtered sound masking, using a parametric equalizer to generate the masking spectrum based on the speech spectrum; Method C is a standard PPO sound masking control, constructing states and rewards solely based on sound pressure level and speech transmission index; Method D is an unmodified TD-MPC2 sound masking strategy network, without setting KalMamba latent state estimation, advection diffusion spectrum reshaping, and implicit sound field feedback correction; this invention adopts the complete process described in the claims. The test set includes 40 R&D discussion audio clips, 20 conference audio clips, and 20 call audio clips, each 30 seconds long, and is repeatedly tested under background noise conditions of 38 dB(A) to 46 dB(A) and reverberation times of 0.38 s to 0.72 s.

[0029] Table 1. Comparison of the effects of sound masking implementation in the R&D office area.

[0030] As shown in Table 1, although Method A is simple to implement, its average STI is 0.49, the speech content still has high intelligibility, and the sound pressure level in the target masking area reaches 51.8 dB(A), which can easily cause continuous auditory interference. Method B uses filtering to approximate the speech frequency, reducing the average STI to 0.42, but the sound pressure leakage in the non-target area is only -4.6 dB, indicating that the masking sound spreads outward significantly, and the abrupt change in the filtering boundary results in a user comfort score of only 3.1. After introducing a common PPO, Method C reduces the average STI to 0.36, but the second-order change amplitude of the policy action reaches 0.26, which is reflected in the significant changes in the intensity of the masking sound, indicating that there is policy oscillation in the general reinforcement learning control. Method D, using the unmodified TD-MPC2, achieved an average STI of 0.33, which is an improvement over Method C. However, due to the lack of KalMamba latent state estimation and implicit sound field feedback correction, the sound pressure leakage in non-target areas remained at -7.1 dB. The cold start convergence rounds for the new space were 22, and the deployment speed across conference rooms was still not ideal.

[0031] The average STI of this invention is 0.27, which is lower than the preset speech intelligibility safety threshold of 0.32, indicating that the intelligibility of the target speech content has been effectively reduced. The sound pressure level in the target masking area is 46.9 dB(A), which is lower than that of methods A and B, indicating that masking is not achieved by simply increasing the sound pressure. The sound pressure leakage in the non-target area reaches -9.4 dB, indicating that the implicit sound field representation model and feedback correction improve the projection capability of the target area. The user comfort score reaches 4.4, mainly due to the gradual protection bandwidth formed by the advection diffusion spectrum reshaping, which reduces the electronic noise caused by the abrupt change in the traditional filtering boundary. The second-order change amplitude of the strategy action is 0.08, indicating that the second-order difference smoothing penalty and model prediction control action selection effectively suppress the abrupt change of the sound masking parameters. The average response delay is 18 ms, indicating that the streaming time-frequency attention fusion processing can quickly generate transient semantic masks with directional labels after the sudden speech occurs. The cold start convergence rounds in the new space are 9, which is significantly lower than that of methods C and D, indicating that the cloud-side strategy latent space category prototype can provide an effective initial sound masking strategy for the new target space.

[0032] In conjunction with the above embodiments, this invention addresses the problem of sudden speech localization and response lag by using transient semantic masks with orientation labels, reduces the acoustic state dimension while preserving psychoacoustic differences by using perceptual latent variables, optimizes the constraints between privacy and security, sound pressure leakage, user comfort, energy consumption, and second-order differential smoothing through an improved TD-MPC2 acoustic masking strategy network, enhances the naturalness of the masking spectrum through advection-diffusion spectrum reshaping, reduces reliance on manual calibration through an implicit sound field representation model, and reduces the cold start cost of cross-spatial deployment through a cloud-side strategy latent space category prototype. Therefore, this invention has good engineering applicability and scalability in scenarios such as R&D offices, confidential meetings, remote collaboration, and open office privacy protection.

[0033] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for adaptive optimization of sound masking strategy based on deep reinforcement learning, characterized in that, Includes the following steps: Step 1: Acquire the ambient sound stream of a multi-channel microphone array in the target space, extract the audio Mel spectrum and multi-channel time delay features, and perform streaming time-frequency attention fusion processing to generate a transient semantic mask with orientation labels; Step 2: Extract the target audio segment based on the transient semantic mask, perform psychoacoustic contrastive coding and information bottleneck compression on the target audio segment, and generate perceptual latent variables; Step 3: Construct a deep reinforcement learning state space based on transient semantic mask and perceptual latent variables, construct the acoustic masking parameter set as the action space, and generate acoustic masking policy reward values ​​based on privacy and security constraints and acoustic operation constraints; Step 4: Input the deep reinforcement learning state space and action space into the improved TD-MPC2 acoustic masking policy network, and generate target acoustic masking action parameters by KalMamba acoustic field hidden state estimation, latent space trajectory prediction, candidate action sequence evaluation and model prediction to control action selection. Step 5: Construct a masking energy concentration field based on the target acoustic masking action parameters and the speech formant frequency band, and perform advection diffusion spectrum reshaping on the masking energy concentration field to generate an acoustic masking spectrum with a gradually changing protection bandwidth; Step 6: Input the sound masking spectrum and target spatial coordinate parameters into the implicit sound field representation model to generate the phase weighting coefficients and amplitude weighting coefficients of the multi-channel loudspeaker, and project the sound masking signal onto the target masking area; Step 7: Based on the sound field feedback data and the sound masking strategy reward value, form a state transition sample, write back to update the improved TD-MPC2 sound masking strategy network, aggregate the strategy latent space category prototype on the cloud side, generate an initial sound masking strategy based on the acoustic characteristics of the new target space, and complete the adaptive optimization of the sound masking strategy.

2. The method of claim 1, wherein, The ambient sound stream is acquired by a multi-channel microphone array at the same sampling clock and carries the channel number, acquisition timestamp, and sound pressure sampling value; The ambient sound stream is streamed and framed according to a preset frame length and preset frame shift to generate a multi-channel audio frame sequence. The center channel of the array is selected as the reference channel, and DC component removal, pre-emphasis, windowing, short-time spectrum transformation and Mel filter bank mapping are performed on the reference channel audio frames to generate the audio Mel spectrum. The reference channel audio frame is paired with the other channel audio frames according to the same frame number, and generalized cross-correlation-phase transformation processing is performed to obtain the sound arrival time difference. The phase difference at the same frequency band position is extracted, and the sound arrival time difference, phase difference, cross-correlation peak intensity and channel pair number are concatenated to generate multi-channel time delay features. Align the audio Mel spectrum and multi-channel delay features according to frame number and acquisition timestamp to generate a time-frequency spatial joint frame sequence; perform streaming time-frequency attention fusion processing on the time-frequency spatial joint frame sequence, encapsulate the audio Mel spectrum into a time-frequency token, encapsulate the multi-channel delay features into a spatial token, and establish a causal buffer window composed of the current frame and its preceding buffer frames; Within the causal cache window, time-frequency tokens are used as attention query terms, and spatial tokens are used as attention key and value terms. Time-frequency spatial attention weights are generated based on the consistency of frequency band energy changes, the stability of sound arrival time difference, and the stability of phase difference. The spatial tokens are then weighted and aggregated to generate a fused token sequence. Based on the statistical results of frequency band energy changes in the background audio frame, the threshold for sudden speech changes is determined. The fusion tokens that reach the threshold for sudden speech changes and whose sound arrival time difference falls within the same preset directional interval are marked as sudden speech tokens. A transient semantic mask is generated based on the frame number, Mel frequency band number and preset directional interval corresponding to the sudden speech token, and a directional label is written into it to generate a transient semantic mask with a directional label.

3. The method of claim 1, wherein, Step two specifically includes: Based on the effective speech regions, frame numbers, Mel frequency band numbers, and location labels in the transient semantic mask with location labels, continuous effective speech regions are merged, and target audio segments are extracted from the ambient sound stream based on the merged start and end frames. Amplitude normalization and endpoint smoothing are performed on the target audio segment. Based on short-time energy, A-weighted sound pressure level, Mel band energy distribution, spectral centroid, low-frequency fluctuations of the audio envelope, and linear predictive coding, a psychoacoustic reference feature set is generated, including energy trajectory, loudness approximate trajectory, high-frequency energy proportion, spectral centroid trajectory, modulation spectrum features, and formant positions. Based on psychoacoustic reference feature groups and location labels, comparative sample pairs are constructed. Two target audio segments with psychoacoustic reference features under the same location label that are within the same preset classification range are marked as positive sample pairs, and two target audio segments with any psychoacoustic reference feature that are within different preset classification ranges are marked as negative sample pairs. The Mel frequency band energy distribution, Mel frequency cepstral coefficients, psychoacoustic reference feature set and orientation label of the target audio segment are concatenated into an encoded input sequence, which is then input into a one-dimensional convolutional coding layer and a gated recurrent coding layer to generate an initial perceptual embedding vector. Based on positive and negative sample pairs, a contrastive learning constraint is performed on the initial perceptual embedding vector to generate a perceptual embedding vector. The perceptual embedding vector is then input into a linear bottleneck layer and compressed into a low-dimensional vector of a preset dimension through a fully connected mapping to generate perceptual latent variables.

4. The method of claim 1, wherein, Step three specifically includes: Based on the effective speech region and perceptual latent variables in the transient semantic mask, the time frame position, Mel frequency band number, orientation label and perceptual latent variables corresponding to the effective speech region are concatenated into the current acoustic state vector and written into the deep reinforcement learning state space. The sound masking frequency band center, bandwidth, bandwidth energy, time envelope, spatial projection intensity, and loudspeaker channel weights are used as a set of sound masking parameters. The action value boundaries, action change step size, and channel weight normalization range are set respectively to generate the action space. Calculate the speech transmission index based on the target audio segment and compare it with the preset speech intelligibility security threshold to generate privacy security reward items or privacy security penalty items; The sound pressure leakage is calculated based on the A-weighted equivalent sound pressure level collected by microphones in the target concealment area and non-target area, and compared with the preset sound pressure leakage threshold to generate a sound pressure leakage constraint term. Based on the A-weighted equivalent sound pressure level, spectral centroid change, short-time energy fluctuation, and loudspeaker output power of the sound masking signal, user comfort constraints and energy consumption constraints are generated. The second-order difference smoothing penalty term is calculated based on the set of acoustic masking parameters at three consecutive time points; Using a preset baseline reward value as the initial value, the reward value for sound masking strategy is generated by adding or subtracting items based on privacy and security reward items, privacy and security penalty items, sound pressure leakage constraint items, user comfort constraint items, energy consumption constraint items, and second-order difference smoothing penalty items. The result after the addition or subtraction is limited to the preset reward value range.

5. The method of claim 1, wherein, The improved TD-MPC2 acoustic masking strategy network includes a KalMamba acoustic field hidden state estimation layer, a latent space trajectory prediction layer based on a latent space dynamics model, a candidate action sequence evaluation layer based on a value prediction network, and a model prediction control action selection layer. The KalMamba acoustic field hidden state estimation layer performs field normalization, sliding time window splicing, and one-dimensional convolution projection on the current acoustic state vector based on the current acoustic state vector and historical acoustic field feedback data to generate an acoustic state input sequence. The acoustic state input sequence is written into a selective state-space scanning structure, and feedback residuals are generated based on historical sound field feedback data. Sound field confidence hidden states are generated through the prediction and correction steps of Kalman filtering. The latent space trajectory prediction layer is based on the sound field confidence latent state and candidate sound masking action parameters. The candidate sound masking action parameters are generated into a latent space action vector through a multilayer perceptron action embedding layer and input into the latent space dynamics model to predict the sound field confidence latent state, speech transmission index prediction value, sound pressure leakage prediction value and user comfort prediction value at the next moment. The candidate action sequence evaluation layer generates candidate sound masking action sequences based on the action space. The candidate sound masking action sequences are then progressively input into the latent space dynamics model for rolling prediction. The single-step prediction reward is calculated based on the predicted values ​​of the speech transmission index, sound pressure leakage, user comfort, and changes in sound masking action parameters at each prediction step. Finally, the value prediction network generates the candidate action sequence evaluation results. The model predicts the control action selection layer to remove candidate sound masking action sequences that exceed the preset speech intelligibility safety threshold, preset sound pressure leakage threshold, preset smoothing threshold, or preset energy consumption threshold, and selects the sound masking action parameter with the highest long-term masking benefit from the remaining candidate sound masking action sequences as the target sound masking action parameter. A temporal difference objective is constructed based on state transition samples and value prediction network output, and the parameters of the KalMamba acoustic field hidden state estimation layer, latent space trajectory prediction layer, and candidate action sequence evaluation layer are updated using mini-batch gradient descent.

6. The method of claim 1, wherein, Step five specifically includes: Based on the acoustic masking frequency band center, bandwidth, frequency band energy and time envelope in the target acoustic masking action parameters, as well as the speech formant frequency band in the transient semantic mask, the speech formant frequency band is converted into a set of frequency sampling points in the short-time Fourier spectrum through Mel frequency band inverse mapping to generate the formant target frequency band; Based on the center of the sound masking frequency band, bandwidth, band energy, and time envelope, a masking energy concentration field is established on the frequency axis of the short-time Fourier spectrum; When performing advection processing on the masking energy concentration field, the advection direction is generated based on the frequency offset direction between the initial energy peak position and the target frequency band of the resonance peak. The masking energy is moved along the frequency axis using first-order upwind finite difference, and the spectral energy is kept continuous by linear interpolation. When performing diffusion processing on the masking energy concentration field after advection processing, the target frequency band of the resonance peak is used as the diffusion center. The second-order central finite difference is used to extend the masking energy to the adjacent frequency sampling points on both sides, and the energy amplitude is reduced according to the distance to generate a gradually changing protection bandwidth that covers the target frequency band of the resonance peak and whose edge energy decreases step by step. The masking energy concentration field after diffusion processing is subjected to Hanning window spectrum smoothing, energy limiting and inter-frame crossfading processing. The processed masking energy concentration field is converted into a short-time Fourier amplitude spectrum and written into the background sound phase or random phase according to the availability of the background sound phase to generate a sound masking spectrum with a gradient protection bandwidth.

7. The method of claim 1, wherein, The target spatial coordinate parameters include the three-dimensional coordinates of the multi-channel loudspeaker, the three-dimensional coordinates of the sampling points in the target masking area, and the three-dimensional coordinates of the monitoring points in the non-target area. The implicit sound field representation model includes a Fourier position coding layer, a multilayer perceptron sound field transmission prediction layer, a complex weight solving layer, and a feedback correction layer. Before training the implicit sound field representation model, logarithmic sweep signals are played sequentially by each channel loudspeaker to collect response signals from the target masking area and non-target areas. Amplitude response samples and phase response samples are generated through short-time Fourier transform and impulse response estimation. The Fourier position coding layer generates a position coding vector based on the three-dimensional coordinates of the loudspeaker, the three-dimensional coordinates of the sampling points, and the frequency sampling points in the sound masking spectrum; the multilayer perceptron sound field transmission prediction layer predicts the amplitude transmission value and the phase transmission value based on the position coding vector; When training the implicit sound field representation model, amplitude response samples and phase response samples are used as sound field transfer supervision data, and the training constraints are: the sound pressure level of the target masking area reaches the target masking sound pressure range, the sound pressure level of the non-target area does not exceed the preset sound pressure leakage threshold, the phase continuity of adjacent frequency sampling points, and the amplitude smoothness of adjacent speaker channels. The complex weighting solution layer uses constrained least squares to solve the complex driving weights of the multi-channel loudspeaker based on the sound masking spectrum, amplitude transfer value, and phase transfer value, and decomposes them into phase weighting coefficients and amplitude weighting coefficients. The acoustic masking signal for each speaker channel is generated based on the phase weighting coefficient and the amplitude weighting coefficient and projected onto the target masking area; The feedback correction layer calculates the residuals based on the measured A-weighted equivalent sound pressure level and speech transmission index after projection, and uses recursive least squares to correct the phase weighting coefficients and amplitude weighting coefficients.

8. The method of claim 1, wherein, Step seven specifically includes: Based on the sound field feedback data, the sound masking strategy reward value, the current acoustic state vector, the target sound masking action parameters, and the acoustic state vector at the next moment, the data is packaged into state transition samples according to the acquisition timestamp and written back to the experience cache of the improved TD-MPC2 sound masking strategy network. The sound field confidence hidden state, target sound masking action parameter distribution and long-term masking benefit are extracted from the updated and improved TD-MPC2 sound masking strategy network to generate a local sound masking strategy representation. Local acoustic features are generated based on reverberation time, background noise A-weighted equivalent sound pressure level, target cover area, distance from target cover area to non-target area, number of loudspeaker channels and loudspeaker spacing. The local acoustic features and local acoustic masking strategies are encoded into strategy latent space vectors, and local strategy latent space category prototypes are generated according to reverberation time. The cloud side performs a sample-weighted average of the prototypes of the latent space categories of the same local strategy to generate the prototypes of the latent space categories of the cloud side strategy and the corresponding acoustic feature centers. When a new target space is accessed, the interpolation weights are determined based on the normalized weighted Euclidean distance between the acoustic features of the new target space and the centers of each acoustic feature, and the cloud-side strategy latent space category prototypes are weighted and summed to generate the initialization vector of the new target space strategy latent space. Based on the new target space strategy, the initial sound field confidence hidden state, initial candidate action sampling center, initial action value range and initial parameters of the value prediction network are generated from the latent space initialization vector, thus forming the initial sound masking strategy.