Selective sound source positioning method and apparatus based on cueing guidance

CN122330815BActive Publication Date: 2026-08-28TRUE SPACE (ZHUHAI) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610787805.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-08-28
Estimated Expiration
2046-06-03

AI Technical Summary

Technical Problem

[0005]本发明的第一目的是提供一种基于提示引导的选择性声源定位方法,解决如何在多声源环境中进行即时引导选择性声源定位的问题

Benefits of technology

[0008]由上述方案可见,本发明通过端对端训练的提示引导定位模型,基于与目标声源关联的提示输入,以及混合音频信号联合估计DoA和声源基数,实现在多声源场景中仅对与提示关联的目标声源进行定位,且可用于活动声源数量未知且随时间变化的场景中。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122330815B_ABST
    Figure CN122330815B_ABST
Patent Text Reader

Abstract

The application provides a selective sound source positioning method and device based on prompt guidance, wherein the method comprises: acquiring a mixed audio signal, an audio prompt and a text prompt; wherein the source of the mixed audio signal comprises a target sound source and an interference sound source; the audio prompt and the text prompt are both associated with the target sound source; processing the mixed audio signal, the audio prompt and the text prompt through a pre-trained prompt guidance positioning model to output a DoA posterior graph and a sound source base number, and determining a DoA track corresponding to the target sound source according to the DoA posterior graph and the sound source base number. The prompt guidance positioning model is obtained through end-to-end training, the DoA and the sound source base number are jointly estimated based on the audio prompt and the text prompt associated with the target sound source and the mixed audio signal, the positioning of the target sound source associated with the prompt is realized in a multi-sound source scene, and the problem of how to perform instant guided selective sound source positioning in a multi-sound source, unknown number of active sound sources and time-varying environment is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sound source localization technology, specifically to a selective sound source localization method and apparatus based on prompting guidance. Background Technology

[0002] Sound source localization (SSL) technology supports various microphone array-based audio and speech applications by estimating the spatial location of acoustic events. For example, in smart speakers, SSL enables direction-aware speech enhancement and far-field automatic speech recognition by guiding microphone array beamformers; in hearing aids, SSL achieves directional noise reduction through spatial filtering. While significant progress has been made in SSL technologies based on signal processing (such as GCCPHAT, MUSIC) or deep neural networks (such as CRNNs), existing methods are inherently semantically blind, treating all sound sources as generic signals and lacking identification capabilities: all active sound sources (including background noise) are non-selectively localized. This is fundamentally different from human auditory perception—when competing speakers and background noise are present, listeners can selectively focus on target sounds, i.e., the "cocktail party problem." See also Figure 1 Traditional SSL systems indiscriminately locate target sound sources and irrelevant sound sources (such as dog barking and background noise), making it impossible for users to focus the system on a specific target (such as speech).

[0003] To associate semantics with location, Sound Event Localization and Detection (SELD) methods can detect and locate events of all known categories. However, these methods perform passive scene analysis and lack interactive sound source filtering mechanisms based on user intent. In contrast, in the field of Target Speaker Extraction (TSE), TSE techniques can extract specific sound source guidance through audio or text cues. However, TSE focuses on waveform reconstruction rather than spatial awareness, and its extraction process often distorts the phase differences between channels, thus discarding crucial spatial cues needed to achieve accurate SSL. This leads to a fundamental contradiction: TSE has semantic awareness but does not consider location information, while traditional SSL can locate but ignores semantic information. Therefore, constructing a unified framework that combines semantic awareness and location awareness remains a pressing issue.

[0004] To address this, cue-guided localization methods conditionally adjust the angle-of-arrival (DoA) estimator based on user-specified cues, enabling it to selectively locate the queried target rather than indiscriminately locating all active sound sources. However, existing cue-guided localization research is primarily unimodal, utilizing only one type of cue—text or audio—and failing to employ both simultaneously in a unified model. Furthermore, its evaluation is often limited to simplified scenarios, with limited coverage of moving targets, noisy mixed audio, or cues corresponding to multiple targets. Summary of the Invention

[0005] The first objective of this invention is to provide a prompt-guided selective sound source localization method to solve the problem of how to perform real-time guided selective sound source localization in a multi-sound-source environment.

[0006] A second objective of this invention is to provide a computer device for a selective sound source localization method based on prompting guidance.

[0007] To achieve the aforementioned first objective, the present invention provides a cue-guided selective sound source localization method, comprising the following steps: acquiring a mixed audio signal; wherein the source of the mixed audio signal includes a target sound source and interfering sound sources; acquiring cue input, the cue input including audio cue and / or text cue; wherein the audio cue is associated with the target sound source; the text cue is associated with the target sound source; processing the mixed audio signal and cue input through a pre-trained cue-guided localization model, outputting a DoA posterior map and a sound source cardinality, and determining the DoA trajectory corresponding to the target sound source based on the DoA posterior map and the sound source cardinality.

[0008] As can be seen from the above scheme, the present invention uses an end-to-end trained prompt-guided localization model, based on prompt input associated with the target sound source, and jointly estimates DoA and sound source cardinality using mixed audio signals, to locate only the target sound source associated with the prompt in a multi-sound-source scenario, and can be used in scenarios where the number of active sound sources is unknown and changes over time.

[0009] A further approach is that the cue-guided selective attention module includes a cue-guided selective attention module and a DoA estimator; the cue-guided selective attention module is used to process the mixed audio signal and cue input to obtain the target-specific representation corresponding to the input data; the DoA estimator is used to output the DoA posterior map and the source cardinality based on the target-specific representation.

[0010] Thus, the present invention processes mixed inputs through a prompt-guided selective attention module to obtain a target-specific representation related to the target sound source, and then uses the target-specific representation for DoA estimation of the target sound source.

[0011] A further approach involves a cue-guided selective attention module comprising an audio encoder, a cue encoder, a fusion layer, and an extraction network. The audio encoder outputs amplitude cues and spatial cues based on the mixed audio signals. The spatial cues include inter-channel phase difference (IPD) and inter-channel intensity difference (ILD). The cue encoder extracts audio cues and text cues to obtain guidance vectors. The fusion layer fuses the amplitude cues and guidance vectors to obtain fused cues. The extraction network extracts target amplitude information and feature stacks based on the guidance vectors and fused cues. The DoA estimator comprises an IPD enhancer, a progressive optimization and temporal modeling module, and a prediction head. The IPD enhancer obtains enhanced IPD based on spatial cues, target amplitude information, and feature stacks. The progressive optimization and temporal modeling module processes enhanced IPD, inter-channel intensity difference (ILD), and target amplitude information to obtain alignment features. The prediction head comprises two parallel heads that process the alignment features to obtain the DoA posterior map and source cardinality.

[0012] As can be seen, the prompt-guided selective attention module generates a stack of extracted features based on text prompts and audio prompts. The extracted feature stack further guides the IPD enhancer to strengthen the spatial cues. The enhanced spatial cues are fused with the target amplitude information, and the DoA posterior map and source cardinality are estimated by combining two parallel heads. This coupled design can effectively focus the spatial cues on the user-specified target sound source for selective localization and achieve stable localization and tracking under conditions of moving speakers and intermittent activity.

[0013] A further approach is to include a fusion layer comprising two FiLM cascaded structures driven by a guide vector, which fuse the amplitude cues and the guide vector to obtain fused cues.

[0014] This demonstrates that a good feature fusion effect can be achieved.

[0015] A further approach is to extract the network by comprising multiple stacked DPRNN blocks; in the multiple stacked DPRNN blocks, cross attention is periodically injected between the current audio feature and the guiding vector to obtain the extraction information embedding corresponding to each DPRNN block, and the stacking of extraction features is determined based on the extraction information embedding.

[0016] Therefore, it can be seen that the effect of extracting feature stacks can be improved, and the extracted feature stacks can better reflect the characteristics of the target sound source.

[0017] A further approach is to use 6 DPRNN blocks.

[0018] Therefore, this quantity can achieve good performance.

[0019] A further approach is that the IPD enhancer includes parallel acoustic and semantic branches; the acoustic branch obtains acoustic features based on target amplitude information and spatial cues; the semantic branch obtains semantic features based on the stacking of extracted features; and the IPD enhancer obtains enhanced IPD based on the acoustic and semantic features.

[0020] This demonstrates that it can achieve better feature extraction and fusion effects related to prompts.

[0021] A further approach is that the progressive optimization and timing modeling module includes multiple FR blocks, multiple TCN blocks, and a bidirectional cyclic BiGRU. The outputs of the multiple FR blocks are connected to the multiple TCN blocks, and the outputs of the multiple TCN blocks are connected to the bidirectional cyclic BiGRU. The bidirectional cyclic BiGRU outputs the alignment features.

[0022] This demonstrates that spatial and spectral features can be tightly coupled, enabling modeling over longer time dimensions.

[0023] A further approach is to suggest that the total loss function used in training the localization model is a weighted sum of selection loss, DoA prediction loss, and sound source cardinality prediction loss, with the selection loss being the average negative SI-SNR across the two microphone channels.

[0024] Therefore, it can be seen that the weights of different losses can be adjusted according to the actual situation.

[0025] To achieve the second objective described above, the present invention provides a computer device comprising a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, it implements the above-described selective sound source localization method based on prompting guidance. Attached Figure Description

[0026] Figure 1 This is a schematic diagram of sound source localization using existing technology.

[0027] Figure 2 This is a schematic diagram of selective sound source localization according to the present invention.

[0028] Figure 3 This is a flowchart of an embodiment of the selective sound source localization method based on prompting guidance of the present invention.

[0029] Figure 4 This is a framework diagram of the prompt-guided localization model in an embodiment of the selective sound source localization method based on prompt guidance of the present invention.

[0030] Figure 5These are heatmaps of frame-by-frame DoA posterior probabilities and comparison diagrams of the estimated DoA trajectories and the actual DoA trajectories overlaid on four prompt-guided localization models trained under different prompt configurations in an embodiment of the selective sound source localization method based on prompt guidance of the present invention.

[0031] Figure 6 This is a polar coordinate diagram of the DoA trajectory in an embodiment of the selective sound source localization method based on prompting guidance of the present invention.

[0032] The present invention will be further described below with reference to the accompanying drawings and embodiments. Detailed Implementation

[0033] The present invention provides a prompt-guided selective sound source localization method that, when the target sound source is mobile and the number of active sound sources is unknown and changes over time, locates only the target sound source corresponding to the prompt based on a prompt given by the user, including audio and / or text, while suppressing interference.

[0034] See Figure 2 In a space containing the target sound source, interfering sound sources, and environmental noise, the localization model can focus solely on the target sound source and provide only the DoA trajectory of the target sound source based on the user's input prompts, including the text prompt "Locate the speech" and the audio prompt "Another audio clip".

[0035] It should be noted that the user-provided prompts can be audio or text. These prompts serve as input to the prompt-guided localization model, which outputs the DoA (DoA) of the target sound source matching the prompt description in each frame. For example, when the prompt input is only the text "piano," the prompt-guided localization model locates the piano sound source in the scene and outputs its DoA trajectory.

[0036] Example of a prompt-guided selective sound source localization method: This embodiment is implemented based on computer program execution; specifically, see [link to relevant documentation]. Figure 3 This includes the following steps: S11: Acquire the mixed audio signal.

[0037] S12: Get audio and text cues.

[0038] S13: By using a prompt-guided localization model to process mixed audio signals, audio prompts, and text prompts, the DoA trajectory corresponding to the target sound source is determined.

[0039] In step S11, the mixed audio signal refers to the dual-channel mixed audio signal collected by the microphone array (including two microphones). The dual-channel mixed audio signal comes from the sound signals of the target sound source, the interference sound source and the ambient noise. Each channel of the mixed audio signal corresponds to one microphone.

[0040] In step S12 above, the audio prompt is associated with the target sound source, and the text prompt is associated with the target sound source.

[0041] An audio cue is an audio clip containing key acoustic features of a target sound source. It provides audio clues to the model, enabling the model to pinpoint the target sound source within a mixed audio signal. Key acoustic features can be timbre or spectral characteristics, etc.

[0042] Text prompts are natural language instructions used by users to describe and specify target sound sources. They provide semantic cues to the model so that the prompts guide the localization model to identify the target sound source from the mixed audio signal.

[0043] In step S13 above, the DoA trajectory of the target sound source refers to the DoA change of the target sound source over time.

[0044] The prompt-guided localization model is a pre-trained neural network model, and the model and its training process will be further introduced below.

[0045] This embodiment is illustrated using a dual-channel recording system. For a dual-channel recording system, the signal received by the m-th microphone (m∈1,2) This can be modeled as the sum of the convolutions of J sound source signals with their respective room impulse responses (RIRs), and then adding ambient noise to this, expressed as: ,in, This represents the signal from the j-th sound source; This represents the number of active sound sources at time t; This represents the additive noise from the m-th microphone; Let RIR represent the distance between the j-th sound source and the m-th microphone at time t, where RIR depends on the time-varying angle of arrival of the sound source. ; This represents convolution.

[0046] This embodiment obtains a parameterized prompting and guidance localization model through training. The cue-guided localization model aims to estimate the DoA trajectory of a target sound source based on multimodal cues. The cue-guided localization model mixes two channels of signals. and audio prompts and text prompts This is mapped to two corresponding frame-level predictions: the DoA posterior map and the source cardinality.

[0047] The DoA posterior map is used to determine the most probable DOA of a target sound source in the current frame. In this embodiment, the DoA posterior map is represented as follows: This represents the frame-level DoA confidence map. Among them, Indicates the number of output frames after time alignment; , representing the number of discretized azimuth intervals from 0° to 179° with a resolution of 1°, thus limiting the prediction to the range of 0° to 180°. The multi-label output format of the DoA posterior map allows multiple azimuth intervals to be activated in each frame, providing continuous confidence values. The azimuth interval with the highest confidence is determined as the most probable DoA for the target sound source in the current frame.

[0048] The source cardinality is used to determine the number of target sound sources in the current frame. In this embodiment, the source cardinality is expressed as follows: ,in, This indicates the number of output frames after time alignment, and "3" indicates the probability distribution for the existence of 0, 1, or 2 active target sound sources.

[0049] The parameters of the localization model are guided by the optimization of the total loss function. The DoA posterior map and source cardinality corresponding to the target sound source are jointly predicted, and its form is defined as follows: .

[0050] The following will combine Figure 4 The prompting and guidance positioning model of this embodiment will be further introduced.

[0051] The cue-guided localization model comprises a cue-guided selective attention module and a DoA estimator. The cue-guided selective attention module generates a target-specific representation of the input data. The DoA estimator jointly predicts the DoA posterior map and the source cardinality based on the target-specific representation.

[0052] The cue-guided selective attention module acts as a selective filter for cue-guided input, combining text and audio cues to extract target-specific representations from the input data. Specifically, the input data for the cue-guided selective attention module includes a mixed audio signal, audio cues, and text cues, and the output is a target-specific representation.

[0053] The cue-guided selective attention module includes an audio encoder, a cue encoder, a fusion layer, and an extraction network. The cue encoder includes a parameter-frozen audio cue encoder and a parameter-frozen text encoder.

[0054] An audio encoder is used to output amplitude and spatial cues based on a mixed audio signal. In this embodiment, the input to the audio encoder is a dual-channel mixed audio signal acquired by two microphones. ,in, Complex spectrum calculated using STFT (Hanning window) ,in, Represents a time frame. The frequency range is represented, and then the complex spectrum is extracted to obtain amplitude clues and spatial clues.

[0055] Amplitude Clues The amplitude components of the complex spectrum It is obtained by sequentially encoding through the Conv1D layer and the ReLU layer, and is represented as: .in, Represents the set of real numbers. Indicates the embedding dimension.

[0056] Spatial cues are obtained from the phase components of the complex spectrum, specifically including the inter-channel phase difference (IPD) and the inter-channel intensity difference (ILD), expressed as: , ,in, It is a positive number used to ensure numerical stability (such as 10). 10 ).

[0057] The cue encoder uses a parameter-frozen Contrastive Language and Audio Pre-trained (CLAP) model to extract audio cues. and text prompts In this model, the audio cue encoder extracts audio cues to obtain audio feature vectors. The CLAP model's text encoder extracts text prompts to obtain text feature vectors. , is represented as: , Then, the audio feature vector and the text feature vector are projected and concatenated to form a guiding vector. , Since the purpose of projection is to ensure that the text feature vector and the audio feature vector have the same dimension, if the text feature vector and the audio feature vector have the same dimension, projection can be skipped and the guiding vector can be obtained directly by concatenation.

[0058] Optionally, the dual-channel mixed audio signal can specifically be a 6-second audio clip captured by a microphone array, or audio cues. It could be a 1-second audio clip associated with the target sound source, for example... Figure 3In this scenario, the target sound source is a cat, and the cat's location needs to be determined. Audio cues can be a 1-second audio clip of a cat meowing.

[0059] The fusion layer is used to integrate the bootstrap vector. With amplitude clues Fusion, obtaining fusion clues In the fusion layer, the following is used: The two FiLM cascaded structures driving the amplitude cues of the encoding Conditionalization is performed. A FiLM cascade structure includes a Norm layer (normalization layer), a FiLM layer (characteristic linear modulation layer), and a Projection layer. The Norm layer connects to the FiLM layer, the FiLM layer connects to the Projection layer, and the Projection layer connects to the Norm layer of the next FiLM cascade structure. Specifically, it can be represented as: ,in, and It is a learnable function; This indicates the input that needs to be modulated; This represents a condition, corresponding to a guiding vector, which is used to... The mapping is to scaling and translation parameters. The two FiLM cascade structures are conditionalized. Specifically, firstly, in the first FiLM cascade structure, the guiding vector... Modulation of amplitude cues across all dimensions Then a projection is performed; secondly, in the second FiLM cascade structure, the guiding vector is used. The projected features are modulated again to obtain the final fusion cues. .

[0060] Extraction network is used to extract based on the guiding vector With fusion clues Extract target amplitude information and feature stacking .

[0061] In extracting network clues, Perform chunking, dividing the length into segments. The data blocks are fragmented with 50% overlap to obtain the fragmented data blocks, and then... The processing is performed using stacked DPRNN blocks (including intra-block and inter-block RNNs). The preferred number of DPRNN blocks is six. Each DPRNN block includes an Intra-block (intra-block RNN) and an Inter-block (inter-block RNN). The Intra-block performs local modeling within the segmented data blocks, while the Inter-block performs global modeling between different data blocks. Within the stacked structure of the DPRNN blocks, the current audio feature is periodically compared with the guiding vector. Inject cross-attention between them. Specifically, let... This represents the frame-level audio features in a DPRNN block. This represents the guiding vector. The guiding vector... Broadcasting over the time dimension yields... And calculate cross attention: , , Given ,as well as , , , These are the learnable query weight matrix, key weight matrix, and value weight matrix, respectively. Indicates the output dimension. Attention weights are represented as... The residual update is represented as: Therefore, for each DPRNN block, a corresponding extracted information embedding can be obtained, and each extracted information embedding is linearly projected onto... The time-frequency feature map in the image is obtained by averaging the embedded information from the two microphone channels. , The extracted features are stacked in a stacked manner. , represented as These mappings highlight the time-frequency regions relevant to the extraction of the target sound source and provide guidance for the IPD enhancer.

[0062] exist In a stack of DPRNN blocks, the output of the last DPRNN block is passed through a projection layer to generate a target mask. Thus, through the target mask Amplitude Clues Filtering is performed to obtain the filtered characteristics corresponding to the target sound source. , is represented as: Finally, it passes through a multilayer perceptron decoder. For filtering features The target amplitude is obtained through reconstruction. Since it has been segmented previously, the target amplitude is concatenated through an Over-and-Add layer to obtain the target amplitude information. , is represented as: , 'm' represents the microphone channel. Since the amplitude information collected by different microphone channels at the same time is at the same height, the target amplitude information corresponding to the first microphone is used. Perform subsequent DoA estimation.

[0063] The output of the aforementioned cue-guided selective attention module is used as the target-specific representation input into the DoA estimator. The DoA estimator includes an IPD enhancer and a progressive optimization and temporal modeling module, which are used to output the DoA posterior map and the source cardinality. The inputs of the DoA estimator include: (1) the target amplitude information corresponding to the first microphone obtained from the cue-guided attention module. (2) Inter-channel intensity difference (ILD) and inter-channel phase difference (IPD) after processing with cosine and sinine functions; (3) Stacking of extracted features from the extraction network This embodiment does not stack the extracted features. It is directly concatenated to the temporal backbone network, and used only internally in the IPD enhancer used to optimize spatial cues. DoA estimator output Frame-level posterior maps within the azimuth angle interval, and classification distribution on the sound source count set. (This embodiment) ),in, This represents the maximum number of target sound sources that can be activated simultaneously.

[0064] The IPD enhancer includes acoustic and semantic branches. The enhanced IPD is obtained by optimizing the inter-channel phase difference (IPD) in the cue-guided selective attention module through the IPD enhancer, denoted as... The IPD enhancer relies on both the spatial cues of the cue-guided selective attention module and the target amplitude information corresponding to the first microphone. .

[0065] The input to the acoustic branch includes: , , The inputs to the acoustic branch are concatenated to form a 4-channel spectrogram representation, which is then processed by the acoustic feature encoder (Ac head) and then efficiently captured by a lightweight depthwise separable convolution (DWS Conv2d) to obtain the acoustic features.

[0066] The semantic branch is specifically designed to handle the stacking of extracted features from the network output. It operates in parallel with the acoustic branch. The semantic head (Sem head) encodes the extracted features into a stacked conditional representation, which then drives gated FiLM modulation to obtain semantic features. Specifically, global pooling is used to derive channel-level scaling and translation parameters, while spatial gating modulates the acoustic features to achieve spatially selective IPD enhancement.

[0067] The fusion and residual prediction (Fuse head & IPD Delta) method fuses acoustic and semantic features and calculates the residual based on the predicted value. .set up This represents the acoustic features generated by the acoustic branch. This represents the semantic features generated by the semantic header. These features are fused by applying FiLM conditionalization and combining it with gated modulation, and are represented as follows: ,in, This represents a spatially gating head that achieves spatially selective IPD enhancement by adjusting FiLM conditional features. Subsequently, a fusion head predicts the cosine-sine residuals (IPD deltas), and the enhanced IPD is achieved through the four-quadrant arctangent function. The calculation yields the following result, which is expressed as: , , , .

[0068] The inputs to the progressive optimization and time series modeling module include ILD and Specifically, for each frame The input is assembled by splicing, and is represented as: Stacking them along the time dimension yields a stacked sequence, represented as... , The feature dimension is represented. The stacked sequence is processed through a progressive optimization and temporal modeling module. First, the stacked sequence is projected (Pre-proj) and processed through a set of Feature Refinement Blocks (FR blocks), the composite operation of which is denoted as... Three independent threads are gradually merged using FR blocks: , The input and ILD are combined to obtain a fused representation. Each FR block is implemented using a depthwise separable one-dimensional convolutional layer (DWConv1d), a normalization layer (Norm), an activation function (ReLU), a squeeze-and-excitation (SE) module, and residual connections. This progressive fusion maps the input to a unified feature space where spatial and spectral information are tightly coupled. Based on this fused representation, multiple dilated TCN (Temporal Convolutional Network) blocks and a bidirectional gated recurrent unit (BiGRU) focus on modeling longer-range temporal patterns. The BiGRU consists of forward GRU layers and backward GRU layers, represented as follows: The temporal alignment layer (Aligner) achieves alignment by adjusting the frame rate, thus obtaining alignment features. , is represented as: .

[0069] For the prediction head, alignment features The decoding output is performed by two parallelizable heads: the Card Head and the DoA Head. Both the Card Head and the DoA Head are fully connected heads that represent the learned frame-by-frame affine projection. Let... and These represent the DoA header and radix header, respectively, used to output logical values ​​and applied line-by-line to alignment features. , is represented as: , ,in, , .

[0070] During inference, the two predictions are strictly coupled frame-by-frame. First, the number of active sound sources is estimated as follows: Subsequently, the orientation angle was estimated from... The former The peak values ​​are derived from the peak values. Specifically, the peak value selection includes: (1) the peak values ​​are derived from the peak values. (1) Perform Gaussian smoothing to handle the discontinuity of the surrounding area from 0° to 180°, while making the angle change smoother; (2) Identify the area exceeding the adaptive threshold. The strict local maxima, where, , The frame mean and standard deviation are... (3) Selection The highest peak value is used as the final direction angle prediction.

[0071] It's important to note that the number of sound sources predicted by the model is not the total number of all active sound sources in the scene, but rather the number of active target sound sources in the current scene. For example, suppose there are two sound sources in the scene (an irrelevant sound source and the target sound source to be located). At time t, both the irrelevant sound source and the target sound source are active. In the model's output of the DoA posterior map and sound source cardinality for that frame, the DoA posterior map outputs a significantly high probability at one azimuth angle (corresponding to the target sound source), and the sound source cardinality (corresponding to the target sound source) is 1. At time t+1, the target sound source is silent, and the irrelevant sound source is active. In the model's output of the DoA posterior map and sound source cardinality for that frame, the DoA posterior map does not output a significantly high probability at any azimuth angle, and the sound source cardinality is 0.

[0072] Since the model outputs frame-level DoA posterior maps and source cardinality for the target sound sources associated with the prompt input, without explicitly distinguishing instance-level identities, in special cases, there may be more than one target sound source in the scene. In this case, the model outputs the DoA trajectories of all target sound sources. During the trajectory drawing stage after inference, if it is determined that there are two target sound sources, a simple association strategy based on temporal continuity is used: peaks with small azimuth changes in adjacent frames are regarded as continuous trajectories of the same target sound source. In this embodiment, the number of target sound sources may be 0, 1, or 2. For example, if the sound source for the prompt input is a piano, and there are two piano sound sources in the scene (referred to as "first piano" and "second piano"), the model does not explicitly distinguish between these two piano sound sources. Even if the two piano sound sources have different timbres, as long as they are both associated with the prompt input, the model will treat both piano sound sources as target sound sources for localization. Assuming that the two piano sound sources are active and moving in the scene, the peak values ​​corresponding to the current frame are 45° and 120°, and the number of sound sources is 2. The corresponding peak values ​​in the next frame adjacent to the current frame are 47° and 122°, respectively. Based on the principle of distinguishing the same target sound source by the small change in azimuth angle, in the two adjacent frames, the DoA trajectory of the first piano moves from 45° to 47°, and the DoA trajectory of the second piano moves from 120° to 122°.

[0073] For the training objective, the model is jointly trained using three loss terms. The loss is chosen as the average negative SI-SNR across the two microphone channels, expressed as: ,in, This represents the target sound source waveform on the first microphone channel reconstructed by the model. This represents the target sound source reference signal on the first microphone channel. This represents the target sound source waveform on the second microphone channel reconstructed by the model. This represents the target sound source reference signal on the second microphone channel.

[0074] During the localization process, binary cross-entropy (BCE) is used for frame-level direction of arrival (DoA) estimation, and cross-entropy (CE) is used for source cardinality prediction, expressed as: , ,in, It is a soft DoA tag built based on the actual azimuth angle. It is a frame-by-frame sound source counting tag.

[0075] The total loss function is a weighted sum of the selection loss, the DoA prediction loss term, and the source basis prediction loss: Among them, fixed And adjust other weights to balance the impact of gradients for each task.

[0076] Choose loss (Negative SI-SNR) typically converges to Scale. In comparison, Based on The average BCE across several frequency bands. Due to the sparsity of the DoA targets (only 2 active out of a maximum of 180 bands), the convergent model will predict a loss value close to zero for approximately 99% of the outputs. Therefore, the average loss... Become very small, for example or To prevent the vanishing of the DoA gradient and ensure that its effect remains comparable to the separation term (i.e. This scale mismatch needs to be compensated for. Therefore, Set as Magnitude.

[0077] The following describes the datasets used for training and evaluation in this embodiment. Specifically, the datasets include synthetic datasets and real-world datasets.

[0078] For the synthetic dataset, training and evaluation are performed on a dataset covering a variety of acoustic conditions. The synthesis process combines clean sound source signals with simulated dynamic reverberation function (RIR) and ambient noise.

[0079] For the sound sources, to ensure the diversity of the target sound events, sound source signals were collected from multiple public datasets. Specifically, sound source signals were collected from the LibriSpeech corpus, and instrument sounds were collected from the CC-Music piano dataset and the GuitarSet dataset. To cover a wider range of general acoustic events, samples from AudioSet and WavCaps were also used.

[0080] For noise sources, realistic and challenging mixed noise samples were created by integrating multiple sets of noise recording data. Specifically, these data came from several datasets, including MSSNSD, WHAM, ESC-50, UrbanSound, QUT-noise, and Musan.

[0081] For the data simulation, all training, validation, and test sets were generated through synthesis based on the aforementioned sound and noise sources. This process relied on a dynamic reverberation chamber generated using the GPURIR library. For each simulation scenario, a room with dimensions of 4×4×2 meters was defined, with a reverberation time (…). The duration is 0.2 seconds. A dual-channel linear microphone array is placed at coordinates [1.9, 2.0, 1.0] and [2.1, 2.0, 1.0], such that the microphone spacing is 20 cm.

[0082] To simulate a moving sound source, a 5-second trajectory was generated within a 180° azimuth range, and uniform resampling was performed at a frequency of 15Hz. Each segment was generated. Discrete target directions are used. Furthermore, a coarser tag frame rate than the acoustic feature frame rate is employed to shorten the sequence while preserving sufficient temporal detail. At a 15Hz sampling rate, for a sound source moving 180° within 5 seconds, the maximum movement angle between adjacent tags is 2.4°.

[0083] Each audio mix sample is generated from one or two dynamic sound sources. Since these sources may be intermittently silent, the number of active sources in any given time frame may be zero, one, or two. These dynamic sound sources are generated by convolving a clean sound source signal with its corresponding dynamic reverberation response (RIR). Finally, the samples are uniformly sampled from […]. Within a 5,500 dB range, a signal-to-noise ratio (SNR) is selected, and randomly selected noise signals are added to the audio mixture samples. Training text prompts corresponding to the target sound source in the audio mixture samples can be manually labeled. Training audio prompts corresponding to the target sound source in the audio mixture samples can be another audio segment corresponding to the target sound source. For example, if the audio mixture samples include a 5-second audio segment of the target sound source, the corresponding training audio prompt is a 1-second audio segment of the target sound source that does not overlap with the aforementioned 5-second audio segment. Finally, a training set consisting of 288.9 hours of audio data, a validation set consisting of 31.1 hours of audio data, and a test set consisting of 18.1 hours of audio data are synthesized.

[0084] For real-world datasets, seven room subsets (air raid shelters, gymnasiums, PB132, PC226, SC203, SE203, and TB103) from the TAU-SRIR dataset were used to verify the effectiveness of the method in real acoustic environments. These subsets cover major typical indoor scenarios (open industrial spaces, sports stadiums, conference rooms, classrooms of various sizes, and lecture halls), including different surface materials and spatial volumes (rock and plastic-coated surfaces, carpet and hard floors, glass partitions), and include two types of sound source trajectories provided by the dataset (circular trajectories for air raid shelters, gymnasiums, PB132, and PC226, and linear trajectories for SC203, SE203, and TB103). The TAU-SRIR dataset provides discrete sound source locations along circular or linear trajectories. This embodiment uses a fixed sound source-microphone geometry (single SRIR) to render each 5-second evaluation segment, acquiring dual-channel signals by selecting symmetrical microphone pairs using a tetrahedral array. This design aims to ensure diversity in reverberation and reflection modes while maintaining a compact and non-redundant evaluation process. Five seconds of clean speech extracted from the LibriSpeech dataset were convolved with selected dual-channel RIRs (symmetric microphone pairs 0 and 2) and then compared with data from […]. The signal-to-noise ratio (SNR) of uniformly sampled background noise is 5,5]dB; the angle is represented in the dual-microphone frame by subtracting the array baseline and folding to [0°, 180°).

[0085] The specific training implementation details are as follows. At 16kHz, Hann STFT ( , ) below, get =513 and = 251; In the audio encoder, = 256; In the prompting and guiding selective attention module, , , 128, , In the DoA estimator, , For training, 200 epochs were used with early stopping (patience value = 10), and the Adam optimizer was used. Gradient clipping (clip norm = 1) was employed. In all experiments, the loss weights in the training objective were used... .

[0086] The following describes the corresponding experimental data for this embodiment, including comparison with the baseline model and ablation experiments.

[0087] For baseline models, the baselines are divided into four categories: (1) Trajectory-by-trajectory models, including: IPDNet, EINV2, embedded-ACCDoA, SALSA-Lite, DCASE2025 Task 3 baseline. These models assign a fixed number of output trajectories to the azimuth / activity trajectory. (2) SELD+DoA models, including: SELDnet, SELDT. These models output prediction results by category without explicit trajectory control. (3) Pure DoA models, including SRPDNN, FNSSL. These models directly estimate the angle of arrival from spatial features and do not include an activity detection branch. (4) Prompt-based methods. In this type of model, the text query SELD model is compared with the prompt-guided localization model proposed in this embodiment.

[0088] All baseline models were evaluated using a binaural non-FOA experimental setup (i.e., using standard 2-channel audio instead of the 4-channel first-order Ambisonics format). The specific adjustments were as follows: (1) Multi-microphone or first-order Ambisonics inputs were simplified to left-right channel pairs, consistent with the stereo / variable array setup used in existing studies. (2) Standard FOA features were replaced with dual-microphone spatial cues (cos / sin(IPD), ILD, GCCPHAT), following the feature replacement protocol used in the DCASE baseline and SALSA-lite. (3) The geometry was constrained to a horizontal plane (0° to 180° azimuth) to eliminate ambiguity. (4) The original prediction head and loss function were preserved with minimal modifications, in line with academic conventions and the design of variable array sound source localization. (5) The inference output was mapped to a uniform frame-by-frame DoA heatmap format to ensure fair comparison.

[0089] For evaluation metrics, frame-level static DoA metrics and trajectory-level dynamic performance metrics are used. Trajectory-level static DoA metrics include MAE (Mean Angular Error), Accuracy, F1 score, and Recall. Trajectory-level dynamic performance metrics include MOTA. (Multiple Object Tracking Accuracy; ID-agnostic, identity-agnostic version), DetA (Detection Accuracy), OSPA-T (Optimal Sub-Pattern Assignment for Tracking). For each frame ,make This represents the set of directional assumptions for prediction. This indicates the reference point for the annotation. The Hungarian assignment method is used to... and Matching, the solution minimizes The circumferential distance on the circle, i.e.: If the error A match is considered a correct match (TP), a non-match is a false positive (FP), and a non-match is a false negative (FN). , , Report accuracy , , .

[0090] MAE is used to measure the average angle deviation between the estimated DoA and the true DoA, expressed as: ,in, express and Hungarian assignment between them.

[0091] MOTA In this embodiment, the ID switching penalty is omitted; since all DoA hypotheses extracted under the same prompt belong to a single category, this embodiment provides an ID-independent variant, represented as follows: .

[0092] DetA measures the proportion of sound sources that are correctly detected within a given angular tolerance threshold, and is expressed as: .

[0093] OSPA-T is used to simultaneously evaluate angular error, source number estimation error, and trajectory continuity. Let as well as The frame count is: Among them, the truncation value Set to 15°, and , The sequence score is averaged over all frames: .

[0094] Specific experimental data are shown in Table 1, which illustrates the overall performance of all baseline models in frame-level static DoA and trajectory-level dynamic performance metrics. Static metrics are frame-level metrics (based on peak matching), and dynamic metrics are trajectory-level metrics. OSPA-T's reporting unit is degrees, with parameters set to p=1 and cutoff distance c=15. Acc, F1, Recall, and MOTA are also included. Both DetA and DetA are reported as percentages. † indicates non-trajectory-by-trajectory processing: the number of trajectories cannot be pre-specified when performing DoA estimation. Global best results are highlighted in bold, and best results for each category are underlined.

[0095] For ease of understanding, the results analysis is performed according to the four baseline models defined above.

[0096] To differentiate between the effects of architecture and input conditions, results under three input conditions are reported: Mixed input, Sel-Joint input, and Clean input. Specifically, Mixed input refers to the original stereo mixed signal with noise superimposed on an unknown number (0 / 1 / 2) of target sound sources, directly input into the baseline model without any extraction front-end processing; Sel-Joint input refers to end-to-end training, where the baseline model uses the same cue-guided selective attention module as in this embodiment at the input end, and its input is the concatenation of the enhancement amplitude of the cue-guided selective attention module and the original mixed phase to form a dual-channel representation; Clean input refers to a clean input condition containing only target sound sources consistent with the cue (without interfering sound sources and noise).

[0097] Among the baseline models for each trajectory, only the IPDNet and EINV2 models remain effective in mixed input scenarios. The IPDNet model has a MAE of 4.6° and an F1 score of 0.24, while the EINV2 model, although having a higher F1 score (0.40), suffers from a significantly increased MAE of 11.37°. This combination of high F1 and large MAE error reflects the design principle of its DP-IPD and ACCDoA models: as long as the target sound source is active, the tracking trajectory tends to be activated / triggered, but its estimated DoA is "pulled" towards a compromise position between the target and the interference source. This results in a higher average angle error, although the estimated value remains within the evaluation tolerance. The remaining trajectory baselines exhibit almost zero F1 and negative or near-zero values ​​in the mixed scenario. It is only competitive when Sel-Joint or Clean input simplifies the scenario.

[0098] Table 1. Model performance on different performance metrics

[0099] Under clean input conditions, the MAE of the IPDNet and EINV2 models is approximately 1.4°. Approximately 0.66–0.69, trajectory-level modeling works well once the text-specified target is effectively separated. However, the optimization goal of these systems is to interpret all active sound sources, rather than tracking a single target: IPDNet, based on DP-IPD, handles any sound source in a given direction in a similar manner, while style-dependent models like ACCDOA, and style systems (EINV2, SALSA-Lite, embedded-ACCDoA, DCASE2025_baseline), only weakly encode text queries. In complex mixed sound sources, trajectories often follow geometrically similar non-target sound sources or segment the target trajectory into short fragments.

[0100] The model in this embodiment (i.e., "Ours" in the table) combines a cue-guided selective attention module front-end with an azimuth DoA heatmap, and its heatmap training uses the same peak matching criterion as the evaluation. This design significantly improves performance under mixed conditions while maintaining the low MAE of the optimal trajectory-level baseline model under clean conditions. This effectively narrows the gap left by the design of a general trajectory-by-trajectory baseline model.

[0101] SELD+DoA † In the baseline models, the SELD+DoA baseline methods (SELDnet, SELDT) can predict frame-level activity and DoA for each class without explicit trajectories. Under mixed conditions, they are highly sensitive to interference, with SELDnet having a MAE of ≈15.8° and SELDT having a MAE of ≈22.1°, and both exhibiting MOTA. ≤ 0 indicates extremely poor detection performance. When the input is Sel-Joint and Clean, its static metrics are significantly improved (MAE ≈ 6.6°), which is consistent with the previous results of SELD in low clutter scenes.

[0102] The improvement in trajectory-level dynamic performance metrics is significantly insufficient. Even with clean input, SELDnet still performs worse than MOTA. =0.12, SELDT only reaches MOTA The value is approximately 0.19, indicating that inter-frame fluctuations in the SED / DoA outputs of each category still lead to trajectory fragmentation during evaluation. In contrast, SELDT performs better on trajectory-level metrics, confirming that its tracking-oriented objective helps stabilize trajectories. However, it is worth noting that both models are trained to account for all event categories rather than a single text-specific target; when multiple source data overlap, they distribute energy across multiple categories and directions rather than locking a single stable trajectory for the target. Therefore, even with Sel-Joint or Clean data as input, their MOTA (Motion Override) is not optimal. It is still far inferior to the best trajectory-level baseline model and the query conditionalization system of this embodiment, the latter of which can explicitly enforce the selectivity and temporal consistency of text queries. For pure DoA † (Non-track-wise) models, pure DoA baseline models (SRPDNN, FNSSL), require no SED head or explicit trajectory, directly estimating the angle of arrival (DoA) through binaural phase cues. The SRPDNN model follows the SRP-DNN paradigm: a causal CRNN predicts the spatial spectrum on a fixed azimuth grid for each frame, ultimately obtaining the DoA through peak extraction. In our binaural horizontal setup, this produced very similar localization results under different input conditions (MAE = 6.24 ° / Mix, 5.56 ° / SelJoint, 5.97 ° / Clean), and MOTA... Only modest changes were observed (≈0.07 / 0.07 / 0.11). This limited variation reflects the fixed spatial grid and lack of explicit temporal modeling: sharper amplitudes would sharpen peaks but only slightly affect their centers, while inter-frame fluctuations would still cause trajectory fragmentation.

[0103] The FNSSL model also estimates DP-IPD, but employs a full-band / narrow-band fusion technique to take advantage of cross-band and temporal structures. From Mix to Sel-Joint and Clean, its mean absolute error (MAE) improves by 7.01° to 5.8°, and MOTA... The performance improved from 0.18 to 0.23 / 0.21, achieving a better static-dynamic tradeoff compared to SRPDNN. However, the FNSSL model remains a purely geometric, single-stage DoA estimator that treats all spatial peaks in the mixed signal equally, lacking both a mechanism to focus on the text query target and independent trajectory tracking capabilities. Therefore, even in Clean, its performance is significantly inferior to our query conditionalization system, which integrates a cue-guided selective attention module front-end with an azimuth heatmap. This performance gap in both static and dynamic metrics confirms that geometric information alone is insufficient for robust, query-target-aligned localization and tracking.

[0104] For the cue-based baseline model, the cue-based benchmark SEL model directly locates the target event based on the text description. This method fuses audio representations with CLAP-based text embeddings through multi-channel audio and text queries, and predicts the discretized azimuth distribution for each frame. Under mixed conditions, while the SEL model benefits from text conditionation, it remains limited by interference and ambiguous audio-video-text associations. Correspondence (MAE = 5.37°, MOTA) = 0.07). Switching to Sel-Joint or Clean input significantly enhances frame-level localization capabilities (MAE is approximately 2.8° in both cases) and improves trajectory metrics (MOTA). (Approximately 0.16 to 0.17). However, the SEL model predicts only one orientation distribution per frame for a given text embedding; when there are multiple sound source matching queries, it often focuses on the dominant target or averages multiple directions, resulting in secondary trajectories being missed or trajectories being unstable, a finding consistent with the results of the original benchmark.

[0105] The model in this embodiment is specifically designed to handle at least two concurrent target sound sources in each text query. When the raw mixed data is used as input, the model achieves a MAE of 0.98° and a MOTA of 0.92. Even when receiving Sel-Joint or Clean input, its performance significantly outperforms the SEL model. This is attributed to the combination of query conditional cueing-guided selective attention module front-end and azimuth DoA heatmap, which presents multiple peaks of the text matching source in each frame, rather than compressing it into a single discrete angle. This results in a clearer and more stable trajectory, especially in frames with two target sources.

[0106] The following describes the ablation experiment in this embodiment.

[0107] First, a system-level ablation experiment was conducted to evaluate the contribution of the main design choices in the cue-guided localization model (SelectTSL). Table 2 categorizes the variants into five types: (1) the coupling method between the selection phase and the angle of arrival (DoA) phase (using extracted target amplitude versus mixed amplitude), (2) the type of spatial cues provided to the DoA estimator (full IPD + ILD, ILD removed, IPD removed, or spatial cues completely removed), (3) the number of extracted information embedded in the DPRNN block output, denoted by EIE; (4) the internal architecture of the DoA module (whether it includes a cross-attention mechanism and a fusion layer FuseLayer), and (5) whether it includes a cardinality head. Only the parameters shown in the "Settings" column are modified in each row, while the other components remain in their original configurations.

[0108] System-level architecture ablation experiments: Coupling and spatial cue variants demonstrate that both target-dependent features and explicit multichannel cues are indispensable. Even with information extraction embedding disabled ( Conditioning the DoA estimator to separate targets still significantly outperforms the mixed-direction-angle variant using mixed amplitude (MAE = 1.21° vs. 3.32°, MOTA). = 0.88 vs. 0.49). Discarding ILD has already led to a significant performance degradation, while removing IPD or all spatial cues produces a huge error (MAE ≥ 3.89°) and MOTA. The significant reduction confirms the crucial role of phase space information in reverberant multi-source scenarios.

[0109] The remaining variants explore the choices of time and structural design. A few EIE blocks are sufficient: for the entire model. Best performance Their performances are similar, while 3 or 4 will monotonously worsen MAE and MOTA. This indicates that the additional taps primarily inject noise. In the backbone network, removing cross-attention in DPRNN is more detrimental to trajectory metrics than simply replacing the FuseLayer with a concatenation, suggesting that cross-modal interactions are more important than the precise fusion operator. Removing cardinality heads is particularly harmful to multi-object tracking: compared to optimal configurations ( Compared to MOTA, The OSPA decreased from 0.89 to 0.58, while the OSPA increased from 3.1 to 6.3, demonstrating that explicit source count prediction is crucial for stabilizing target trajectories. Finally, training the selection module and DoA head separately as a frozen two-stage pipeline also lagged behind end-to-end optimization (MAE=1.28°, F1=0.90, MOTA). =0.82), which indicates that the cue-guided selective attention module needs to be jointly trained to generate embedding vectors based on extracted information that truly meet the DoA objective.

[0110] Table 2. Model performance on different indicators in ablation experiments

[0111] The following details the ablation experiment using only prompts. Based on Table 3, to analyze the roles of text and audio prompts in the fusion layer, ablation experiments using only prompts were designed: while keeping the mixed signal encoded, one conditional branch was masked to all-zero embeddings, forcing the model to rely entirely on the remaining modalities. For the plain text conditionalization experiment, the text branch was kept active while the audio embeddings were cleared. Rewritten text was generated for each category using Qwen2.5-7B and categorized as CLAP. The similarity between CLAP and standard text was divided into three levels: high, medium, and low ([0.85, 1.00], [0.70, 0.85], [0.50, 0.70]). Performance monotonically decreased with decreasing similarity: compared to standard plain text (Text-only) (MAE = 1.12°, MOTA...). = 0.86), the highest level corresponding to MAE has deteriorated to 1.94°, MOTA = 0.74; the lowest level's MAE drops to 3.48°, MOTA = 0.43, which indicates that semantically similar cues are crucial in the absence of audio cues.

[0112] For pure audio conditionalization, text embeddings are disabled, and a 1-second audio cue is input via an audio cue branch. Four strategies were compared: extracting the first 1 second from segments of the same category (Audio-Only / Front), randomly sampling 1-second segments (Pos-Rand), time-stretching the first 1 second (Aug-Front), and extracting the first 1 second from other segments of the same category (XClip-Front). The best pure audio performance (MAE = 1.57°, MOTA) was obtained using the segment start point. = 0.81); random or cross-fragment cues significantly reduce its effectiveness (MAE≈2.30°, MOTA). The value is approximately 0.51, while simple time stretching falls between the two. This indicates that under pure audio conditionation, the model is most sensitive to temporal misalignment and cross-instance mismatch, while moderate temporal distortion is less harmful.

[0113] Table 3. Performance of text and audio ablation on different performance metrics

[0114] The following details the ablation experiments of the DoA estimator. The proposed DoA estimator is ablated to evaluate the IPD enhancer, semantic-acoustic branch, and feature refinement module (FRB). Table 4 summarizes the MAE, frame-level F1, and MOTA results. And OSPA-T; in the following sections we will focus on MAE and MOTA. As a representative static indicator.

[0115] For the IPD enhancer and semantic-acoustic design (A1–A4), removing the IPD enhancer (A1) results in a significant performance degradation relative to the full model (MAE = 0.98° → 2.10°, MOTA). =0.92 →0.69), indicating that denoising and optimizing the original IPD is crucial for robust localization. Using only the semantic branch (A2) or only the acoustic branch (A3) is also suboptimal: both lag behind the full system, with the semantic branch alone outperforming the acoustic branch alone, suggesting that textual conditional embeddings carry strong class cues but still benefit from fusion with spatial features. Replacing the enhanced IPD with the direct IPD input (A4) performs the worst in this group (MAE = 2.71°, MOTA). = 0.53), further highlighting the importance of the IPD enhancement module.

[0116] For FRB variants (B1 to B7), experimental results show that iterative optimization is necessary, but must be carefully designed. Complete removal of the FRB (B1) significantly degrades performance (MAE = 3.80°, MOTA). =0.38), indicating that processing via the backbone network alone is insufficient. Adjusting the FRB depth (B2–B4) shows that the default depth ×2 achieves the optimal balance: configurations that are too shallow or too deep fail to improve MAE and MOTA. This could even worsen them. Ableasing the internal components (B5–B7) by removing convolutional layers, SE branches, or residual connections further deteriorated both metrics, confirming that the complete FRB structure containing all three components is the most effective.

[0117] Table 4. Performance of the DoA estimator on different performance metrics in ablation experiments.

[0118] For the ablation experiment of the cardinality head: as mentioned earlier, the DoA estimation is decoupled from the source cardinality prediction: the DoA head generates a frame-level posterior map over the Θ azimuth interval, while the independent cardinality head predicts... The distribution of . During the reasoning phase, take . And select The former The local maxima are used as DoA estimates (“Ours” in Table 5), thus the cardinality head provides a discrete prior for the number of sound sources without explicitly injecting cardinality embeddings. To verify whether tighter coupling is beneficial, this design is compared with two alternatives that feed cardinality information back to the DoA head. In the “Card-head attention” variant, the predicted count is mapped to the embedding of a multi-head attention block that queries the angle of arrival feature, and its attention outputs are summed via residual connections. In the “Embed” variant, the three-way cardinality probabilities are used to generate an embedding vector through a small MLP, which is concatenated with the DoA head input.

[0119] Table 5. Ablation experiments on the interaction between the cardinality header and the DoA header.

[0120] Table 5 shows that both variants were worse than the proposed scheme, despite using more parameters; MAE increased from 0.98° to 1.52° / 1.87°, while MOTA... The values ​​decreased from 0.92 to 0.83 / 0.72. Compared to the "no cardinality head" configuration in Table 4, all three designs benefited from a dedicated source cardinality predictor, but directly injecting cardinality embeddings into the DoA features often distorted the spatial structure. Therefore, adopting... The chosen decoupling design provides the best overall trade-off while keeping the architecture and decoding rules simple.

[0121] The robustness and generalization evaluation are described in detail below.

[0122] To study robustness under different acoustic conditions and varying acoustic complexity, three difficulty levels (A / B / C) were defined by progressively increasing room size and widening the reverberation time range; specific parameter ranges are shown in Table 6. Level A uses a compact room with a shorter reverberation time, Level B moderately increases both, and Level C covers the largest and most asymmetrical room with the longest and most varied reverberation time. As shown in Table 6, performance gradually decreases from A to C. MAE increases from 2.36° to 3.03°, and MOTA... The value decreased from 0.61 to 0.46, with other static and trajectory-level metrics showing a similar trend. Larger, more asymmetrical rooms introduce stronger spatial ambiguity, while longer, more varied rooms blur binaural cues; these factors collectively define the robustness range of our model with increasing acoustic variability.

[0123] Table 6. Performance and acoustic configurations for different difficulty levels

[0124] To assess robustness under different motion dynamics, four velocity ranges (A, B, C, and D) are defined to control the instantaneous angular velocity of the source. Range A corresponds to slow motion within ±5°, range B corresponds to medium-speed motion within ±15°, and ranges C and D correspond to fast motion within ±30° and ±50°, respectively.

[0125] Table 7 shows that tracking performance decreases with increasing motion speed and irregularity. MAE remains low during slow and medium-speed motion (1.20° in A, 1.14° in B), but rises to 1.55° and 2.32° in C and D, respectively; while MOTA... The value dropped from 0.96 to 0.53, and other metrics showed a similar downward trend. Higher angular velocities lead to greater changes in orientation between frames, disrupting the temporal smoothness upon which the model depends. Therefore, gradual or medium-speed motions are handled well, while sudden high-speed rotations remain challenging.

[0126] Table 7. Performance within different speed ranges

[0127] For robustness to motion and cue conditions, the robustness was further tested in the presence of spatial dynamics and the possibility of missing text cues, see Table 8.

[0128] For motion scenarios, the motion of the background and target is independently controlled, forming four states: stat-stat (static noise, static sound source), mov-stat (dynamic noise, static sound source), stat-mov (static noise, dynamic sound source), and mov-mov (dynamic noise, dynamic sound source). Regarding cue data, additional uncued audio segments are constructed (i.e., descriptive text is removed, but the target sound source is still retained in the mixed audio), and the proportion of uncued samples is adjusted to conduct practical analysis. .

[0129] In the exercise group, the best performance was achieved in the completely static condition (stat-stat) (MAE = 0.15°, MOTA). =0.98), with only background dynamic noise having a weaker effect (mov–stat), while a moving target sound source (stat–mov, mov–mov) will increase the MAE to about 0.97° and reduce the MOTA. Reduced to approximately 0.92. In the silent group, the system remains usable even when text is frequently missing: with r The change in MAE remained between 2.53° and 1.00°, and MOTA The value remained between 0.60 and 0.92, indicating that the model can still perform target localization and tracking primarily based on audio cues even in the absence of descriptive text.

[0130] Table 8. Ablation and robustness under different motion and cueing conditions

[0131] For real-world data, the generalization ability was further evaluated on a subset of real rooms in TAU-SRIR, using the same protocol as in the simulation. As shown in Table 9, the average performance for each room was MAE = 2.62° and MOTA = 2.62°. =0.77 and OSPA-T=2.85, indicating that it has strong migration capability to the measured space.

[0132] The room-level trends are related to the measured room attributes and trajectory design. PB132 is a small, carpeted classroom using a circular trajectory, which implies a shorter effective reverberation time and a stable direct-to-reverberation ratio; this results in stronger correlation and a higher score (MAE of 0.82°, MOTA). The MAE is 0.91, while OSPA-T is 1.25°. SA203 is a sloping-floor lecture hall with linear trajectories at multiple distances, which introduces stronger morning / evening reflections and greater DRR variations, resulting in increased angular deviation and trajectory fragmentation (MAE = 6.89°, MOTA = 0.91, OSPA-T = 1.25°). = 0.64, OSPAT = 4.26°). SE203 (a large classroom with a hard floor and linear trajectories) exhibits different failure modes: frame-level DoA remains clear (MAE = 0.79°), but repetitive crossovers and mirror clutter cause association breakdowns, resulting in the highest OSPA-T (4.74°). Overall, the model maintains accurate DoA estimates in different real-world rooms, and the performance differences can be explained by room geometry and trajectory complexity.

[0133] Table 9. Real-world performance

[0134] It should be noted that the cue-guided localization model used in the above training and experiments is a full model trained using both audio cues and text cues. To better understand the effects of different cue designs, three system variants with different cue input scenarios were also trained in the experiments, resulting in a total of four models with different cue input scenarios: (1) full model (using both audio cues and text cues); (2) text-only model (using only text cues); (3) audio-only model (using only audio cues); and (4) no-cue model. The frame-by-frame angle-of-arrival (DoA) posterior distribution generated by the four independently trained models was visualized for analysis. Figure 5Each row of data evaluates four models based on the same mixed recording, and the rightmost coordinate graph overlays their decoded DoA trajectories with the ground truth (GT) values.

[0135] In the first row, only the full (text + cue) model was able to track challenging non-stationary trajectories, while the plain text model and the plain audio model failed to provide reliable trajectories, and the uncue model generated completely wrong trajectories.

[0136] In the second row, the complete model and the plain text model closely match the real trajectory, while the plain audio model and the unannotated model show significant deviations and trajectory fragmentation.

[0137] In the third row, the full model and the pure audio model basically captured the motion pattern, while the pure text model and the unannotated model again failed to track the source signal.

[0138] These qualitative results indicate that text and audio cues provide complementary information, and using both together yields the most robust localization performance.

[0139] In addition, Figure 6 The DoA trajectory estimated by the entire model is visualized in polar coordinates, with the radial axis encoding time (0-75 frames) and the angular axis encoding DoA (0-180°). Figure (a) shows a two-source recording with wide-angle motion; the model follows the target across the entire range. Figure (b) shows a single-source sequence with smooth motion, where the predictions form a continuous trajectory that closely matches the ground truth. Figure (c) shows a two-source sequence containing quiescent periods that form three non-overlapping active segments; the model quickly re-locks onto the correct orientation whenever the target reactivates. These examples demonstrate that cue-guided localization models can handle two-source, single-source, and quiescent-period scenarios while maintaining temporally consistent trajectories.

[0140] In summary, this invention provides a cue-based multi-channel selective SSL method that learns target-aware representations by utilizing cue conditional selectivity for DoA estimation. The cue-guided localization model combines textual and audio cues with a target-aware multi-spatial cue table, enabling selective estimation of the DoA trajectory of the target sound source. Extensive experiments on large-scale synthetic benchmarks and real-world recording data in dynamic multi-source scenarios with varying numbers of moving speakers demonstrate that the proposed method consistently outperforms strong baseline methods and exhibits robust generalization capabilities to real-world acoustic environments.

[0141] Computer device embodiment: The computer device in this embodiment includes a processor and a memory. The memory stores a computer program, and when the processor executes the computer program, it implements the above-described embodiment of the selective sound source localization method based on prompting guidance.

[0142] A computer device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that a computer device may include more or fewer components, or a combination of certain components, or different components; for example, a computer device may also include input / output devices, network access devices, buses, etc.

[0143] For example, a processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microcontroller or any conventional processor. The processor is the control center of a computer device, connecting all parts of the computer device through various interfaces and lines.

[0144] The memory can be used to store computer programs and / or modules. The controller implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory, and by accessing data stored in the memory. For example, the memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (e.g., sound receiving function, sound-to-text function, etc.), etc.; the data storage area may store data created based on the use of the mobile phone (e.g., audio data, text data, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0145] Finally, it should be emphasized that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A selective sound source localization method based on prompting guidance, characterized in that, Includes the following steps: Acquire a mixed audio signal; wherein the source of the mixed audio signal includes a target sound source and an interference sound source; Obtain prompt input, the prompt input including audio prompts and / or text prompts; wherein, the audio prompts are associated with the target sound source; the text prompts are associated with the target sound source; The pre-trained prompt-guided localization model processes the mixed audio signal and the prompt input, outputs a DoA posterior map and a sound source cardinality, and determines the DoA trajectory corresponding to the target sound source based on the DoA posterior map and the sound source cardinality. The cue-guided localization model includes a cue-guided selective attention module and a DoA estimator. The cue-guided selective attention module processes the mixed audio signal and the cue input to obtain a target-specific representation corresponding to the input data. The DoA estimator outputs the DoA posterior map and the sound source cardinality based on the target-specific representation. The cue-guided selective attention module includes an audio encoder, a cue encoder, a fusion layer, and an extraction network. The audio encoder outputs amplitude cues and spatial cues based on the mixed audio signal. The spatial cues include inter-channel phase difference (IPD) and inter-channel intensity difference (ILD). The cue encoder extracts the audio cue and the text cue to obtain a guidance vector. The fusion layer fuses the amplitude cues and the guidance vector to obtain a fusion cue. The extraction network extracts target amplitude information and extracted feature stacks based on the guidance vector and the fusion cue. The DoA estimator includes an IPD enhancer, a progressive optimization and temporal modeling module, and a prediction head. The IPD enhancer is used to obtain an enhanced IPD based on the spatial cues, the target amplitude information, and the extracted feature stacking. The progressive optimization and temporal modeling module processes the enhanced IPD, the inter-channel intensity difference (ILD), and the target amplitude information to obtain alignment features. The prediction head includes two parallel heads, which process the alignment features to obtain the DoA posterior map and the source cardinality.

2. The selective sound source localization method based on prompting and guidance as described in claim 1, characterized in that: The fusion layer includes two FiLM cascaded structures driven by the guide vector. The two FiLM cascaded structures fuse the amplitude cue and the guide vector to obtain the fusion cue.

3. The selective sound source localization method based on prompting guidance as described in claim 2, characterized in that: The extraction network includes multiple stacked DPRNN blocks; in the multiple stacked DPRNN blocks, cross attention is periodically injected between the current audio feature and the guiding vector to obtain the extraction information embedding corresponding to each DPRNN block, and the extraction feature stack is determined based on the extraction information embedding.

4. The selective sound source localization method based on prompting guidance as described in claim 3, characterized in that: The number of DPRNN blocks is 6.

5. The selective sound source localization method based on prompting guidance as described in claim 1, characterized in that: The IPD enhancer includes acoustic and semantic branches that run in parallel; The acoustic branch obtains acoustic features based on the target amplitude information and the spatial cues; the semantic branch obtains semantic features based on the stacking of the extracted features. The IPD enhancer obtains the enhanced IPD based on the acoustic features and the semantic features.

6. The selective sound source localization method based on prompting guidance as described in claim 4, characterized in that: The progressive optimization and timing modeling module includes multiple FR blocks, multiple TCN blocks, and a bidirectional cyclic BiGRU. The outputs of the multiple FR blocks are connected to the multiple TCN blocks, and the outputs of the multiple TCN blocks are connected to the bidirectional cyclic BiGRU. The bidirectional cyclic BiGRU outputs the alignment feature.

7. The selective sound source localization method based on prompting guidance as described in any one of claims 1 to 6, characterized in that: The total loss function used in the training of the prompt-guided localization model is a weighted sum of selection loss, DoA prediction loss and sound source cardinality prediction loss. The selection loss is the average negative SI-SNR on the two microphone channels.

8. A computer device comprising a processor and a memory, characterized in that: The memory stores a computer program, which, when executed by the processor, implements the selective sound source localization method based on prompting guidance as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and device for selectively positioning sound source based on visual prompt, medium and product

    CN120161404A

  • Apparatus, Method or Computer Program for estimating an inter-channel time difference

    US20210012784A1