Three-dimensional sound event positioning and detection method based on audio-guided visual attention
A three-dimensional sound event localization and detection method that guides visual attention through audio, through the adaptive fusion of audio and visual features, solves the performance limitation problem of the existing audio-visual SELD model under complex acoustic conditions, and improves the sound event localization and detection accuracy and source distance estimation capability.
Patent Information
- Application Number
- CN202510731807.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-16
AI Technical Summary
Existing audio-visual SELD models have difficulty in effectively fusing audio and visual information under complex acoustic conditions, resulting in limited performance in sound event localization and detection, especially in source distance estimation.
A three-dimensional sound event localization and detection method with audio-guided visual attention is adopted. Features are extracted through audio encoder and visual encoder, combined with audio-guided visual attention module and fusion module, and the sound source coordinate estimation branch to achieve adaptive fusion of audio and visual features and effective utilization of multimodal information.
It improves the performance of localization and detection of sound events in three-dimensional space, especially the precision and accuracy in complex scenes, solves the problem of insufficient fusion of audio and visual information in existing technologies, and enhances the ability of source distance estimation.
Smart Images

Figure CN120654180A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio-visual fusion and sound event positioning detection, and in particular to a three-dimensional sound event positioning and detection method based on audio-guided visual attention. Background Art
[0002] Sound Event Localization and Detection Evaluated (SELD) is the task of identifying and locating sound events in a given environment. It combines sound event detection (SED) and direction of arrival (DOA) estimation. With technological advancements, SELD is becoming increasingly important in areas such as smart home systems, audio surveillance, robotics, and immersive audio experiences, as these applications require understanding both the type of sound and its source in three-dimensional space. In recent years, deep neural network (DNN) methods have demonstrated great potential for SELD, with much research focusing on acoustic feature extraction, output representation format design, and network architecture optimization. Researchers have explored various approaches to improve SELD performance in acoustic feature extraction. For example, some studies have combined log-Mel spectrograms with generalized cross-correlation phase transforms between microphone pairs and intensity vectors in log-Mel space for acoustic feature extraction. Other studies have proposed using spatially enhanced log-spectral features and bio-inspired gammatone auditory filters to improve SELD performance.
[0003] Modeling methods for output representation include multi-task learning frameworks for SED and DOA estimation, Cartesian Direction of Arrival (ACCDOA) representations coupled with activity probabilities, Cartesian Direction of Arrival (DOA) representations coupled with multiple activity probabilities (multi-ACCDOA representations), and trajectory-related output formats. Furthermore, recent studies have proposed introducing multiple SELD outputs based on angle-distance via position-guided detection.
[0004] The design of network architecture is also one of the important directions of SELD research. Architectures such as convolutional recurrent neural networks (CRNN), ResNet-Conformer, CST-Former, MFF-EINV2 and SELD-Mamba have been proposed one after another to model the spatiotemporal relationship between different sounds to further improve the performance of SELD. Faced with complex acoustic conditions, such as high noise and strong reverberation scenes, the effect of using only audio modalities for sound detection and positioning is difficult to meet actual needs. Audio-visual (AV) SELD was first introduced in the DCASE 2023 challenge and expanded to a three-dimensional (3D) SELD task including sound source distance estimation (SDE) in the DCASE 2024 challenge, which greatly increased the difficulty of sound localization and detection.
[0005] However, existing audio-visual SELD models fail to outperform audio-only models, prompting researchers to explore more effective methods to extract and utilize useful information from both audio and video modalities. In addition, neural network-based SDE remains a relatively understudied field, with most methods adopting classification-based strategies and fewer using regression methods.
[0006] In summary, despite significant progress in the field of SELD, challenges remain in effectively integrating audio and visual information, improving source distance estimation accuracy, and optimizing network architecture. Addressing these challenges and promoting the further development of SELD technology in practical applications is of great significance.
[0007] In view of this, the present invention is proposed. Summary of the Invention
[0008] The purpose of the present invention is to provide a three-dimensional sound event localization and detection method based on audio-guided visual attention. By introducing an audio-guided visual attention mechanism and two audio-visual fusion methods, combined with the sound source coordinate estimation branch, the positioning and detection performance of sound events in three-dimensional space are effectively improved, thereby solving the above-mentioned technical problems existing in the prior art.
[0009] The purpose of the present invention is achieved through the following technical solutions:
[0010] A 3D sound event localization and detection method based on audio-guided visual attention is adopted, which is composed of an audio encoder, a visual encoder, an audio-guided visual attention module, a fusion module, a context network and three output branches, including:
[0011] The audio features of the audio data are extracted through the audio encoder and output to the audio-guided visual attention module after context aggregation;
[0012] The visual features of the video data are extracted through the visual encoder and output to the audio-guided visual attention module;
[0013] The audio-guided visual attention module guides the focus of visual features through audio features, obtains the focused visual features of the visual area related to the sound source, and outputs the focused visual features as the visual sound localization result through an output branch;
[0014] The audio features and visual features are fused through the fusion module, and the multimodal features obtained after fusion are input into the context network to obtain the sound event detection results and sound source coordinate estimation results, which are output through two output branches respectively.
[0015] Compared with the prior art, the method for locating and detecting three-dimensional sound events based on audio-guided visual attention provided by the present invention has the following beneficial effects:
[0016] By introducing an audio-guided visual attention mechanism through an audio-guided visual attention module, visual features related to the sound source position are adaptively learned, thereby enhancing the localization and detection performance of sound events in three-dimensional space. Through audio-visual fusion, the sound source coordinate estimation branch is used to simultaneously estimate the direction and distance of the sound source, thereby effectively processing the sound event localization and detection tasks in three-dimensional space. This solves the problem that existing sound event localization and detection methods fail to effectively capture the complementarity of multimodal data when fusing audio and visual information, and have deficiencies in source distance estimation, resulting in limited performance in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 A flowchart of a method for locating and detecting three-dimensional sound events based on audio-guided visual attention provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0019] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the specific content of the present invention. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments, and do not constitute a limitation of the present invention. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0020] First, the following terms may be used in this article:
[0021] The term “and / or” means that either or both of them can be realized at the same time. For example, X and / or Y includes both “X” or “Y” and “X and Y”.
[0022] The terms "include," "comprises," "contains," "has," or other similar expressions should be interpreted as non-exclusive. For example, "including certain technical features (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, procedures, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products, or manufactured articles, etc.) should be interpreted as including not only the technical features explicitly listed, but also other technical features known in the art that are not explicitly listed.
[0023] The term "consisting of" excludes any technical features not explicitly listed. If used in a claim, this term renders the claim closed, excluding any technical features other than those explicitly listed, except for conventional impurities associated with them. If this term appears only in a clause of a claim, it limits only the elements explicitly listed in that clause; elements listed in other clauses are not excluded from the claim as a whole.
[0024] Unless otherwise specified or limited, the terms "mounted," "connected," "connect," and "fixed" should be interpreted broadly. For example, they can refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediary; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in this document based on specific circumstances.
[0025] The terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings and are only for the convenience and simplification of description, and do not explicitly or implicitly indicate that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore should not be understood as a limitation to this document.
[0026] The scheme provided by the present invention is described in detail below. The contents not described in detail in the examples of the present invention belong to the prior art known to professionals in this field. If specific conditions are not specified in the examples of the present invention, they are carried out according to conventional conditions in the field or conditions recommended by the manufacturer. If the manufacturer of the reagents or instruments used in the examples of the present invention is not specified, they are all conventional products that can be purchased commercially.
[0027] like Figure 1 As shown, an embodiment of the present invention provides a 3D sound event localization and detection method based on audio-guided visual attention, which uses a 3D sound event localization and detection method composed of an audio encoder, a visual encoder, an audio-guided visual attention module, a fusion module, a context network, and three output branches, including:
[0028] The audio features of the audio data are extracted through the audio encoder and output to the audio-guided visual attention module after context aggregation;
[0029] The visual features of the video data are extracted through the visual encoder and output to the audio-guided visual attention module;
[0030] The audio-guided visual attention module guides the focus of visual features through audio features, obtains the focused visual features of the visual area related to the sound source, and outputs the focused visual features as the visual sound localization result through an output branch;
[0031] The audio features and visual features are fused through the fusion module, and the multimodal features obtained after fusion are input into the context network to obtain the sound event detection results and sound source coordinate estimation results, which are output through two output branches respectively.
[0032] Preferably, in the above method, the audio encoder adopts an 18-layer ResNet network consisting of 4 convolutional blocks and 2 convolutional layers, wherein each convolutional block contains 4 3×3 convolutional layers, and each convolutional block is sequentially connected to a batch normalization layer and a ReLU activation function;
[0033] In the first three convolution blocks, maximum pooling is performed along the frequency bins with kernel sizes of 4, 4, and 2, respectively, and no pooling operation is performed in the time dimension;
[0034] A convolutional layer with a kernel size of 1×1 is used after the four convolution blocks, outputting 256 channels;
[0035] The audio encoder generates a tensor of shape 5T×512.
[0036] Preferably, in the above method, the audio encoder uses the logarithmic Mel-spectrogram and intensity vector in the first-order Ambisonics data as input data, the input data includes 7 channels and 64 frequency points, and outputs the result with a time resolution of 100 ms.
[0037] Preferably, in the above method, the visual encoder adopts a pre-trained ResNet-18 model that can extract deep visual features of video frames. The video frames are input into the ResNet-18 model at a frame rate of 10fps, and the feature map before the global average pooling layer is extracted, and the feature map shape is T×512×7×7.
[0038] Preferably, in the above method, the audio encoder performs context aggregation on the audio features and outputs the results to the audio-guided visual attention module in the following manner, including:
[0039] Contextually aggregate the audio features by averaging them over a predetermined time window to match the frame rate of the visual features.
[0040] The average audio features within a predetermined time window are output to the audio-guided visual attention module.
[0041] Preferably, in the above method, the predetermined time window is four time windows including 10 audio feature frames, 20 audio feature frames, 30 audio feature frames, and 40 audio feature frames.
[0042] Preferably, in the above method, the audio-guided visual attention module guides the focus of the visual features through the audio features in the following manner to obtain the focused visual features of the visual area related to the sound source, including:
[0043] The spatial attention weight m is calculated by the following formula using the input visual features and the audio features after context fusion: t ,for:
[0044] m t =σ2(W o σ1(c t )) (1);
[0045] In formula (1), is the attention weight of the visual feature space of the t-th frame, represents a set of real numbers, H represents the height of the visual feature, and W represents the width of the visual feature; σ1(·) and σ2(·) are the hyperbolic tangent function and the sigmoid function respectively; represents the learnable parameters; c t The calculation formula is:
[0046]
[0047] In formula (2), and Both represent learnable parameters, and d represents the dimensional information of audio features and video features after transformation through the fully connected layer; is the audio feature of the tth frame after passing through the audio encoder, d a is the audio feature dimension; and are two transformation matrices that pass through a fully connected layer with ReLU activation; is the visual feature map of the tth frame after the video encoder, is the visual feature after shape transformation, d v is the channel dimension of the visual feature map;
[0048] The weighted visual features are calculated as follows: The attention visual features of the visual area related to the sound source are obtained as follows:
[0049]
[0050] In formula (3), represents matrix multiplication,
[0051] Preferably, in the above method, the context network adopts an 8-layer Conformer network.
[0052] Preferably, in the above method, the fusion module fuses the audio features and the visual features of interest in the following manner, including:
[0053] The audio features and the visual features are concatenated to obtain a multimodal feature vector, which is then subjected to dimensionality reduction processing by a fully connected layer to obtain the input features of the input context network.
[0054] Alternatively, the audio features and the attention visual features are respectively reduced in dimension through separate fully connected layers to obtain audio features and attention visual features that match in dimension. The audio features and attention visual features that match in dimension are directly added together to obtain fused features as the input features of the input context network.
[0055] Preferably, in the above method, the detection model in the method is trained on the STARSS23 dataset and the synthetic dataset after audio channel enhancement, and is trained through multiple rounds of iterations until convergence. During the training process, the model parameters are adjusted by minimizing the total loss. To optimize, we use the Adam optimizer to update the parameters, and adopt a three-stage learning rate strategy during training;
[0056] The total loss for:
[0057] In formula (1), The binary cross entropy loss for sound event detection is:
[0058]
[0059] In formula (2), o 1tn is the true value of the activity probability of the nth type of sound event in the tth frame; is the predicted value of the activity probability of the nth type of sound event in the tth frame; T is the number of time frames; N is the number of sound event categories;
[0060] In formula (1), is the mean square error loss of the sound source coordinate estimation, which is:
[0061]
[0062] In formula (3), o 2tn is the true value of the sound source coordinates of the nth type of sound event in the tth frame; is the predicted value of the sound source coordinates of the nth type of sound event in the tth frame; 1tn is the true value of the activity probability of the nth type of sound event in the tth frame;
[0063] In formula (1), λ is the weight used to balance the mean square error loss of the visual sound localization task, The mean square error loss for the visual sound localization task is:
[0064]
[0065] In formula (4), o 3tn is the true localization map of the sound source of the nth type of sound event in the tth frame; It is the sound source prediction localization map of the n-th type sound event in the t-th frame.
[0066] In summary, the method of the embodiment of the present invention effectively improves the positioning and detection performance of sound events in three-dimensional space by introducing an audio-guided visual attention mechanism and an audio-visual fusion method, combined with a sound source coordinate estimation branch.
[0067] In order to more clearly demonstrate the technical solution and technical effects provided by the present invention, the solution provided by the embodiment of the present invention is described in detail with reference to specific embodiments below.
[0068] Example 1
[0069] like Figure 1 As shown, an embodiment of the present invention provides a three-dimensional sound event localization and detection method based on audio-guided visual attention, which adopts a three-dimensional sound event localization and detection model composed of an audio encoder, a visual encoder, an audio-guided visual attention module, a fusion module, an 8-layer Conformer as a context network and three output branches. These modules work together to achieve accurate positioning and detection of sound events. The audio encoder and the visual encoder extract the audio features of the audio data and the visual features of the video data respectively. The audio-guided visual attention module guides the focus of the visual features through the audio features after context fusion to obtain the focus visual features of the visual area related to the sound source. The fusion module effectively fuses the audio features and the focus visual features, and inputs them into the context network for processing. The three output branches respectively output the processed sound event detection results, the sound source coordinate estimation results and the visual sound localization task.
[0070] In the above model, ResNet-Conformer serves as the backbone network, providing the model with powerful feature extraction and processing capabilities. The audio and visual encoders are responsible for extracting high-level features from audio and video data, respectively. The audio-guided visual attention module uses audio information to guide the focus of visual features, enhancing the feature representation of visual areas related to sound sources. The fusion module effectively fuses audio and visual features to fully utilize the complementarity of multimodal information. The three output branches are used to handle sound event detection, sound source coordinate estimation, and visual sound localization (VSL), respectively, ensuring that the model can comprehensively and accurately locate and detect sound events.
[0071] The audio-guided visual attention module introduces an audio-guided visual attention mechanism. Through audio context aggregation and spatial attention calculation, the model can adaptively focus on visual areas related to the sound source location. Specifically, audio features are first context-aggregated to match the frame rate of visual features. By averaging features within a time window, audio features are obtained that are consistent with the frame rate of visual features. Then, the audio and visual features are transformed separately through fully connected layers, and spatial attention weights are calculated. These spatial attention weights are used to weight the visual features, highlighting the visual areas associated with the audio.
[0072] The fusion module adopts two audio-visual fusion methods to give full play to the complementary advantages of multimodal features. The first method is feature splicing, which splices the audio and visual features in the feature dimension to form a multimodal feature vector. This method can retain the independent information of the two modalities, but it will increase the feature dimension. Therefore, after splicing, it is processed by a fully connected layer for dimensionality reduction to meet the subsequent Conformer network input requirements. The second method is feature summation, which directly adds the audio and visual features to obtain fused features. Before summing, the audio and visual features are respectively reduced in dimensionality through separate fully connected layers to ensure that they match in dimension. These two fusion methods combine audio and visual information from different angles, providing the model with diverse feature representations, which helps to improve the model's ability to locate and detect sound events in complex scenes.
[0073] The audio encoder uses an 18-layer ResNet network consisting of four convolutional blocks and two convolutional layers. Each convolutional block contains four 3×3 convolutional layers followed by batch normalization and ReLU activation. A 3×3 convolutional layer is used before the convolutional block to convert the number of channels to 24. In the first three blocks, max pooling is performed along the frequency bins with kernel sizes of 4, 4, and 2, respectively, while no pooling is performed in the temporal dimension. A convolutional layer with a kernel size of 1×1 is used after the convolutional block, outputting 256 channels. Ultimately, the audio encoder generates a tensor of shape 5T×512. The audio encoder takes as input the log-mel spectrogram and intensity vector (IV) from first-order ambisonics data. These features effectively capture the spectral and spatial information of the audio signal. The input data contains 7 channels and 64 frequency bins, and the model outputs the results with a temporal resolution of 100ms. For an audio clip containing T frame labels, a frame shift of 20ms is used to extract acoustic features, and the input features of 5T frames are obtained.
[0074] The visual encoder uses a pre-trained ResNet-18 model to extract deep visual representations of video frames. To match the frame rate of the labels, video frames are fed into the ResNet-18 at 10 fps. Feature maps are extracted before the global average pooling layer. These feature maps are rich in spatial information and have a shape of T × 512 × 7 × 7. During training, the parameters of the visual encoder remain frozen to ensure the stability of the visual features and the effectiveness of pre-training.
[0075] The audio-guided visual attention module above calculates the spatial attention weight of the visual area and highlights the visual area associated with the audio. The specific calculation method is as follows:
[0076] First, the audio features are context-aggregated to match the frame rate of the visual features by averaging the features within a time window. The time window length of audio context aggregation for visual sound localization uses four context lengths of 10, 20, 30, and 40 audio feature frames. The average audio features in the context window are used as the input to the audio-guided visual attention module. The audio and visual features of the tth frame are represented as and First, the visual features are reshaped into The specific spatial attention weight m t It can be calculated as:
[0077] m t =σ2(W o σ1(c t )) (1);
[0078] in,
[0079] In the above formulas (1) and (2), the spatial attention weight and They are two transformation matrices implemented by fully connected layers with ReLU activation; and denotes the learnable parameters, σ1(·) and σ2(·) are the hyperbolic tangent function and the sigmoid function, respectively.
[0080] By using the sigmoid function to calculate the attention weights, we can better locate the spatially overlapping sounds. Finally, the weighted attention visual features It can be calculated as:
[0081]
[0082] In the above formula (3), represents matrix multiplication,
[0083] By exploiting audio-visual synchrony, the audio-guided visual attention module is expected to focus on visual areas related to the audio.
[0084] The contextual network described above uses an 8-layer Conformer network, a model that combines CNN and Transformer architectures to effectively capture local and global dependencies between audio and visual features. In this paper, the 8-layer Conformer network is used to further process the fused multimodal features to enhance the model's ability to locate and detect sound events.
[0085] The audio-guided visual attention module used in the above method is trained according to the following criteria. It is trained on the STARSS23 dataset and the synthetic dataset after audio channel enhancement, and it is trained through multiple rounds of iteration until convergence. During the training process, the model parameters are adjusted by minimizing the total loss The Adam optimizer is used to update the parameters, and a three-stage learning rate strategy is adopted during training.
[0086] The total loss function of training for:
[0087]
[0088] in, Binary Cross Entropy (BCE) loss for Sound Event Detection (SED); is the mean square error (MSE) loss of sound source coordinate estimation (SCE); MSE loss for the auxiliary visual sound localization (VSL) task;
[0089] The calculation method of each loss is as follows:
[0090] (1) SED loss
[0091]
[0092] (2) SCE loss
[0093]
[0094] (3) VSL loss
[0095]
[0096] In the above formulas, o 1tn is the true value of the activity probability of the nth type of sound event in the tth frame; is the predicted value of the activity probability of the nth type of sound event in the tth frame; 2tn is the true value of the sound source coordinates of the nth type of sound event in the tth frame; is the predicted value of the sound source coordinates of the nth type of sound event in the tth frame; 3tn is the true localization map of the sound source of the nth type of sound event in the tth frame; is the sound source prediction localization map of the n-th type of sound event in the t-th frame; λ is used to balance the VSL loss The weight of .
[0097] Example 2
[0098] This embodiment provides a three-dimensional sound event localization and detection method based on audio-guided visual attention. Applied to audio surveillance systems, this method improves monitoring accuracy and helps monitoring personnel quickly identify and respond to potential security threats by precisely locating sound sources. For example, in public places, the system can monitor and locate unusual sounds, such as quarrels or screaming, in real time, notifying security personnel for intervention.
[0099] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0100] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims. The information disclosed in the background technology section of this article is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or any form of implication that the information constitutes prior art already known to those skilled in the art.
Claims
1. A three-dimensional sound event localization and detection method based on audio-guided visual attention, characterized in that: The 3D sound event localization and detection model is composed of an audio encoder, a visual encoder, an audio-guided visual attention module, a fusion module, a context network, and three output branches, including: The audio features of the audio data are extracted through the audio encoder and output to the audio-guided visual attention module after context aggregation; The visual features of the video data are extracted through the visual encoder and output to the audio-guided visual attention module; The audio-guided visual attention module guides the focus of visual features through audio features, obtains the focused visual features of the visual area related to the sound source, and outputs the focused visual features as the visual sound localization result through an output branch; The audio features and visual features are fused through the fusion module, and the multimodal features obtained after fusion are input into the context network to obtain the sound event detection results and sound source coordinate estimation results, which are output through two output branches respectively.
2. The method for locating and detecting three-dimensional sound events based on audio-guided visual attention according to claim 1, characterized in that: The audio encoder uses an 18-layer ResNet network consisting of 4 convolutional blocks and 2 convolutional layers, where each convolutional block contains 4 3×3 convolutional layers, and each convolutional block is sequentially connected to a batch normalization layer and a ReLU activation function; In the first three convolution blocks, maximum pooling is performed along the frequency bins with kernel sizes of 4, 4, and 2, respectively, and no pooling operation is performed in the time dimension; A convolutional layer with a kernel size of 1×1 is used after the four convolution blocks, outputting 256 channels; The audio encoder generates a tensor of shape 5T×512.
3. The method for locating and detecting three-dimensional sound events based on audio-guided visual attention according to claim 1 or 2, characterized in that: The audio encoder takes the logarithmic Mel spectrogram and intensity vector in the first-order Ambisonics data as input data. The input data contains 7 channels and 64 frequency points, and outputs the result with a time resolution of 100ms.
4. The method for locating and detecting three-dimensional sound events based on audio-guided visual attention according to claim 1, characterized in that: The visual encoder uses a pre-trained ResNet-18 model that can extract deep visual features of video frames. The video frames are input into the ResNet-18 model at a frame rate of 10fps, and the feature map before the global average pooling layer is extracted. The feature map shape is T×512×7×7.
5. The method for locating and detecting three-dimensional sound events based on audio-guided visual attention according to claim 1, 2 or 4, characterized in that: The audio encoder contextually aggregates the audio features and outputs them to the audio-guided visual attention module in the following manner, including: Contextually aggregate the audio features by averaging them over a predetermined time window to match the frame rate of the visual features. The average audio features within a predetermined time window are output to the audio-guided visual attention module.
6. The method for locating and detecting three-dimensional sound events based on audio-guided visual attention according to claim 5, characterized in that: The predetermined time windows include four time windows including 10 audio feature frames, 20 audio feature frames, 30 audio feature frames, and 40 audio feature frames.
7. The method for locating and detecting three-dimensional sound events based on audio-guided visual attention according to claim 5, characterized in that: The audio-guided visual attention module guides the focus of visual features through audio features to obtain the focused visual features of the visual area related to the sound source in the following manner, including: The spatial attention weight m is calculated by the following formula using the input visual features and the audio features after context fusion: t ,for: m t =σ2(W o σ1((c t ))(1); In formula (1), is the attention weight of the visual feature space of the t-th frame, represents a set of real numbers, H represents the height of the visual feature, and W represents the width of the visual feature; σ1(·) and σ2(·) are the hyperbolic tangent function and the sigmoid function respectively; represents the learnable parameters; c t The calculation formula is: In formula (2), and Both represent learnable parameters, and d represents the dimensional information of audio features and video features after transformation through the fully connected layer; is the audio feature of the tth frame after passing through the audio encoder, d a is the dimension of audio features; and are two transformation matrices that pass through a fully connected layer with ReLU activation; is the visual feature map of the tth frame after the video encoder, is the visual feature after shape transformation, d v is the channel dimension of the visual feature map; The weighted visual features are calculated as follows: The attention visual features of the visual area related to the sound source are obtained as follows: In formula (3), represents matrix multiplication, 8. The method for locating and detecting three-dimensional sound events based on audio-guided visual attention according to claim 1, 2 or 4, characterized in that: The context network adopts an 8-layer Conformer network.
9. The method for locating and detecting three-dimensional sound events based on audio-guided visual attention according to claim 1, 2 or 4, characterized in that: The fusion module fuses the audio features and the visual features of interest in the following manner, including: The audio features and the visual features are concatenated to obtain a multimodal feature vector, which is then subjected to dimensionality reduction processing by a fully connected layer to obtain the input features of the input context network. Alternatively, the audio features and the attention visual features are respectively reduced in dimension through separate fully connected layers to obtain audio features and attention visual features that match in dimension. The audio features and attention visual features that match in dimension are directly added together to obtain fused features as the input features of the input context network.
10. The method for locating and detecting three-dimensional sound events based on audio-guided visual attention according to claim 1, 2 or 4, characterized in that: The detection model in the method is trained on the STARSS23 dataset and the synthetic dataset after audio channel enhancement, and is trained through multiple rounds of iterations until convergence. During the training process, the model parameters are adjusted by minimizing the total loss. To optimize, we use the Adam optimizer to update the parameters, and adopt a three-stage learning rate strategy during training; The total loss for: In formula (1), The binary cross entropy loss for sound event detection is: In formula (2), o 1tn is the true value of the activity probability of the nth type of sound event in the tth frame; is the predicted value of the activity probability of the nth type of sound event in the tth frame; T is the number of time frames; N is the number of sound event categories; In formula (1), is the mean square error loss of the sound source coordinate estimation, which is: In formula (3), o 2tn is the true value of the sound source coordinates of the nth type of sound event in the tth frame; is the predicted value of the sound source coordinates of the nth type of sound event in the tth frame; 1tn is the true value of the activity probability of the nth type of sound event in the tth frame; In formula (1), λ is the weight used to balance the mean square error loss of the visual sound localization task, The mean square error loss for the visual sound localization task is: In formula (4), o 3tn is the true localization map of the sound source of the nth type of sound event in the tth frame; It is the sound source prediction localization map of the n-th type sound event in the t-th frame.