An Audio-Visual Speaker Tracking Method Based on Semantic-Spatial Feature Fusion
By adopting semantic-spatial feature fusion technology in the audio-visual speaker tracking method and using the cross attention module to fusion audio-visual information, the problem of insufficient features of the existing method in complex environments is solved, and more accurate and robust sound source tracking is achieved.
Patent Information
- Application Number
- CN202411361389.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-09-27
AI Technical Summary
Existing audio-visual speaker tracking methods may fail or be insufficient to distinguish between targets and distractors in the presence of occlusion, illumination changes, complex backgrounds, noise, reverbs, and multiple sound sources, and the dependence of single-level features ignores the complementarity between different levels of features.
The audio-visual speaker tracking method based on semantic-spatial feature fusion is adopted, and the semantic feature encoding is extracted through the keyword recognition network of audio-visual fusion, and the image frame and acoustic signals are processed by combining the dual-stream network structure of visual branches and auditory branches, and the cross attention module is used to realize the information interaction and fusion of semantic-spatial features.
It improves the ability to express audio-visual features, achieves more accurate sound source tracking, and enhances robustness and accuracy in complex environments.
Smart Images

Figure CN119227003B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multimodal perception, and particularly relates to an audiovisual speaker tracking method based on semantic-spatial feature fusion. Background Art
[0002] In recent years, with the development of artificial intelligence and multimodal learning technologies, audiovisual speaker tracking, as a fundamental task in multimodal perception and intelligent systems, has become a research topic attracting much attention. Audiovisual speaker tracking aims to accurately track the position of a speaker by combining visual and auditory information, and has broad application prospects in fields such as video conferencing, intelligent monitoring, and human-computer interaction.
[0003] Currently, researchers have proposed a variety of speaker tracking methods, which can be roughly divided into methods based on visual features, methods based on auditory features, and audiovisual fusion methods. Methods based on visual features mainly utilize image data captured by a camera to achieve target tracking by detecting and tracking the facial features or other significant features of the speaker. These methods perform well in environments with sufficient light and clear vision, but their performance will significantly decline in the presence of occlusion, light changes, or complex backgrounds. Methods based on auditory features mainly utilize audio data captured by a microphone array to locate and track the speaker by analyzing the time-frequency features and spatial cues in the sound signal. These methods work well in situations with low noise and few environmental echoes, but in a noisy environment or in the presence of multiple sound sources, the stability and accuracy of the acoustic features will be greatly affected. Audiovisual fusion methods attempt to combine visual and auditory information to make up for the deficiencies of single-modal features. By fusing multimodal data, the robustness and accuracy of speaker tracking can be improved.
[0004] However, existing audiovisual speaker tracking methods mainly focus on using the spatial position information in visual and auditory data, such as establishing an appearance model of the target and performing similarity searches in various regions of the tracking space, or using spectral analysis and time-frequency analysis to extract the spatial cues of sound sources in multi-channel signals. These methods have high requirements for the quality and stability of visual features and acoustic cues. In the case of occlusion, deformation, and complex backgrounds in images, and noise, reverberation, and multiple sound sources in the environment, the spatial position information may become invalid or insufficient to distinguish the target from the interference. In addition, existing trackers usually rely on single-level features and ignore the complementarity between different-level features. High-level features have strong semantic information and robustness, but lack detailed information and spatial accuracy; low-level features have high spatial resolution and detailed information, but lack semantic information and anti-interference ability. Considering that multimodal signals contain rich semantic features, the advantages of different-level features can be fully utilized to provide more context information for the tracking target. Summary of the Invention
[0005] In view of the above technical problems, the present invention provides an audiovisual speaker tracking method based on semantic-spatial feature fusion, which improves the expression ability of audiovisual features and realizes more accurate sound source tracking.
[0006] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0007] An audiovisual speaker tracking method based on semantic-spatial feature fusion, comprising the following steps:
[0008] S1. Obtain the original audio-visual signal, and use an audiovisual fusion keyword recognition network to extract the semantic feature encoding in the audiovisual signal;
[0009] S2. Adopt a two-stream network structure including a visual branch and an auditory branch to process the image frames and acoustic signals respectively;
[0010] S3. Use a cross-attention module to realize the information interaction and fusion between two different sequences of semantic-spatial features.
[0011] The method for extracting the semantic feature encoding in the audiovisual signal in S1 is as follows:
[0012] S11. Convert the original audio-visual signal into a high-dimensional feature representation F a and F v ;
[0013] S12. Use an encoder based on Transformer to obtain the semantic features related to the keywords.
[0014] The processing process of the semantic feature encoder in S12 is expressed as:
[0015] Encoder(X) = X att + FFN(X att )
[0016] X att = X + MHA(Q, K, V)
[0017] FFN(X att ) = max(0, X att W1 + b1)W2 + b2
[0018] where X is the input of the encoder, that is, [CLS a ; F a or [CLS v ; F vAdd the positional encoding; FFN(·) is a fully connected feed-forward network used to enhance the model's fitting ability, including two linear transformation layers with a ReLU activation function in between; W and b represent the weight matrix and the basis vector respectively, and the subscripts indicate different layers; CLS a and CLS v are the CLS tokens for classification.
[0019] The method of using a two-stream network structure including a visual branch and an auditory branch to process image frames and acoustic signals respectively in S2 is as follows:
[0020] S21. Input a pair of images into the visual spatial feature extraction network, including the reference template I tpl and the search area image I s ; both are input into a fully convolutional network with shared weights for feature extraction;
[0021] S22. Perform a cross-correlation operation based on the convolutional operator to obtain the response map; the operation process of the visual network is defined as follows:
[0022] f v = Conv v (I tpl ) * Conv v (I s )
[0023] where Conv v are two fully convolutional networks with the same structure and shared weights, and * represents the cross-correlation operation;
[0024] S23. Use the camera model to project the acoustic cues onto the image plane;
[0025] S24. Then use a fully convolutional network structure similar to the visual network to embed the sound signal into the consistent localization space containing the position context; the operation process of the auditory network is defined as:
[0026] f a = Conv a (R Ω )
[0027] where R Ω is the manually extracted stGCF acoustic cue.
[0028] The method of using the cross-attention module to achieve information interaction and fusion between two different sequences of semantic-spatial features in S3 is as follows:
[0029] S31. Use the features of one modality as Q, and the features of the other modality as K and V for attention calculation;
[0030] S32. And use multi-head attention in residual form to integrate information from different sequences.
[0031] The method of using multi-head attention in residual form to integrate information from different sequences in S32 is as follows:
[0032] The definition of the CA mechanism is as follows:
[0033]
[0034] Among them, Q A is a linear transformation of A + E A and Q B is a linear transformation of B + E B E A and E B are position encodings used to supplement spatial position information. The definitions of K and V are similar. LN(·) represents layer normalization;
[0035] After that, two-step fusion is used to enhance the feature distinctiveness of different tracking targets in a multi-person scenario, including intra-modal and inter-modal stages. First, semantic-spatial feature fusion is performed within each modality, and then audiovisual fusion is performed between different modalities. The fusion in the first stage can be formally expressed as:
[0036] f sv = f v + MHA(Q sv , K v , V v )
[0037] f sa ' = f a + MHA(Q sa , K a , V a )
[0038] Among them, Q sv and Q sa are respectively linear transformations of the auditory and visual semantic feature encodings output by the encoder of the AVKS network. K v and V v are linear transformations of the visual spatial feature f v . K a and V a are linear transformations of the auditory spatial feature f a . The cross-modal fusion process in the second stage is expressed as:
[0039] f sa = f sa '+ MHA(Q sv ', K sa ', V sa ')
[0040] f av = CA(f sa , f sv )
[0041] Among them, K sa ′ and V sa ′ are linear transformations of f sa ′, and Q sv ′ is another linear transformation of the visual semantic feature encoding.
[0042] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0043] The visual semantic feature encoding of the present invention reflects the speaking state of the target, and the attention map in the auditory semantic feature encoder reflects the positions where the keywords appear in the input sequence. The present invention adopts a cross-attention mechanism to explore the correlation and complementarity between different-level features and different-modal features, and promotes the information interaction between different information sources. The semantic-space feature fusion mechanism can adaptively focus on valuable information, learn multi-level and cross-modal spatio-temporal consistent feature representations, further improve the expression ability of audio-visual features, and thus achieve more accurate tracking. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can also be obtained based on the provided drawings.
[0045] The structures, ratios, sizes, etc. shown in this specification are only used to cooperate with the content disclosed in the specification for those who are familiar with this technology to understand and read, and are not used to limit the limiting conditions for the implementation of the present invention. Therefore, they do not have a substantial technical meaning. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.
[0046] Figure 1 It is a schematic diagram of the network model of the present invention.
[0047] Figure 2 It is a schematic diagram of the AVKS network structure of the present invention.
[0048] Figure 3 It is a schematic diagram of four fusion methods of the experimental design of the present invention.
[0049] Figure 4Schematic diagram for visualizing the semantic-spatial features of the visual and auditory branches of the present invention. Detailed implementation manners
[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. These descriptions are only for further explaining the features and advantages of the present invention, rather than limiting the claims of the present invention; based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0051] The following will further describe in detail the specific implementation manners of the present invention with reference to the accompanying drawings and embodiments. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0052] The present invention provides an Audio-Visual Tracker (AVSS-Tracker) based on semantic-spatial feature fusion. Through an audio-visual semantic feature encoding method based on an audio-visual keyword recognition network that can extract semantic information related to the target from audio-visual signals, and a two-stage semantic-spatial feature fusion mechanism, the cross-attention mechanism is used to mine the correlation and complementarity of features at different levels and modalities, significantly improving the feature expression ability and tracking accuracy. In the actual application process, as Figure 1 shown, it includes the following steps:
[0053] S1. Obtain the original audio-visual signal, and use an audio-visual fusion keyword recognition network (Audio-Visual Keyword Spotting, AVKS) to extract the semantic feature encoding in the audio-visual signal. The original audio-visual signal is converted into a high-dimensional feature representation F a and F v . The audio-visual features are further encoded into higher-level semantic representations, and a Transformer-based encoder is used to obtain the semantic features related to the keywords. The processing process of the semantic feature encoder is as follows:
[0054] Encoder(X) = X att + FFN(X att )
[0055] X att = X + MHA(Q, K, V)
[0056] FFN(X att ) = max(0, X att W1 + b1)W2 + b2
[0057] Among them, X is the input of the encoder, that is, [CLS a ; F a or [CLS v ; F v plus the position encoding. FFN(·) is a fully connected feed-forward network used to enhance the model fitting ability, including two linear transformation layers, with a ReLU activation function in between. W and b represent the weight matrix and the basis vector respectively, and the subscripts represent different layers.
[0058] CLS a and CLS v are the CLS tokens for classification.
[0059] S2. Adopt a two-stream network structure including a visual branch and an auditory branch to process image frames and acoustic signals respectively.
[0060] As Figure 2 , Figure 4 shown, the input of the visual spatial feature extraction network is a pair of images, including the reference template I tpl and the search area image I s . Both are respectively input into a fully convolutional network with shared weights to extract features. Then perform a cross-correlation operation based on the convolutional operator to obtain the response map. The operation process of the visual network is defined as follows:
[0061] f v =Conv v (I tpl )*Conv v (I s )
[0062] The response value represents the probability of the template at each position in the search area, where Conv v are two fully convolutional networks with the same structure and shared weights, and * represents the cross-correlation operation.
[0063] Auditory spatial feature extraction is based on the stGCF acoustic map extraction algorithm, and uses the camera model to project acoustic cues onto the image plane. Then use a fully convolutional network structure similar to the visual network to embed the sound signal into a consistent localization space containing position context. The operation process of the auditory network is defined as follows:
[0064] f a =Conv a (R Ω )
[0065] Among them, R Ω is the manually extracted stGCF acoustic cue.
[0066] S3. Implement information interaction and fusion between two different sequences using a Cross Attention (CA) module. Use the features of one modality as Q, and the features of the other modality as K and V for attention calculation. And use multi-head attention (MHA) in residual form to integrate information from different sequences. The definition of the CA mechanism is as follows:
[0067]
[0068] Among them, Q A is a linear transformation of A + E A , and Q B is a linear transformation of B + E B . E A and E B are position encodings used to supplement spatial position information. The definitions of K and V are similar. LN(·) represents layer normalization.
[0069] After that, two-step fusion is used to enhance the feature distinctiveness of different tracking targets in a multi-person scenario, including two stages: intra-modal and inter-modal. First, perform semantic-spatial feature fusion within each modality, and then perform audio-visual fusion between different modalities. The fusion in the first stage can be formally expressed as:
[0070] f sv = f v + MHA(Q sv , K v , V v )
[0071] f sa ′ = f a + MHA(Q sa , K a , V a )
[0072] Among them, Q sv and Q sa are respectively linear transformations of the auditory and visual semantic feature encodings output by the encoder of the AVKS network. K v and V v are linear transformations of the visual spatial feature f v . K a and V a are linear transformations of the auditory spatial feature f a . The cross-modal fusion process in the second stage is expressed as:
[0073] f sa = f sa ′ + MHA(Q sv ′, K sa ′, Vsa ′)
[0074] f av = CA(f sa , f sv )
[0075] where K sa ′ and V sa ′ are linear transformations of f sa ′, and Q sv ′ is another linear transformation of the visual semantic feature encoding.
[0076] Dataset
[0077] AV16.3 Dataset: This dataset is a real indoor multi-speaker audio-visual corpus, specifically designed to test the localization and tracking algorithms for pure audio, pure video, and audio-visual fusion. The dataset contains 43 audio-visual sequences with durations ranging from 14 seconds to 9 minutes, including 1 to 3 moving speakers, and is recorded using a distributed sensor platform. The audio data is recorded by two 8-microphone uniform circular arrays with a radius of 10 cm placed on the conference table at a sampling rate of 16 kHz. The video is captured by 3 digital cameras at a sampling rate of 25 Hz, with each frame being a color image of 360×288 pixels. The experiments in this chapter use the data from one camera and one microphone, and are tested on 9 sequences in three single-speaker scenarios (seq08, 11, 12) of AV16.3-SOT.
[0078] CAV3D Dataset: This dataset is a 3D speaker tracking audio-visual corpus collected by a co-located sensor platform. The sensing platform consists of a camera and an 8-element circular microphone array with a diameter of 20 cm, which records 8-channel audio signals at 96 kHz (24-bit) and video signals at 15 frames per second (fps) with a resolution of 768×1024 pixels. In addition, 4 corner cameras are configured for 3D trajectory calibration. The experiments use two subsets, CAV3D-SOT2 (6 scenarios, one speaker and one non-speaking participant) and CAV3D-MOT (5 scenarios, multiple simultaneous speakers).
[0079] Experimental Setup
[0080] The proposed AVKS and AVSS-Tracker use a phased training strategy. First, the classifiers of the visual and auditory branches of AVKS are pre-trained, and then the single-modal models are combined for training the audiovisual fusion model. Cross-entropy is used as the loss function for training, and it is executed for 35 epochs with an initial learning rate of 0.01. During testing, 5-second audiovisual signals are intercepted as the input to the semantic feature encoding network. The image samples for training AVSS-Tracker are image patches with a size of 120×120 cropped from the AV16.3 single-person sequences seq01, 02, and 03. The image samples are randomly cropped around the ground-truth bounding boxes and completely contain the ground-truth boxes. The audio samples are stGCF spectrograms extracted from the corresponding sampling regions. A total of 44,985 groups of audiovisual sample pairs are collected. The parameters of the semantic feature encoding part are fixed to train AVSS-Tracker, and the SGD optimizer is used to train for 50 epochs with a batch size of 16 and a fixed learning rate of 0.0001. All experiments are based on the PyTorch environment and are conducted on an NVIDIA RTX 4090Ti graphics card.
[0081] This task involves multi-object tracking. Therefore, in addition to the two-dimensional MAE (unit: pixel) and localization accuracy (ACC, unit: %), the MOT (Multiple Object Tracking) metric is adopted as the evaluation index, including MOTA (MOT Accuracy, unit: %) and MOTP (MOT Precision, unit: pixel).
[0082] Evaluation of the effectiveness of semantic-spatial feature fusion
[0083] To evaluate the effectiveness of the semantic-spatial feature fusion module, comparative experiments on different fusion strategies are conducted on two sequences, CAV3D-SOT2 and AV16.3-SOT. Four fusion methods are designed for the experiment, as Figure 3 shown. Method 1 does not add semantic features and only uses the cross-attention mechanism to fuse audiovisual spatial features. Method 2 directly adds the audiovisual features element-wise after performing intra-modal semantic-spatial feature fusion, lacking cross-modal information interaction. Method 3 adds a cross-modal audiovisual cross-attention module on the basis of Method 2. Method 4 adopts a two-stage fusion strategy, adding the fusion of visual semantic features and auditory spatial features on the basis of Method 3. New networks are trained according to different fusion methods, and the tracking performance is shown in Table 2. It can be seen that the methods using both spatial-feature fusion and cross-modal fusion are better than the two one-stage fusion methods, demonstrating the effectiveness of semantic features and the effectiveness of the cross-modal attention mechanism.
[0084] The above only elaborates in detail on the preferred embodiments of the present invention. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention, and all such changes should be included within the protection scope of the present invention.
Claims
1. An audio-visual speaker tracking method based on semantic-spatial feature fusion, characterized in that: The following steps are involved: S1, obtain the original audio and video signal, use the audio-visual fusion keyword recognition network to extract the semantic feature coding in the audio-visual signal; The method for extracting semantic feature coding from the audio-visual signal in S1 is: S11. Convert the original audio and video signals into high-dimensional feature representation F a With F v ; S12, use Transformer-based encoder to obtain semantic features related to keywords; S2, using a two-stream network structure including a visual branch and an auditory branch to process image frames and acoustic signals respectively; S3, using the cross attention module to achieve information interaction and fusion between two different sequences of semantic-spatial features; The method of using the cross attention module in S3 to realize the information interaction and fusion between two different sequences of semantic-spatial features is: S31. Take the features of one modality as Q and the features of the other modality as K and V for attention calculation; S32 and use residual multi-head attention to integrate information from different sequences; The CA mechanism is defined as follows: Among them, Q A It's A+E A The linear transformation of Q B It's B+E B The linear transformation of E A and E B It is used to supplement the position coding of spatial position information. The definitions of K and V are similar. LN(·) represents layer normalization. Then, a two-step fusion is used to enhance the feature distinguishability of different tracking targets in multi-person scenes, including intra-modality and inter-modality stages. First, semantic-spatial feature fusion is performed within each modality, and then audio-visual fusion is performed between different modalities. The fusion of the first stage can be formally expressed as: f sv =f v +MHA(Q sv ,K v ,V v ) f sa ′=f a +MHA(Q sa ,K a ,V a ) Among them, Q sv and Q sa are the linear transformations of the auditory and visual semantic feature encodings output by the encoder of the audio-visual fusion keyword spotting network, K v and V v is the visual spatial feature f v Linear transformation, K a and V a is the auditory spatial feature f a The linear transformation of ; the cross-modal fusion process in the second stage is expressed as: f sa =f sa ′+MHA(Q sv ′,K sa ′,V sa ′) f av =CA(f sa ,f sv ) Among them, K sa ′ and V sa′ Yes sa ′, Q sv ′ is another linear transformation of visual semantic feature encoding.
2. The audio-visual speaker tracking method based on semantic-spatial feature fusion according to claim 1, characterized in that: The processing process of the semantic feature encoder in S12 is expressed as: Encoder(X)=X att +FFN(X att ) X att =X+MHA(Q,K,V) FFN(X att )=max(0,X att W1+b1)W2+b2 Where X is the input of the encoder, that is, [CLS a ; F a ] or [CLS v ; F v ] plus position encoding; FFN(·) is a fully connected feedforward network used to enhance the model fitting ability, including two linear transformation layers with a ReLU activation function in between; W and b represent the weight matrix and basis vector respectively, and the subscripts represent different layers; CLS a and CLS v is the CLS tag for classification.
3. The audio-visual speaker tracking method based on semantic-spatial feature fusion according to claim 1, characterized in that: The method of using a dual-stream network structure including a visual branch and an auditory branch in S2 to process image frames and acoustic signals respectively is: S21, input a pair of images into the visual space feature extraction network, including the reference template I tpl and search area image I s ; Both are input into a fully convolutional network with shared weights to extract features; S22, performing a cross-correlation operation based on a convolution operator to obtain a response map; the operation process of the visual network is defined as follows: f v =Conv v (I tpl )*Conv v (I s ) Among them, Conv v These are two fully convolutional networks with the same structure and shared weights. * indicates the cross-correlation operation. S23, projecting acoustic cues onto the image plane using a camera model; S24, then use a fully convolutional network structure similar to the visual network to embed the sound signal into a consistent positioning space containing the position context; the operation process of the auditory network is defined as: f a =Conv a (R Ω ) Among them, R Ω are manually extracted acoustic cues of stGCF.