Visual-Audio Multimodal Object Tracking Method and Device Based on Cross-Modal Transformer

By using a cross-modal Transformer to fusion visual and audio information in target tracking, the lack of visual-audio fusion method in the prior art is solved, and higher robustness and tracking accuracy are achieved.

CN118378123BActive Publication Date: 2025-05-30WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410390051.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-02
Publication Date
2025-05-30
Estimated Expiration
2044-04-02

AI Technical Summary

Technical Problem

The lack of methods of fusion of vision and audio in the prior art leads to inaccurate target tracking results.

Method used

The visual-audio multimodal target tracking method based on cross-modal Transformer is adopted. By extracting images and audio tokens, cross-modal alignment and in-modal alignment, combined with multi-layer Transformer encoder for feature extraction and fusion, and finally predicting target coordinates using classification and bounding box regression.

Benefits of technology

The robustness of the system and the perceived ability of the target are improved, and the spatial and temporal consistency of the target is better captured by integrating visual and audio information, improving the accuracy of tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118378123B_ABST
    Figure CN118378123B_ABST
Patent Text Reader

Abstract

The present invention discloses a visual-audio multimodal object tracking method and device based on a cross-modal Transformer. First, information tokens from two modalities, namely vision and audio, are obtained, and two multimodal feature alignment methods are introduced to optimize the embeddings of the two modalities before feature extraction and fusion between multimodals. Before inputting into the encoder, the aligned audio features are injected into the embeddings of the search region image and the template region image, thereby improving the learning ability of the encoder. Subsequently, through the processing of the encoder layer, the features between multimodals are fully fused and learned. Finally, by using the methods of classification and bounding box regression, the output of the last layer of the encoder is utilized to accurately predict the coordinates of the target. The multimodal fusion method of the present invention has higher robustness compared to a single modality, and can improve the system's perception ability of the target. The fusion of vision and audio can better capture the spatio-temporal consistency of the target, thus improving the accuracy of tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a visual-audio multi-modal object tracking method and device based on cross-modal Transformer. Background Art

[0002] Single-object tracking is one of the research focuses in the field of computer vision. Its core task is to obtain the position information of a single object in each frame of a video, providing a basis for in-depth analysis and understanding of the motion behavior and laws of the object, and laying a foundation for further research. In recent years, driven by artificial intelligence, significant progress has been made in the field of object tracking, making it one of the popular research directions in computer vision. With the rise of deep learning, the Transformer network has emerged in the field of object tracking. Especially in recent years, introducing visual Transformer into the field of object tracking has achieved remarkable results.

[0003] Transformer object tracking methods.

[0004] Currently, the object tracking methods using Transformer include two types: based on CNN-Transformer and completely based on Transformer. Among them, the CNN-Transformer-based tracker shows more excellent characteristics in tracking performance than the CNN-based tracker. According to the empirical evaluation of challenging benchmark datasets, the completely Transformer-based tracker is significantly better than other methods while maintaining a moderate efficiency score. However, compared with the CNN-based tracker, most Transformer trackers have a decrease in tracking speed due to the large model training network.

[0005] The completely Transformer-based trackers include a two-stream tracking framework and a single-stream tracking framework. The two-stream two-stage complete Transformer tracker uses a backbone Transformer model to extract the features of the target template and the search area, and uses another Transformer for feature fusion and enhancement. Finally, the target is located through a small prediction network. These two-stream two-stage trackers demonstrate a simple and neat tracking architecture and are superior to the CNN-based and CNN-Transformer-based trackers in performance, fully leveraging the attention mechanism of Transformer to achieve more accurate object tracking.

[0006] On the other hand, single-stream single-stage trackers adopt a fully Transformer architecture to combine the feature extraction and feature fusion processes. These trackers segment the target template and the search region image into tokens, concatenate their positional embeddings, and then input them into the Transformer for processing. Since these trackers use a single Transformer network to extract features, they can effectively integrate the features of the template and the search region, thus enabling effective identification and differentiation of the background and the target. Based on these advantages, the fully Transformer-based single-stream single-stage trackers demonstrate excellent performance on various benchmark datasets.

[0007] Multi-modal tracking methods.

[0008] With the continuous development of technology and the increasing demand, researchers have gradually realized that it is difficult to meet the comprehensive requirements for target tracking using only image data in some cases. Therefore, target tracking has gradually expanded from data limited to the image modality to scenarios that integrate multiple modalities including images, text, audio, etc., in order to capture the features of the target more comprehensively.

[0009] Through in-depth research, the applicant found that introducing visual-text multi-modal fusion methods into the field of target tracking can, to a certain extent, improve the robustness and performance of the system. However, compared with the progress in this field, there is still a significant room for improvement in applying the fusion of the two modalities of vision and audio to target tracking. The methods in the prior art do not currently have a method for fusing vision and audio, resulting in inaccurate target tracking results. Summary of the Invention

[0010] To address the above problems, the present invention innovatively proposes a vision-audio target tracking method based on Transformer and multi-modal fusion to fill the gap in the field of target tracking where there is a lack of methods for fusing vision and audio using cross-modal Transformers, providing more comprehensive multi-modal processing technology support for target tracking.

[0011] To achieve the above object, in the first aspect of the present invention, a vision-audio multi-modal target tracking method based on cross-modal Transformer is provided, including:

[0012] S1: Input the search region image, the template region image, the audio segment corresponding to the image frame of the search region image, and the audio segment corresponding to the image frame of the template region image;

[0013] S2: Extract image tokens from the input search region image and template region image;

[0014] S3: Extract audio tokens from the audio segments of the image frames corresponding to the input search area image and the audio segments of the image frames corresponding to the template area image;

[0015] S4: Perform cross-modal alignment on the extracted image tokens and audio tokens to obtain the cross-modally aligned image tokens and audio tokens. Among them, the cross-modally aligned image tokens include the cross-modally aligned search area image tokens and the cross-modally aligned template area image tokens; Perform intra-modal alignment on the cross-modally aligned search area image tokens and the cross-modally aligned template area image tokens to obtain the cross-modally aligned image tokens;

[0016] S5: Perform a modality mixing operation on the cross-modally aligned audio tokens and the cross-modally aligned image tokens, and use a multi-layer Transformer encoder for feature extraction and fusion to obtain the fused multi-modal features. The fused multi-modal features include the search area image features and the template area image features;

[0017] S6: Obtain the coordinates of the target in the search area image according to the search area image features output by the last layer of the encoder.

[0018] In one implementation, step S2 includes:

[0019] S2.1: Respectively divide the input search area image and template area image into non-overlapping image block groups;

[0020] S2.2: Apply linear projection to the divided image block groups to generate the corresponding search area image tokens and template area image tokens;

[0021] S2.3: Add learnable position embeddings to maintain the position information of the image blocks.

[0022] In one implementation, step S3 includes:

[0023] S3.1: Perform downsampling processing on the audio segments of the image frames corresponding to the input search area image and the audio segments of the image frames corresponding to the template area image;

[0024] S3.2: Divide the downsampled audio segments into short-time frames and add Hamming windows;

[0025] S3.3: Apply the short-time Fourier transform to each short-time frame to obtain the amplitude spectrum and energy spectrum, and apply the Mel filter bank to obtain the Mel filter bank features;

[0026] S3.4: Extract the first L Mel cepstral coefficients, then obtain D-dimensional audio tokens through a linear projection layer, and then add learnable position embeddings;

[0027] S3.5: Add the CLS token of the class token at the start position of the audio tokens sequence to represent the class information of the entire audio sequence.

[0028] In one implementation, step S4 performs cross-modal alignment on the extracted image tokens and audio tokens to obtain the cross-modally aligned image tokens and audio tokens, including:

[0029] Calculate B 2 cosine similarities between the tokens of image and audio pairs, where the image and audio pairs from the same time are called positive pairs, and the image and audio pairs from different times are called negative pairs;

[0030] By maximizing the cosine similarities of B positive pairs and minimizing the cosine similarities of B 2 -B negative pairs to train the audio encoder, and the symmetric cross-entropy loss is used for calculation during the training process;

[0031] Obtain the tokens of the cross-modally aligned image and audio through the trained audio encoder.

[0032] In one implementation, step S4 performs intra-modal alignment on the cross-modally aligned search area image tokens and cross-modally aligned template area image tokens, including:

[0033] Fuse the tokens of the search area and template area in the visual modality after cross-modal alignment, where the search area image and template area image from the same time are regarded as positive pairs, and the images from different times are regarded as negative pairs;

[0034] Calculate through the contrastive loss to obtain the tokens of the cross-modally aligned image.

[0035] In one implementation, step S5 includes:

[0036] S5.1: Inject the cross-modally aligned audio tokens into the two image tokens after cross-modal alignment in the visual modality respectively to obtain the search area visual embedding and template area visual embedding injected with audio information;

[0037] S5.2: Concatenate the visual embeddings of the search area and the template area embedded with audio information, and then send them to the L-layer Transformer encoder for joint processing. Each layer of the encoder includes a multi-head self-attention module and a feed-forward network. The multi-head self-attention module adopts the multi-head self-attention mechanism, which allows the model to interact with the input sequence from multiple attention perspectives and capture different feature relationships in the input sequence. The feed-forward network introduces non-linear transformation of features and obtains the fused multi-modal features based on the output of the multi-head attention module.

[0038] In one implementation, step S6 includes:

[0039] S6.1: Use the classification branch to reshape the image features of the search area output by the last layer of the Transformer encoder into a 2D feature map with the original spatial resolution, and then pass it through a 1-layer fully convolutional network to obtain the target classification score map, local offset, and normalized bounding box size. The point with the highest score in the target classification score map is the target center point, and the local offset is used to compensate for the discretization error caused by the reduced resolution.

[0040] S6.2: Use the bounding box regression branch to predict the center coordinate offset and size of the object. In the prediction process, the loss is calculated by weighting the l1 loss and the GIoU loss.

[0041] S6.3: Convert the visual-audio multi-modal object tracking model based on cross-modal Transformer into a multi-task optimization problem, and optimize the cross-modal alignment loss, intra-modal loss, classification loss, and regression loss simultaneously.

[0042] S6.4: Complete the parameter update by backpropagation through calculating the loss, further correct the position of the predicted bounding box, and finally obtain the coordinates of the target center point and the size of the bounding box.

[0043] Based on the same inventive concept, the second aspect of the present invention provides a visual-audio multi-modal object tracking device based on cross-modal Transformer, including:

[0044] An input module, configured to input the search area image, the template area image, the audio segment corresponding to the image frame of the search area image, and the audio segment corresponding to the image frame of the template area image.

[0045] An image tokens extraction module, configured to extract image tokens from the input search area image and template area image.

[0046] An audio tokens extraction module, configured to extract audio tokens from the audio segments corresponding to the image frames of the input search area image and the template area image.

[0047] A cross-modal alignment module for cross-modally aligning the extracted image tokens and audio tokens to obtain cross-modally aligned image tokens and audio tokens. Among them, the cross-modally aligned image tokens include cross-modally aligned search area image tokens and cross-modally aligned template area image tokens; perform intra-modal alignment on the cross-modally aligned search area image tokens and the intra-modal aligned template area image tokens to obtain cross-modally aligned image tokens;

[0048] A feature extraction and fusion module for performing a modal mixing operation on the cross-modally aligned audio tokens and the cross-modally aligned image tokens, and using a multi-layer Transformer encoder for feature extraction and fusion to obtain fused multi-modal features. The fused multi-modal features include search area image features and template area image features;

[0049] A bounding box prediction head module for obtaining the coordinates of the target in the search area image based on the search area image features output by the last layer of the encoder.

[0050] Based on the same inventive concept, the third aspect of the present invention provides a computer-readable storage medium with a computer program stored thereon. When the program is executed by a processor, it implements the method described in the first aspect.

[0051] Based on the same inventive concept, the fourth aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method described in the first aspect.

[0052] Compared with the prior art, the advantages and beneficial technical effects of the present invention are as follows:

[0053] The visual-audio multimodal object tracking method based on cross-modal Transformer disclosed in the present invention first obtains information tokens (image tokens and audio tokens) from two modalities, visual and audio, and introduces two multimodal feature alignment methods, namely cross-modal alignment and intra-module alignment, to optimize the embeddings of the two modalities before feature extraction and fusion between multimodals. Before inputting into the encoder, the aligned audio features are injected into the embeddings of the search-region image and the template-region image, thereby improving the learning ability of the encoder. Subsequently, through the processing of the encoder layer, the features between multimodals are fully fused and learned. Finally, by using the methods of classification and bounding box regression, the coordinates of the target are accurately predicted using the output of the last encoder layer. The multimodal fusion method is adopted, which has higher robustness compared with a single modality, and the perception ability of the system for the target can be improved by fusing the information of these two modalities. This is because visual and audio information provide the system with features of different aspects of the target. Finally, since visual and audio information can complement each other in space and time, the fusion of visual and audio can better capture the spatio-temporal consistency of the target, thus improving the accuracy of tracking. Description of the Drawings

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0055] Figure 1 It is the structural diagram of the model adopted by the method of the present invention;

[0056] Figure 2 It is the flowchart for extracting audio tokens in the embodiment of the present invention;

[0057] Figure 3 It is the schematic diagram of modality alignment in the embodiment of the present invention;

[0058] Figure 4 It is the structural diagram of the Transformer encoder adopted in the embodiment of the present invention. Detailed Embodiments

[0059] Through in-depth research on the prior art, the applicant has found that introducing a visual-text multimodal fusion method into the field of object tracking significantly improves the robustness and performance of the system. However, compared with the progress in this field, there is still considerable room for improvement in applying the fusion of the visual-audio two modalities to object tracking. Visual and audio information provide features of different aspects of the object, where vision mainly captures the appearance, shape, and motion trajectory, while audio provides information about the environment around the object and the sound of the object itself. Fusing the information of these two modalities can significantly improve the system's perception ability of the object, enabling it to more comprehensively understand the state of the object. In addition, visual and audio information complement each other in space and time. For example, when observing the change in the position of an object in a video, audio information may provide additional clues about the movement of the object. Through this information fusion, the spatio-temporal consistency of the object can be better captured, thereby improving the accuracy of tracking. The above prior knowledge proves that applying the method of fusing the visual and audio two modalities in object tracking can effectively improve the robustness and performance of the system.

[0060] The main idea of the present invention is as follows:

[0061] First, obtain information tokens from both the visual and audio modalities, and introduce two multimodal feature alignment methods, namely cross-modal alignment and intra-module alignment, to optimize the embeddings of the two modalities before feature extraction and fusion between multimodals. Before inputting into the encoder, inject the aligned audio features into the embeddings of the search area image and the template area image, thereby improving the learning ability of the encoder. Subsequently, through the processing of the encoder layer, the features between multimodals are fully fused and learned. Finally, adopt the methods of classification and bounding box regression, and use the output of the last layer of the encoder to accurately predict the coordinates of the object. Advantages of the present invention: First, this multimodal fusion method has higher robustness compared to a single modality. Second, fusing the information of these two modalities can improve the system's perception ability of the object because visual and audio information provide features of different aspects of the object for the system. Finally, since visual and audio information can complement each other in space and time, the fusion of visual and audio can better capture the spatio-temporal consistency of the object and improve the accuracy of tracking.

[0062] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0063] Embodiment 1

[0064] The embodiment of the present invention provides a visual-audio multi-modal object tracking method based on a cross-modal Transformer, including:

[0065] S1: Input the search area image, the template area image, the audio segment of the image frame corresponding to the search area image, and the audio segment of the image frame corresponding to the template area image;

[0066] S2: Extract image tokens from the input search area image and template area image;

[0067] S3: Extract audio tokens from the audio segment of the image frame corresponding to the input search area image and the audio segment of the image frame corresponding to the template area image;

[0068] S4: Perform cross-modal alignment on the extracted image tokens and audio tokens to obtain the cross-modal aligned image tokens and audio tokens. Among them, the cross-modal aligned image tokens include the cross-modal aligned search area image tokens and the cross-modal aligned template area image tokens; Perform intra-modal alignment on the cross-modal aligned search area image tokens and the intra-modal aligned template area image tokens to obtain the cross-modal aligned image tokens;

[0069] S5: Perform a modality mixing operation on the cross-modal aligned audio tokens and the cross-modal aligned image tokens, and use a multi-layer Transformer encoder for feature extraction and fusion to obtain the fused multi-modal features. The fused multi-modal features include the search area image features and the template area image features;

[0070] S6: Obtain the coordinates of the target in the search area image according to the search area image features output by the last layer of the encoder.

[0071] Figure 1 It is the model structure diagram adopted by the method of the present invention; The input of the model constructed by the present invention is the search area image, the template area image and the audio segment of the corresponding image frame. The output is the coordinate representation of the target in the search area image.

[0072] Specifically, S4 is to perform cross-modal alignment, including cross-modal alignment and intra-modal alignment. In order to address the problem that directly inputting image and audio features into a unified encoder layer may lead to poor self-attention learning performance, the present invention proposes two methods for multi-modal feature alignment, namely cross-modal alignment and intra-module alignment. This aims to mitigate the negative impact brought by the modality data difference in multi-modal learning.

[0073] Before inputting the tokens of images and audio into the encoder layer, cross-modal alignment is first performed to ensure that visual and audio features can be effectively embedded. Although cross-modal alignment is effective, it ignores important intra-modal supervision signals. To better learn time-invariant features, the present invention also performs intra-modal alignment to fully utilize the intra-modal temporal supervision information. In this part of intra-modal alignment, it focuses on fusing the tokens in the search region and the template region of the visually modality after cross-modal alignment.

[0074] In step S5, the features of the search region image and the template region image are extracted and fully fused through a multi-layer encoder, and the search region image and template region image features with strong ability to distinguish the target and the background can be obtained. The present invention extracts the features of the search region image output by the last layer of the encoder and inputs them into the bounding box prediction head module, and the coordinates of the target in the search region image can be learned. The bounding box prediction head is decomposed into two branches: classification and bounding box regression.

[0075] In one implementation, step S2 includes:

[0076] S2.1: Respectively segment the input search region image and template region image into non-overlapping groups of image patches;

[0077] S2.2: Apply linear projection to the segmented groups of image patches to generate corresponding search region image tokens and template region image tokens;

[0078] S2.3: Add learnable position embeddings to maintain the position information of the image patches.

[0079] In the specific implementation process, when the input pair of image groups is the search region image and the template region image First, they are segmented into non-overlapping groups of image patches with a resolution size of P×P, and the number of segmented groups of image patches is N x and N z .

[0080]

[0081]

[0082] x refers to the tensor of the search region image, R represents the real number field, H x refers to the height of the search region image, set to 320, W x refers to the width of the search region image, set to 320. z refers to the tensor of the template region image, H z refers to the height of the template region image, set to 128, W z refers to the width of the template region image, set to 128. P refers to the width and height of an image patch, set to 16. Nx Denotes the number of groups of image patches obtained by dividing the search area image, N z Denotes the number of groups of image patches obtained by dividing the template area image.

[0083] Then apply linear projection to these groups of image patches to generate tokens of the image modality data, called and Then add learnable position embeddings to maintain the position information of the image patches. Randomly initialize two matrices as the position encodings of the search area image and the template area image, which are updated during the training process, and and then add them to the tokens to obtain and

[0084]

[0085]

[0086] D denotes the dimension of the tokens tensor, set to 512, H x Denotes the search area image tokens obtained by linear projection, H z Denotes the template area image tokens obtained by linear projection, Posi x and Posi z Denote the position encodings of the search area image and the template area image respectively, Denotes the search area image tokens with position encoding added, Denotes the template area image tokens with position encoding added.

[0087] In one implementation, step S3 includes:

[0088] S3.1: Downsample the audio segments corresponding to the image frames of the input search area image and the audio segments corresponding to the image frames of the template area image;

[0089] S3.2: Divide the downsampled audio segments into short-time frames and add Hamming windows;

[0090] S3.3: Apply the short-time Fourier transform to each short-time frame to obtain the amplitude spectrum and the energy spectrum, and apply the Mel filter bank to obtain the Mel filter bank features;

[0091] S3.4: Extract the first L Mel cepstral coefficients, then obtain D-dimensional audio tokens through a linear projection layer, and then add learnable position embeddings;

[0092] S3.5: Add the CLS token of the class token at the beginning of the audio tokens sequence to represent the class information of the entire audio sequence.

[0093] Specifically, the input is a one-dimensional long audio waveform. In this embodiment, the audio waveform is first downsampled, and then the audio segments are divided into short-time frames with a time unit of 8 milliseconds. Next, a Hamming window of 32 milliseconds is added to each frame signal. Then, the short-time Fourier transform is applied to each frame to obtain the amplitude spectrum and the energy spectrum. Then, the Mel filter bank is applied to each frame to obtain the Mel filter bank features. Finally, the first L Mel cepstral coefficients are extracted, and then D-dimensional audio tokens are obtained through a linear projection layer, and learnable position embeddings are added. Finally, the CLS token of the class token is added at the beginning of the tokens sequence to represent the class information of the entire audio sequence.

[0094] Please refer to Figure 2 , which is the flowchart for extracting audio tokens in the embodiment of the present invention.

[0095] 1. Short-time Fourier transform

[0096] Apply the short-time Fourier transform to each frame to obtain the amplitude spectrum and the energy spectrum. The specific calculation formula of the Fourier transform is as follows:

[0097] Euler's formula:

[0098] e jωn = cos(ωn) + j sin(ωn)

[0099] The calculation method of the k-th point of the discrete Fourier transform (DFT):

[0100]

[0101] ω is the angular frequency; x[n] is the n-th sampling point of the time-domain waveform; X[k] is the k-th point of the Fourier spectrum;

[0102] N is the number of samples, and K is the size of the DFT, where K ≥ N.

[0103] Using Euler's formula:

[0104]

[0105] X[k] = X real [k] - jX imag [k]

[0106]

[0107]

[0108] X real [k] is the real part of X[k], and X imag [k] is the imaginary part of X[k].

[0109] Amplitude spectrum:

[0110]

[0111] X magnitude [k] is the amplitude spectrum of X[k].

[0112] Energy spectrum:

[0113] X power [k] = X real [k] 2 + X imag [k] 2

[0114] X power [k] is the energy spectrum of X[k].

[0115] 2 Obtain Mel filter bank features

[0116] Apply the Mel filter bank to each frame to simulate the perceptual characteristics of the human ear for audio. In this embodiment, a Mel filter bank with 26 triangular filters is defined.

[0117] Mel frequency calculation:

[0118]

[0119] f is the frequency, and Mel(f) represents the Mel frequency.

[0120] Mel band energy calculation:

[0121]

[0122] f(·) represents the Mel frequency after Mel scale transformation. The center frequency of each filter is f(m), and the starting and ending frequencies are f(m - 1) and f(m + 1) respectively. H m (k) represents the Mel band energy of the m-th filter bank, and k represents the frequency index of the discrete Fourier transform, i.e., the k-th point.

[0123] Mel filter bank feature calculation:

[0124] FilterBanks(k) = X power [k]H m (k)

[0125] FilterBanks(k) is the Mel filter bank feature.

[0126] 3. Obtain Mel-Frequency Cepstral Coefficients

[0127] First, sum the features of multiple Mel filter banks and then perform a logarithmic transformation because the way the human ear perceives audio conforms more to a logarithmic scale, and the logarithmic transformation helps improve the robustness of the features.

[0128]

[0129] N is the number of sampling points, m represents the m-th filter bank, and M represents the number of filter banks.

[0130] Then, perform operations such as discrete cosine transform to obtain the Mel-Frequency Cepstral Coefficients. Finally, select the first L coefficients as the final Mel-Frequency Cepstral Coefficients.

[0131]

[0132] C(n) is the Mel-Frequency Cepstral Coefficient, and the L-th order represents the order of the Mel-Frequency Cepstral Coefficient, taking 16.

[0133] Then, obtain D-dimensional audio tokens through a linear projection layer. Then, add learnable position embeddings. Randomly initialize a matrix as the position encoding of the audio, which is updated during the training process.

[0134] H a =Linear(C(n))

[0135]

[0136]

[0137] H a represents the audio tokens, N a is the length of the audio tokens, t is the time length of the audio segment corresponding to each frame of the image, D is the dimension size of the audio tokens, which is the same as the dimension of the image tokens. Linear(·) represents the linear projection layer. Posi a is the position encoding of the audio, is the audio tokens after adding the position encoding.

[0138] Finally, add the CLS token of the class label at the beginning of the tokens sequence to represent the class information of the entire audio sequence.

[0139] In one implementation, step S4 performs cross-modal alignment on the extracted image tokens and audio tokens to obtain the cross-modal aligned image tokens and audio tokens, including:

[0140] Calculate B 2 The cosine similarity between the tokens of B image and audio pairs, where the image and audio pairs from the same time are called positive pairs, and the image and audio pairs from different times are called negative pairs;

[0141] By maximizing the cosine similarity of B positive pairs and minimizing the cosine similarity of B 2 -B negative pairs to train the audio encoder, and the training process is calculated using the symmetric cross-entropy loss;

[0142] Obtain the tokens of the cross-modally aligned images and audio through the trained audio encoder.

[0143] In the specific implementation process, in order to address the problem that directly inputting the features of the image and audio modalities into a unified encoder layer may lead to poor self-attention learning performance, this embodiment proposes two methods for multi-modal feature alignment, namely cross-modal alignment and intra-module alignment. This aims to mitigate the negative impact brought by the modality data differences in multi-modal learning.

[0144] Please refer to Figure 3 , which is the schematic diagram of modality alignment in the embodiment of the present invention. In this figure, the token pairs with the same number are called positive pairs, and different ones are called negative pairs. Alignment is mainly achieved by maximizing the cosine similarity of B positive pairs and minimizing the cosine similarity of B2 - B negative pairs.

[0145] This embodiment maps the tokens of the search area image the tokens of the template area image and the audio segment tokens to 256-dimensional vectors, which are f x , f z and f a .

[0146]

[0147]

[0148]

[0149] f x represents the search area image tokens obtained through linear projection, f z represents the template area image tokens obtained through linear projection, f a represents the audio tokens obtained through linear projection, and Linear(·) represents linear projection.

[0150] Before inputting the tokens of images and audio into the encoder layer, it is crucial to perform the step of cross-modal alignment to ensure that visual and audio features can be effectively embedded. The specific operation includes calculating the cosine similarity between B×B visual and audio pairs. Here, the images and audio from the same time are called positive pairs, while the image and audio pairs from different times are called negative pairs. The goal of the present invention is to make the positive pairs closer while increasing the distance between the negative pairs. Therefore, by maximizing the cosine similarity of B positive pairs and minimizing the cosine similarity of B 2 -B negative pairs to train the audio encoder. This process is calculated using the symmetric cross-entropy loss. By optimizing the loss of cross-modal alignment, the visual and audio embeddings can be well aligned in the feature space, and the tokens of the images and audio after cross-modal alignment can be obtained.

[0151] The loss function is as follows:

[0152]

[0153]

[0154] L x2a represents the loss of cross-modal alignment between the search area image tokens and the audio tokens, and L z2a represents the loss of cross-modal alignment between the template area image tokens and the audio tokens. Both represent the tokens vectors of the search area image, the template area image, and the audio corresponding to the i-th frame. B is the batch size of cross-modal alignment, and in this embodiment, B is set to 32. q represents all indices that are not equal to i within the same batch.

[0155] Finally, the sum of the cross-modal alignment losses is

[0156]

[0157] L cross represents the sum of the cross-modal alignment losses.

[0158] By optimizing the loss of cross-modal alignment, the visual and audio embeddings can be well aligned in the feature space, and the tokens of the images and audio after cross-modal alignment can be obtained.

[0159] In one embodiment, step S4 performs intra-modal alignment on the cross-modal aligned search area image tokens and the cross-modal aligned template area image tokens, including:

[0160] Fuse the tokens in the search region and the template region of the visual modality after cross-modal alignment. Among them, the search region image and the template region image from the same time are regarded as positive pairs, and the images from different times are regarded as negative pairs;

[0161] Calculate through the contrastive loss to obtain the tokens of the image after cross-modal alignment.

[0162] In the specific implementation process, although cross-modal alignment is effective, it ignores important intra-modal supervision signals. To better learn time-invariant features, the present invention proposes an intra-modal alignment module to make full use of the intra-modal time supervision information. In this part, the present invention focuses on fusing the tokens in the search region and the template region of the visual modality after cross-modal alignment. The search region image and the template region image from the same time are regarded as positive pairs, while the images from different times are regarded as negative pairs, and the calculation is performed through the contrastive loss, and the calculation method is the same as the previous cross-modal alignment. Finally, the tokens of the image after cross-modal alignment and intra-modal alignment will be obtained.

[0163]

[0164]

[0165] L x2z represents the loss of intra-modal alignment between the search region image and the template region image, and L z2x represents the loss of intra-modal alignment between the template region image and the search region image.

[0166] Finally, the sum of the intra-modal alignment losses is:

[0167]

[0168] L intra represents the sum of the intra-modal alignment losses.

[0169] Finally, the tokens of the image after cross-modal alignment and intra-modal alignment (i.e., the tokens of the image after cross-modal alignment) will be obtained.

[0170] In one implementation, step S5 includes:

[0171] S5.1: Inject the audio tokens after cross-modal alignment into the two image tokens after cross-modal alignment in the visual modality respectively to obtain the search region visual embedding and the template region visual embedding injected with audio information;

[0172] S5.2: Concatenate the visually embedded search region and the visually embedded template region into which the audio information is injected, and then send them to the L-layer Transformer encoder for joint processing. Each layer of the encoder includes a multi-head self-attention module and a feed-forward network. The multi-head self-attention module uses the multi-head self-attention mechanism to allow the model to interact with the input sequence from multiple attention perspectives and capture different feature relationships in the input sequence. The feed-forward network introduces a non-linear transformation of the features and obtains the fused multi-modal features based on the output of the multi-head attention module.

[0173] Please refer to Figure 4 , which is the structural diagram of the Transformer encoder adopted in the embodiments of the present invention.

[0174] Step S5.1 Modal mixing operation

[0175] Inject the audio tokens after cross-modal alignment into two visual tokens respectively. The specific method is: first perform a linear projection on the audio tokens, and then multiply the projection result with the visual tokens in the way of Hadamard product.

[0176] Inject the audio tokens after cross-modal alignment into two visual tokens respectively. All the tokens mentioned in the present invention hereinafter are the tokens after being aligned in step S3. The embedding process is as follows:

[0177]

[0178]

[0179] represents the search region image tokens input to the 0th layer of the encoder, represents the template region image tokens input to the 0th layer of the encoder. ⊙ represents the Hadamard product of matrices, and Linear(·) represents the linear projection layer.

[0180] Step S5.2, Encoder feature extraction and fusion

[0181] See Figure 4 Transformer encoder structure diagram.

[0182] Step S5.2 is to extract and fuse features through an encoder. The visual embeddings of the search area and the template area injected with audio information are concatenated together and then fed into the L-layer Transformer encoder for joint processing. Each layer of the encoder mainly includes two modules, the multi-head self-attention module and the feed-forward network. First is the multi-head self-attention mechanism, whose function is to allow the model to interact with the input sequence from multiple attention perspectives and capture different feature relationships in the input sequence. The multi-head mechanism enables the model to flexibly learn the complex relationships in the input data. Especially for multi-modal data, it helps to more comprehensively capture the associations between multi-source information such as vision and audio. Then it is fed into the feed-forward network to introduce non-linear transformation of features, enhance the representation ability of the model, and enable it to better adapt to the complex multi-modal data features. Among them, layer normalization (Layer Normalization) is applied to standardize the input before the data is input into the multi-head self-attention module and the feed-forward network module to ensure the stability of each layer and accelerate the convergence of training. The specific operations of the encoder are as follows:

[0183]

[0184]

[0185]

[0186]

[0187]

[0188] Q represents the query vector of the attention mechanism, K represents the key vector of the attention mechanism, and V represents the value vector of the attention mechanism. d represents the dimension size of Q, W is the weight matrix of each attention head in the multi-head self-attention layer, and u is the number of heads in the multi-head self-attention layer. W 1 and W 2 are the weight matrices of the feed-forward network, b 1 and b 2 are the bias vectors of the feed-forward network. and are the inputs of the l-th layer encoder, and are the outputs of the l-th layer encoder. is the output after the self-attention mechanism module, is the output after the multi-head self-attention mechanism module, is the output after the feed-forward network.

[0189] In one implementation, step S6 includes:

[0190] S6.1: Using the classification branch, reshape the search region image features output by the last layer of the Transformer encoder into a 2D feature map with the original spatial resolution, and then pass it through a single-layer fully convolutional network to obtain the target classification score map, local offset, and normalized bounding box size. Among them, the point with the highest score in the target classification score map is the target center point, and the local offset is used to compensate for the discretization error caused by the reduced resolution;

[0191] S6.2: Use the bounding box regression branch to predict the center coordinate offset and size of the object. During the prediction process, the loss is calculated by weighting the l1 loss and the GIoU loss;

[0192] S6.3: Convert the vision-audio multi-modal object tracking model based on cross-modal Transformer into a multi-task optimization problem, and simultaneously optimize the cross-modal alignment loss, intra-modal loss, classification loss, and regression loss;

[0193] S6.4: Complete parameter update by backpropagating through the calculated loss, further correct the position of the predicted bounding box, and finally obtain the coordinates of the target center point and the size of the bounding box.

[0194] Specifically, for step S6.1 classification

[0195] First, reshape the search region image features output by the last layer of the encoder into a 2D feature map with the original spatial resolution, and then pass it through a single-layer fully convolutional network FCN. The FCN network consists of 4 layers of Conv-BN-ReLU layers, and each Conv-BN-ReLU layer has three outputs, namely the score map for target classification local offset and normalized bounding box size The classification loss is used to enhance the model's ability to distinguish objects from the background, and is represented by the Gaussian-weighted focal loss. The classification loss, that is, the Gaussian-weighted focal loss, is:

[0196]

[0197]

[0198] DP is the score map for target classification. DP xy is the score corresponding to the xy coordinates in the predicted target classification score map, is the heat map generated using a Gaussian kernel for the true target center, is the corresponding low-resolution equivalent position of the true target center, is the score corresponding to the xy coordinates of this heat map, σ dp is the standard deviation parameter corresponding to the Gaussian kernel function. L clsis the classification loss, α and β are hyperparameters, and in this embodiment, α = 2 and β = 4 are set.

[0199] Step S6.2 Bounding box regression

[0200] The bounding box regression branch is used to predict the center coordinate offset of the object and the size of the object according to the output of the FCN fully convolutional network. The loss of the regression branch is calculated by weighting the l1 loss and the GIoU loss. The l1 loss is to calculate the difference between the predicted coordinates and the true coordinates in each coordinate system, and then sum and divide by the number of coordinate systems. Different from IoU which only focuses on the overlapping area, GIoU not only focuses on the overlapping area but also on other non-overlapping areas, and can better reflect the coincidence degree of the two. The calculation process is as follows:

[0201]

[0202]

[0203]

[0204] L giou = 1 - GIoU

[0205]

[0206] The superscript p in the upper right corner represents the true coordinates, x p represents the true x coordinate, y p represents the true y coordinate. L 1 is the l1 loss, IoU represents the intersection over union of the area of the true target box and the predicted target box, U represents the intersection of the area of the true target box and the predicted target box, A c represents the area of the minimum closed region of the two boxes. Area i represents the area of the predicted box in the i-th frame, represents the area of the true box in the i-th frame. L giou is the GIoU loss, L reg is the loss of the bounding box regression, λ giou and are hyperparameters, and in this embodiment, λ giou = 2 and

[0207] Step S6.3 Loss function calculation

[0208] To train the model of the present invention in an end-to-end manner, it is converted into a multi-task optimization problem, and the cross-modal alignment loss, the intra-modal loss, the classification loss, and the regression loss are optimized simultaneously.

[0209] L = λ cross L cross + λintra L intra +L reg +L cls

[0210] where λ cross = 1, λ intra = 1.

[0211] Step S6.4 Target Coordinate Output

[0212] By calculating the loss to perform backpropagation to complete parameter update, further correct the position of the prediction box to make the accuracy higher. The coordinates of the final target are represented as follows:

[0213] (x d , y d ) = argmax (x,y) DP xy

[0214] (x, y, w, h) = (x d + O(0, x d , y d ), y d + O(1, x d , y d ), S(0, x d , y d ), S(1, x d , y d ))

[0215] (x d , y d ) are the coordinates of the point with the highest score in the target classification score map DP of the prediction output. (x, y) represents the coordinates of the center of the predicted target, and (w, h) represents the width and height of the predicted target. is the local offset of each point in the two coordinate dimensions. O(0, x d , y d ) represents the local offset corresponding to the coordinate (x d , y d ) in the x - coordinate dimension, and O(1, x d , y d ) represents the local offset corresponding to the coordinate (x d , y d ) in the y - coordinate dimension. is the normalized bounding box size in the two coordinate dimensions. S(0, x d , y d ) represents the width of the bounding box corresponding to the coordinate (x d , y d ), and S(1, x d , y d ) represents the width of the bounding box corresponding to the coordinate (xd , y d The height of the bounding box of

[0216] Generally speaking, the advantages and beneficial technical effects of the present invention are as follows:

[0217] Enhanced perception ability: Visual and audio information provide different aspects of the target. Vision can capture the appearance, shape, and movement trajectory of the target, while audio provides sound information about the environment around the target and the target itself. Fusing the information of these two modalities can improve the system's perception ability of the target, enabling it to understand the state of the target more comprehensively.

[0218] Improved robustness: A single modality may be affected by certain environmental conditions or interference, while multi-modal fusion can improve the robustness of the system. When one modality is interfered with, the other modality may still provide valid information, thus helping to maintain the tracking performance of the target.

[0219] Spatio-temporal consistency: Visual and audio information can complement each other spatio-temporally. For example, when observing the position change of a target in a video, the audio information may provide additional clues about the movement of the target. By fusing these two types of information, the spatio-temporal consistency of the target can be better captured.

[0220] Rich semantic understanding: Audio information can provide semantic information about the target, such as the sound characteristics emitted by the target. Fusing this semantic information can enable the system to better understand the behavior and state of the target, helping to track and predict more accurately.

[0221] Embodiment 2

[0222] Based on the same inventive concept, this embodiment discloses a visual-audio multi-modal target tracking device based on a cross-modal Transformer, including:

[0223] An input module for inputting the search area image, the template area image, the audio segment corresponding to the image frame of the search area image, and the audio segment corresponding to the image frame of the template area image;

[0224] An image tokens extraction module for extracting image tokens from the input search area image and template area image;

[0225] An audio tokens extraction module for extracting audio tokens from the audio segment corresponding to the image frame of the input search area image and the audio segment corresponding to the image frame of the template area image;

[0226] The cross-modal alignment module is used to perform cross-modal alignment on the extracted image tokens and audio tokens to obtain the cross-modal aligned image tokens and audio tokens. Among them, the cross-modal aligned image tokens include the cross-modal aligned search area image tokens and the cross-modal aligned template area image tokens; perform intra-modal alignment on the cross-modal aligned search area image tokens and the modal-aligned template area image tokens to obtain the cross-modal aligned image tokens;

[0227] The feature extraction and fusion module is used to perform modal mixing operations on the cross-modal aligned audio tokens and the cross-modal aligned image tokens, and use a multi-layer Transformer encoder for feature extraction and fusion to obtain the fused multi-modal features. The fused multi-modal features include search area image features and template area image features;

[0228] The bounding box prediction head module is used to obtain the coordinates of the target in the search area image based on the search area image features output by the last layer of the encoder.

[0229] The visual-audio multi-modal object tracking device based on cross-modal Transformer is the Figure 1 model in. The feature extraction and fusion module uses the Figure 4 Transformer encoder in.

[0230] Since the device introduced in the second embodiment of the present invention is the device used to implement the visual-audio multi-modal object tracking method based on cross-modal Transformer in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the device, so it will not be elaborated here. Any device used in the method of the first embodiment of the present invention belongs to the scope protected by the present invention.

[0231] Embodiment Three

[0232] Based on the same inventive concept, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed, it implements the method described in Embodiment One.

[0233] Since the computer-readable storage medium introduced in the third embodiment of the present invention is the computer-readable storage medium adopted in the method for cross-modal Transformer-based visual-audio multi-modal object tracking in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the computer-readable storage medium, so it will not be elaborated here. Any computer-readable storage medium adopted in the method of the first embodiment of the present invention falls within the scope of protection of the present invention.

[0234] Embodiment 4

[0235] Based on the same inventive concept, the present application also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the above program, the method in Embodiment 1 is implemented.

[0236] Since the computer device introduced in the fourth embodiment of the present invention is the computer device adopted in the method for cross-modal Transformer-based visual-audio multi-modal object tracking in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the computer device, so it will not be elaborated here. Any computer device adopted in the method of the first embodiment of the present invention falls within the scope of protection of the present invention.

[0237] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0238] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0239] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention. Obviously, those skilled in the art can make various changes and variations to the embodiments of the present invention without departing from the spirit and scope of the embodiments of the present invention. Thus, if these modifications and variations of the embodiments of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and variations.

Claims

1. A visual-audio multimodal target tracking method based on a cross-modal Transformer, characterized in that: include: S1: input a search area image, a template area image, an audio clip of an image frame corresponding to the search area image, and an audio clip of an image frame corresponding to the template area image; S2: extract image tokens from the input search area image and template area image; S3: extracting audio tokens from the audio clip of the image frame corresponding to the input search area image and the audio clip of the image frame corresponding to the template area image; S4: performing cross-modal alignment on the extracted image tokens and audio tokens to obtain cross-modal aligned image tokens and audio tokens, wherein the cross-modal aligned image tokens include cross-modal aligned search area image tokens and cross-modal aligned template area image tokens; performing intra-modal alignment on the cross-modal aligned search area image tokens and the modality aligned template area image tokens to obtain cross-modal aligned image tokens; S5: Perform modal mixing operation on the audio tokens after cross-modal alignment and the image tokens after cross-modal alignment, and use a multi-layer Transformer encoder to extract and fuse features to obtain fused multi-modal features. The fused multi-modal features include search area image features and template area image features. S6: Obtain the coordinates of the target in the search area image according to the search area image features output by the last layer encoder; Among them, step S4 performs cross-modal alignment on the extracted image tokens and audio tokens to obtain cross-modal aligned image tokens and audio tokens, including: Calculate B 2 The cosine similarity between the tokens of the image and audio pairs, where the image and audio pairs from the same time are called positive pairs, and the image and audio pairs from different times are called negative pairs; By maximizing the cosine similarity of B pairs, while minimizing B 2 -The cosine similarity of B negative pairs is used to train the audio encoder, and the training process is calculated using the contrastive loss; Get cross-modal aligned image and audio tokens through the trained audio encoder; Step S4 performs intra-modality alignment on the search area image tokens after cross-modality alignment and the template area image tokens after cross-modality alignment, including: The tokens of the search area and template area in the cross-modal aligned visual modality are fused, where the search area image and template area image from the same time are regarded as positive pairs, and the search area image and template area image from different times are regarded as negative pairs; The tokens of the cross-modal aligned images are obtained by calculating the contrast loss.

2. The visual-audio multimodal target tracking method based on a cross-modal Transformer according to claim 1, characterized in that: Step S2 includes: S2.1: Segment the input search area image and template area image into non-overlapping image block groups respectively; S2.2: Apply linear projection to the segmented image block group to generate corresponding search area image tokens and template area image tokens; S2.3: Add learnable position embeddings to maintain the position information of image patches.

3. The visual-audio multimodal target tracking method based on a cross-modal Transformer as claimed in claim 1, characterized in that: Step S3 includes: S3.1: downsampling the audio clip of the image frame corresponding to the input search area image and the audio clip of the image frame corresponding to the template area image; S3.2: Divide the downsampled audio clip into short time frames and add a Hamming window; S3.3: Apply short-time Fourier transform to each short-time frame to obtain amplitude spectrum and energy spectrum, and apply Mel filter bank to obtain Mel filter bank features; S3.4: Extract the first L Mel-frequency cepstral coefficients, and then use a linear projection layer to obtain D-dimensional audio tokens, and then add a learnable position embedding; S3.5: Add a CLS token with a class tag at the beginning of the audio tokens sequence to indicate the category information of the entire audio sequence.

4. The visual-audio multimodal target tracking method based on a cross-modal Transformer as claimed in claim 1, characterized in that: Step S5 includes: S5.1: Inject the cross-modal aligned audio tokens into the two cross-modal aligned image tokens in the visual modality respectively to obtain the search area visual embedding and the template area visual embedding of the injected audio information; S5.2: The search area visual embedding and the template area visual embedding of the injected audio information are concatenated and then sent to an L-layer Transformer encoder for joint processing. Each layer of the encoder includes a multi-head self-attention module and a feedforward network. The multi-head self-attention module uses a multi-head self-attention mechanism to allow the model to interact with the input sequence from multiple attention perspectives and capture different feature relationships in the input sequence. The feedforward network introduces nonlinear transformation of features and obtains fused multimodal features based on the output of the multi-head attention module.

5. The visual-audio multimodal target tracking method based on a cross-modal Transformer as claimed in claim 1, characterized in that: Step S6 includes: S6.1: Use the classification branch to reshape the search area image features output by the last layer of the Transformer encoder into a 2D feature map of the original spatial resolution, and then pass it through a one-layer fully convolutional network to obtain the target classification score map, local offset, and normalized bounding box size. The point with the highest score in the target classification score map is the target center point, and the local offset is used to compensate for the discretization error caused by the reduced resolution. S6.2: Use the bounding box regression branch to predict the center coordinate offset and size of the object. The loss in the prediction process is Weighted calculation of loss and GIoU loss; S6.3: Convert the cross-modal Transformer-based visual-audio multimodal object tracking model into a multi-task optimization problem, and optimize the cross-modal alignment loss, intra-modal alignment loss, classification loss, and regression loss simultaneously; S6.4: The parameter update is completed by calculating the loss and then transferring it backwards to further correct the position of the predicted box, and finally the coordinates of the target center point and the size of the bounding box are obtained.

6. A visual-audio multimodal target tracking device based on a cross-modal Transformer, characterized in that: include: An input module, used for inputting a search area image, a template area image, an audio clip of an image frame corresponding to the search area image, and an audio clip of an image frame corresponding to the template area image; Image tokens extraction module, used to extract image tokens from the input search area image and template area image; An audio tokens extraction module, used to extract audio tokens from the audio clips of the image frames corresponding to the input search area images and the audio clips of the image frames corresponding to the template area images; A cross-modal alignment module is used to perform cross-modal alignment on the extracted image tokens and audio tokens to obtain image tokens and audio tokens after cross-modal alignment, wherein the image tokens after cross-modal alignment include search area image tokens after cross-modal alignment and template area image tokens after cross-modal alignment; the search area image tokens after cross-modal alignment and the template area image tokens after modality alignment are intra-modally aligned to obtain image tokens after cross-modal alignment; The feature extraction and fusion module is used to perform modal mixing operations on the audio tokens after cross-modal alignment and the image tokens after cross-modal alignment, and use a multi-layer Transformer encoder to extract and fuse features to obtain fused multi-modal features. The fused multi-modal features include search area image features and template area image features. The bounding box prediction head module is used to obtain the coordinates of the target in the search area image according to the search area image features output by the last layer encoder; Among them, the cross-modal alignment module is also used for: Calculate B 2 The cosine similarity between the tokens of the image and audio pairs, where the image and audio pairs from the same time are called positive pairs, and the image and audio pairs from different times are called negative pairs; By maximizing the cosine similarity of B pairs, while minimizing B 2 -The cosine similarity of B negative pairs is used to train the audio encoder, and the training process is calculated using the contrastive loss; Get cross-modal aligned image and audio tokens through the trained audio encoder; The Cross-Modal Alignment module is also used to: The tokens of the search area and template area in the cross-modal aligned visual modality are fused, where the search area image and template area image from the same time are regarded as positive pairs, and the search area image and template area image from different times are regarded as negative pairs; The tokens of the cross-modal aligned images are obtained by calculating the contrast loss.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

8. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 5 is implemented.