Audio-visual matching method based on relation awareness attention correction network
Through the method of correcting attention network based on relationship perception, the problem of lack of fine feature decomposition and interference terms in audio-visual matching is solved, and the fine matching of cross-modal features and the robustness of the model is achieved.
Patent Information
- Application Number
- CN202510181683.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-07-19
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to finely adjust subtle features in audio-visual matching, which may lead to focusing only on some attribute features or ignoring key identity representation features, and interference items such as background noise and multi-person conversations will reduce the accuracy of audio identity information.
The method based on relationship perception correction attention network is adopted, and the parallel adaptive intramodal correction attention module and the relationship perception intermodal correction attention module are combined with feature orthogonal constraints and dynamic residual weight combination strategies to perform the optimal matching of cross-modal features, and modal-independent features are generated through the adversarial network.
Effectively explore the intrinsic connections between cross-modal semantic features, eliminate the pseudo-correlation of interfering information, improve the accuracy and robustness of audio-visual matching, and enhance the generalization ability of the model.
Smart Images

Figure CN120086408A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer vision and cross-modal biometric matching technologies, and particularly to an audiovisual matching method based on a relationship-aware correction attention network. Background Art
[0002] Neuroscientists have confirmed that both voices and faces contain unique identity information suitable for identification. Therefore, people usually naturally associate faces with voices. For example, just by hearing a familiar voice, we can match it with the corresponding face, which is the so-called "identifying a person by voice". This has promoted the rise of a new research field (audiovisual matching), which aims to study how different biometric modalities interact and use neuroimaging to provide a scientific basis for these interactions. Audiovisual matching technology has been widely applied in many research and application scenarios such as criminal investigation security, intelligent monitoring systems, and intelligent voice recording. A major challenge in this field is to mimic the human's perception ability of the correlation of multimodal biometric signals. This involves extracting features from face images and audio segments to establish the connection between cross-modal identity features.
[0003] Existing research has shown that adversarial training can significantly reduce the problem of heterogeneity between modalities, while subtle feature extraction helps to reduce the impact of identity differences. Nevertheless, such methods may inadvertently capture irrelevant features, thereby impairing the generalization ability of the model. In addition, during an interview, background noise, the content of the interviewee's speech, and the overlapping sounds of multiple people talking will all reduce the accuracy of audio identity information. These interference items can lead to overfitting of global features in a single modality, further increasing the complexity of the audiovisual matching task.
[0004] To solve this problem, one method is to use modal decomposition to distinguish identity-related features and modality-related features, and process these features separately to learn effective feature representations and filter out interference features. Another method is to randomly discard keys before calculating the attention matrix to reduce the model's overfitting to interference features. These methods all require adjusting parameters according to experience and specific tasks to achieve the best performance. Therefore, using a soft threshold layer can be used as an adaptive disentanglement method to achieve the purpose of feature decomposition.
[0005] However, as a disentanglement technique, the soft threshold layer fails to completely cut off the connection between identity features and interference features due to the lack of effective separation constraints. To address this problem, two strategies are usually adopted: one is attribute decomposition, which uses an attribute-guided strategy to identify the hierarchical relationship between different attribute features; the other is data decomposition, which adjusts the sample importance of face images and sound segments through a dynamic weighting strategy to avoid the network overfitting to samples containing a large amount of perturbation information.
[0006] However, neither of the above two methods can finely adjust subtle features, which may lead to only focusing on some attribute features or ignoring key identity representation features. Therefore, it is necessary to design an effective intervention strategy to assist the soft threshold layer to achieve effective feature decomposition, so as to establish a matching relationship for effective audiovisual modality representation features. Summary of the Invention
[0007] Object of the Invention: The object of the present invention is to solve the deficiencies existing in the prior art and provide an audiovisual matching method based on a relationship-aware correction attention network.
[0008] Technical Solution: An audiovisual matching method based on a relationship-aware correction attention network of the present invention includes the following steps:
[0009] Step 1: First, obtain an anchor audio segment and the corresponding k face images Take the anchor audio as the recognition item and the face images as the gallery matching items; then respectively extract features from the anchor audio segment and the face images to obtain the original audio segment feature f i a and the face image feature i represents the i-th data tuple; k>1 represents the matching situation, and k = 1 represents a special matching situation;
[0010] Step 2: Send the feature f of the audio segment i a and the feature of the face image into the relationship-aware correction attention network, and calculate the similarity between the audio segment and a single face image respectively through the relationship-aware network to obtain their respective attention matrices and The relationship-aware correction attention network includes a parallel adaptive intra-modal correction attention module RIRA and a relationship-aware inter-modal correction attention module AIRA;
[0011] The relationship-aware inter-modal correction attention module AIRA uses the adaptive attention correction unit AARU to respectively adjust and to obtain the attention matrices and Then, perform first-order inter-modal interaction and second-order inter-modal interaction on the two attention matrices in sequence, and finally obtain the features and
[0012] The adaptive intra-modal correction attention module RIRA uses the adaptive attention correction unit AARU to respectively correct and to respectively obtain the intra-modal corrected attention matrices and Then, respectively, for and use the decoder to extract relevant modal features and
[0013] Step 3: First, introduce the feature orthogonality constraint L orth , so that the modality-independent features and modality-related features complement each other, ensuring the feature orthogonality between the two modalities; use the dynamic residual weight combination method to combine the dynamically selected intra-modal and inter-modal enhanced features and global features for the optimal matching between cross-modal features. The features of the two modalities with optimal matching are respectively denoted as and
[0014] Step 4: Use the audio features and the face image features through the adversarial network discriminator to generate modality-independent features
[0015] Furthermore, in Step 1, the correlation estimation between modalities is also calculated through to evaluate the semantic correlation between different modalities. The calculation formula is as follows:
[0016]
[0017] The above formula maps the semantics of audio and visual images to the same shared space through the convs convolutional layer;
[0018] Next, use as the pseudo-label, and perform further supervised learning by assigning it the true label to calculate the relationship-aware loss L RA as follows:
[0019]
[0020] Among them, "epoch" represents the current number of iterations, and "N" represents the total number of iterations. σ represents the sigmoid activation function; the temperature control parameter is set to 5; the directional guidance interaction depends more on the positive and negative of the correlation rather than its specific numerical value. Therefore, a simulated annealing weight is introduced to reduce the impact of the loss on the model.
[0021] To focus on the relevant semantics in cross-modal interaction and use context-aware correlation to guide this interaction, the present invention introduces an inter-modal correction attention mechanism. The specific working content of the relationship-aware inter-modal correction attention module AIRA is as follows:
[0022] First, for the similarity between the audio segment and the face image, an attention matrix is obtained and The expressions are as follows:
[0023]
[0024] In the above formula, and respectively refer to the attention matrices of the image and the audio;
[0025] Then, the cross-modal semantic threshold is learned through the Adaptive Attention Rectification Unit (AARU) to reduce the influence of interference information, and the attention matrix is adjusted to obtain the adjusted attention matrix and The expressions are as follows:
[0026]
[0027] In the above formula, ⊙ represents element-wise multiplication, clamp represents a truncation operation, and sign represents the sign function, which is defined as follows:
[0028]
[0029] Next, a first-order cross-modal interaction is performed on the audio and the face image, and the expressions are as follows:
[0030]
[0031] and respectively represent the audio segment feature and the k-th face image feature in the first-order cross-modal interaction. The heterogeneity between modalities makes it difficult to fully explore the cross-modal semantic connections only through the first-order cross-modal interaction;
[0032] Furthermore, the similarity matrix of the first-order cross-modal interaction is sparsified using the hyperparameter δ to identify the closest local semantic associations, and the semantic connections are enhanced through the intra-modal attention mechanism; finally, features are extracted through second-order attention interaction, and the expressions are as follows:
[0033]
[0034] In the above formula, decodes is a decoding convolutional layer used to restore the features shared by the audio and the face image to the original feature dimension, and diffuse the feature regions of concern through and
[0035] To perform modality alignment, adaptive intra-modal attention is introduced here. The working content of the Adaptive Intra-modal Rectification Attention Module (RIRA) is as follows:
[0036] To reduce the background interference within the modality and strengthen the representation of foreground features, the attention matrix after in-modal correction is first obtained and The expressions are as follows:
[0037]
[0038] where conv p is a dimensionality reduction convolution operation, and ε is a positive threshold for online learning; for the elements in matrices and that are positive and less than ε, their values are set to zero through the clamp function, and the elements that are negative and greater than -ε are set to zero through the clamp function;
[0039] Then, they are decoded respectively to obtain the following features:
[0040]
[0041] where δ is the constant term mentioned above, which is set to 10 in the experiment; decode p is an upsampling convolution layer used to restore the modality-related features to the original extraction dimension.
[0042] Furthermore, the expression of the feature orthogonality constraint L orth in step 3 is as follows:
[0043]
[0044] where G is the max pooling. M represents the number of mini-batch data tuples.
[0045] To adjust the global characteristics between the two modalities and enhance the semantics inside and outside the modality, an optimal matching feature is obtained through the dynamic residual weight combination strategy, and the expression is as follows:
[0046]
[0047] where m s represents the modality-related feature weight, and mp represents the modality-unrelated feature weight; these two weights are calculated through SENet; to reduce the variance of the modality features, instance normalization (IN) is adopted.
[0048] Furthermore, after the discriminator generates modality-independent features in step 4, the non-linear discrimination ability of the multi-layer perceptron is used to identify potential matching items, and the cross-entropy method is adopted to calculate the loss value;
[0049]
[0050] where Cm represents the matching classification; l i is the matching identity tag;
[0051] To enhance the robustness of audiovisual matching and improve the efficiency of identity association recognition, the relative distance stretching metric loss is used to enlarge the distance between identity-related features and narrow the interval of non-identity-related features. The cross-modal relative distance stretching metric loss is as follows:
[0052]
[0053] where the hyperparameters μ 1 and μ 2 are set to 10 and 4 respectively, and the parameter θ is set to 1.2; represents the Euclidean distance between the anchor sample and the positive sample , is the Euclidean distance between the anchor sample and the negative sample , while is the Euclidean distance between the positive sample and the negative sample ; the label indicates whether the identity matches. When it does not match, is 1, and when it matches, is 0; the label helps calculate the distance between each negative sample in the candidate samples. To achieve the relative distance stretching constraint, the negative sample closest to the positive sample is selected from among the numerous candidate negative samples, and its distance is emphasized.
[0054] Furthermore, the total loss is:
[0055] L = L cls + L dis + L gen + L RDSM ,
[0056] L cls is the cross-entropy calculation loss, and L RDSM is the cross-modal relative distance stretching metric loss.
[0057] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0058] (1) The relationship-aware correction attention network proposed by the present invention can effectively explore the internal connection between cross-modal semantic features. This ability is achieved through the synergistic effect of the relationship-aware inter-modal correction attention and the adaptive intra-modal correction attention designed as a parallel structure.
[0059] (2) The adaptive attention correction unit of the present invention can eliminate the pseudo-correlation caused by interference information and guide the mutual enhancement of cross-modal semantic features.
[0060] (3) The relative distance stretching metric loss in the present invention utilizes the distance relationship between the anchor, positive samples, and negative samples to further push away the mismatched features, thereby promoting a robust audiovisual matching embedding representation. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 is a schematic diagram of the overall process of the present invention;
[0062] Figure 2 is a schematic diagram of the network model of the present invention;
[0063] Figure 3 is a schematic diagram of the dynamic residual weight combination in the present invention;
[0064] Figure 4 is a comparison of the quantitative comparison results of the prior art in the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0065] The technical solution of the present invention will be described in detail below, but the protection scope of the present invention is not limited to the described embodiments.
[0066] As Figure 1 shown, the audiovisual matching method based on the relationship-aware correction attention network of the present invention includes the following steps:
[0067] Step 1: First, obtain an anchor audio clip and the corresponding k face images Take the anchor audio as the recognition item and the face images as the gallery matching items; then extract features from the anchor audio clip and the face images respectively to obtain the original audio clip feature f i a and the face image feature i represents the i-th data tuple;
[0068] Step 2: Feed the feature f of the audio clip i a and the feature of the face image into the relationship-aware correction attention network. The relationship-aware network calculates the similarity between the audio clip and a single face image respectively to obtain their respective attention matrices and The relationship-aware correction attention network includes a parallel adaptive intra-modal correction attention module RIRA and a relationship-aware inter-modal correction attention module AIRA;
[0069] The relation-aware inter-modal correction attention module AIRA uses the adaptive attention correction unit AARU to separately adjust and to obtain the attention matrices and Then, the first-order inter-modal interaction and the second-order inter-modal interaction are successively performed on the two attention matrices, and finally the features and
[0070] The self-adaptive intra-modal correction attention module RIRA uses the adaptive attention correction unit AARU to separately correct and to obtain the intra-modal corrected attention matrices and Then, the decoder is used to extract the relevant modal features from and respectively, obtaining and
[0071] Step 3: First, introduce the feature orthogonality constraint L orth , so that the modality-independent features and the modality-related features complement each other; use the dynamic residual weight combination method to combine the dynamically selected intra-modal and inter-modal enhanced features and the global features for the optimal matching between cross-modal features. The features of the two modalities with the optimal matching are respectively denoted as and
[0072] Step 4: Since the present invention ultimately aims to generate modality-independent features, due to the heterogeneity between the two modalities, the influence cannot be completely eliminated only through interaction and feature embedding, which results in the learned feature differences interfering with the final discrimination result; the present invention uses the audio feature and the face image feature by the discriminator of the adversarial network to generate the modality-independent feature to reduce the influence of modality heterogeneity.
[0073] In this embodiment, the data can be divided into a training set, a validation set, and a test set, and these data are applied according to their respective task characteristics. After adjusting the image data to a size of 3×224×244, it is input into the ResNet network for feature extraction, and the model parameters pre-trained on the ImageNet dataset are utilized. We use the Python module Librosa to analyze the audio signal and extract features through Mel-frequency cepstral coefficients (MFCC). MFCC is a set of about 10 - 20 features that generally describe the envelope shape of the spectrum. MFCC simulates the characteristic waveform of human voices and further extracts high-level semantic information through the ResNet network. Finally, the corresponding modal semantic representations are extracted from the original audio and image information.
[0074] As Figure 2 shown, in the present invention, both the cross-modal rectified attention mechanism and the adaptive intra-modal rectified attention mechanism are to effectively focus on the relevant semantics in cross-modal interactions and use context-aware correlation to guide such interactions. In this embodiment, the semantic correlation between different modalities is estimated to evaluate the semantic correlation between different modalities, and the calculation formula is as follows:
[0075]
[0076] In the above formula, the semantics of audio and visual images are mapped to the same shared space through the convs convolutional layer;
[0077] Next, is used as a pseudo-label, and further supervised learning is carried out by assigning it a true label, and the relationship-aware loss L RA is as follows:
[0078]
[0079] where epoch represents the current iteration number, N represents the total number of iterations, and σ represents the sigmoid activation function. The obtained pseudo-label is also used during testing to guide the interaction of cross-modal features.
[0080] For the audiovisual matching method based on the relationship-aware rectified attention network in this embodiment, the specific working content of the relationship-aware cross-modal rectified attention module AIRA is as follows:
[0081] First, for the similarity between the audio segment and the face image, the attention matrices and are obtained, and the expressions are as follows:
[0082]
[0083] In the above formula, and respectively refer to the attention matrices of images and audio;
[0084] Then, the cross-modal semantic threshold is learned through the Adaptive Attention Rectification Unit (AARU) (to reduce the influence of interference information and improve the accuracy of audiovisual matching), and the attention matrix is adjusted to obtain the adjusted attention matrix and The expressions are as follows:
[0085]
[0086] In the above formula, ⊙ represents the element-wise product, clamp represents the truncation operation, and sign represents the sign function, which is defined as follows:
[0087]
[0088] Since the correlation between audio and face images is established on all relevant segments, the correlation of each segment is adjusted and the attention map is updated to focus on learning the correlation between cross-modal semantics. This process is called first-order inter-modal interaction; furthermore, the expression for the first-order inter-modal interaction between audio and face images is as follows:
[0089]
[0090] Moreover, the hyperparameter δ is used to sparsify the similarity matrix of the first-order inter-modal interaction, and the semantic connection is enhanced through the intra-modal attention mechanism; finally, the features are extracted through the second-order attention interaction, and the expression is as follows:
[0091]
[0092] In the above formula, represents the features of the k-th face image, decodes is a decoding convolutional layer used to restore the features shared by audio and face images to the original feature dimension, and the hyperparameter δ sparsifies the similarity matrix of the first-order inter-modal interaction to more accurately locate the closely related local semantic regions.
[0093] To address the challenge of significant information differences between audiovisual modalities, eliminate the influence of the external environment on the information differences between audio and face images, and achieve cross-modal interaction and alignment to obtain a modality-independent feature representation rich in semantics, this implementation adopts the Adaptive Intra-modal Rectification Attention Module (RIRA) for effective feature representation, and the work content is as follows:
[0094] First, obtain the attention matrix after intra-modal correction and The expressions are as follows:
[0095]
[0096] Among them, conv p is a dimensionality reduction convolution operation, and ε is a positive threshold for online learning; for the elements in the matrices and that are positive and less than ε, the clamp function is used to set them to zero, and the elements that are negative and greater than -ε are set to zero by the clamp function;
[0097] Then, they are decoded separately to obtain the following features:
[0098]
[0099] Among them, δ is a constant term, and decode p is an upsampling convolution layer.
[0100] To ensure that the modality-independent features and the relevant features are complementary, the feature orthogonality constraint L orth is used, and the expression is as follows:
[0101]
[0102] Among them, G is the max pooling; M represents the number of mini-batch data tuples
[0103] As Figure 3 shown, since the cross-modal matching effect in the matching process strongly depends on the global context relevance between modalities, this relevance is determined by the global and local enhanced features, which are indirectly extracted from each modality. Therefore, it is crucial to adjust the global features between different modalities, the enhanced semantics within and between modalities, in order to establish a tight correlation. The optimal matching feature expression obtained by the dynamic residual weight combination strategy in this embodiment is as follows:
[0104]
[0105] Among them, m s represents the feature weight related to the modality, and mp represents the feature weight independent of the modality.
[0106] Since interaction cannot fundamentally solve the heterogeneity problem between cross-modal features, this embodiment introduces the adversarial idea, and combines the discriminator with the audio feature f i a and the face image feature to generate modality-independent features In this process, the non-linear discrimination ability of the multi-layer perceptron is used to identify potential matching items, and the cross-entropy method is used to calculate the loss value;
[0107]
[0108] where Cm represents the matching classification; l i is the matching identity tag.
[0109] And the relative distance stretching metric loss is used to enlarge the distance between identity-related features and narrow the interval of non-identity-related features. The cross-modal relative distance stretching metric loss is as follows:
[0110]
[0111] where the hyperparameters μ 1 and μ 2 are set to 10 and 4 respectively, and the parameter θ is set to 1.2; represents the Euclidean distance between the anchor sample and the positive sample , is the Euclidean distance between the anchor sample and the negative sample , while is the Euclidean distance between the positive sample and the negative sample ; the label indicates whether the identity matches. If it does not match it is 1, and if it matches it is 0.
[0112] The total training loss of the present invention is:
[0113] L = L cls + L dis + L gen + L RDSM ,
[0114] L cls is the cross-entropy calculation loss, and L RDSM is the cross-modal relative distance stretching metric loss. During the training process, the shared layer is updated jointly by all training data, while the fully connected layer of each branch is iteratively updated according to the corresponding modal features.
[0115] From Figures 1 to 3 It can be seen that the entire relationship-aware correction attention network of the present invention includes relationship-aware correction attention, adaptive intra-modal correction attention, and relative distance stretching metric loss. The relationship-aware network enhances the interaction of cross-modal semantic features through correlation guidance, eliminates the pseudo-correlation in the interference information, and restricts the distribution of different identity samples, so that the model can learn stable cross-modal representation features to achieve more accurate audiovisual cross-modal matching.
[0116] Embodiment
[0117] To further verify the performance of the technical solution of the present invention, in this embodiment, a public dataset is used to compare the present invention with the prior art.
[0118] The VoxCeleb dataset is a public audio dataset that contains more than 149,354 audio segments involving 1,225 different speakers. The VGGFace dataset is a public face image dataset that contains 137,060 facial images extracted from the same videos. In this embodiment, individuals whose names start with "A" or "B" are assigned to the validation set, individuals whose names start with "C", "D", or "E" are assigned to the test set, and individuals whose names start with letters from "F" to "Z" are assigned to the training set.
[0119] The network in this embodiment was trained for 50 epochs with a batch size of 50 on an NVIDIA GeForce RTX 3090. It was trained using the stochastic gradient descent optimization method with a learning rate of 10 -5 , with the input face image size of 3×224×224 and the audio sequence of 1×160000.
[0120] For the convenience of quantitative evaluation, this embodiment uses the accuracy metric to measure the performance.
[0121] Accuracy:
[0122] The quantitative comparison results are shown in Table 1
[0123] Table 1 Quantitative experiments on two datasets
[0124]
[0125] Next, qualitative evaluation is carried out:
[0126] The audiovisual matching performance was compared on the public datasets VoxCeleb and VGGFace.
[0127] The results are as Figure 4 shown. The adversarial-based methods are generally superior to the feature-embedding-based methods in audio-visual matching performance because adversarial techniques can better eliminate the heterogeneity between modalities and thus extract relevant cross-modal features. However, due to the interference factors in the data, it is difficult to rely solely on adversarial modeling to enhance the generalization ability of the model, which also hinders the further improvement of the model performance. The present invention is superior to the existing adversarial-based methods in terms of validation, binary matching, and multi-way matching, and significantly improves the performance in both the V-F and F-V scenarios.
[0128] In addition, a 1:k multi-directional cross-modal matching experiment was conducted in this embodiment. As the number of matching candidates increases, the matching difficulty correspondingly rises. Nevertheless, the technical solution of the present invention still demonstrates strong competitiveness.
[0129] Furthermore, compared with the second-ranked method, the average performance of the present invention in the V-F and F-V scenarios has been significantly improved. This indicates that the relationship-aware rectified attention network can effectively bridge the connection between cross-modal features by perceiving the correlation and exclude interference information. Here, V-F refers to identifying multiple candidate face image databases based on audio clip recognition to achieve identity matching, such as Figure 4 (a) in Figure 4 ; conversely, it is called F-V, such as
[0130] (b) in
Claims
1. An audio-visual matching method based on a relation-aware rectified attention network, characterized in that: The following steps are involved: Step 1: Get an anchor audio clip And the corresponding k face images The anchor audio is used as the recognition item and the face image is used as the gallery matching item; then the features of the anchor audio segment and the face image are extracted respectively to obtain the original audio segment features f i a and facial image features i represents the i-th data tuple; Step 2: Transform the audio clip feature f i a and features of face images The audio clips are sent to the relation-aware corrected attention network, and the similarity between the audio clips and the single face image is calculated through the relation-aware network to obtain their respective attention matrices. and The relation-aware corrective attention network includes a parallel adaptive intra-modality corrective attention module RIRA and a relation-aware inter-modality corrective attention module AIRA; The relation-aware inter-modal correction attention module AIRA uses the adaptive attention correction unit AARU to and Adjust to get the attention matrix and Then, the two attention matrices are subjected to first-order inter-modal interaction and second-order inter-modal interaction in turn, and finally the feature and The adaptive intra-modality correction attention module RIRA uses the adaptive attention correction unit AARU to and Correction is performed to obtain the corrected attention matrix within the modality and Then respectively and Use decoder to extract relevant modality features and Step 3: First introduce the feature orthogonality constraint L orth , so that the features unrelated to the modality and the features related to the modality complement each other; the dynamic residual weight combination method is used to combine the dynamically selected modality internal and inter-modality enhanced features with the global features to perform the optimal matching between cross-modality features. The features of the two modalities with the best matching are recorded as and Feature orthogonality constraint L orth The expression is as follows: Among them, G is the maximum pooling, and M represents the number of small batch data tuples; The optimal matching feature expression obtained by the dynamic residual weight combination strategy is as follows: Among them, m s represents the feature weight related to the modality, and mp represents the feature weight independent of the modality; Step 4: Exploiting Audio Features with Adversarial Network Generator and facial image features To generate modality-independent features 2. The audio-visual matching method based on relation-aware corrected attention network according to claim 1, characterized in that: Step 1 also calculates the inter-modal correlation estimate pass Evaluate the semantic relevance between different modalities. The calculation formula is as follows: The above formula maps the semantics of audio and visual images to the same shared space through the convs convolutional layer; Next, Used as a pseudo-label, further supervised learning is performed by assigning its true label and calculating the relationship-aware loss L RA as follows: Among them, epoch represents the current number of iterations, N represents the total number of iterations, and σ represents the sigmoid activation function.
3. The audio-visual matching method based on relation-aware corrected attention network according to claim 1, characterized in that: The specific working content of the relation-aware inter-modal correction attention module AIRA is: First, for the similarity between the audio clip and the face image, the attention matrix is obtained and The expression is as follows: In the above formula, and They refer to the attention matrices of images and audio respectively; Then, the adaptive attention correction unit AARU is used to learn the cross-modal semantic threshold to reduce the influence of interference information, and the attention matrix is adjusted to obtain the adjusted attention matrix and The expression is as follows: In the above formula, ⊙ represents the product of elements, clamp represents the truncation operation, and sign represents the sign function, which is defined as follows: Next, the first-order modal interaction between audio and face image is performed, and the expression is as follows: Furthermore, the similarity matrix of the first-order inter-modal interaction is sparsely processed using the hyperparameter δ, and the semantic connection is enhanced through the intra-modal attention mechanism; finally, features are extracted through the second-order attention interaction, and the expression is as follows: In the above formula, decoders is a decoding convolutional layer, which is used to restore the features shared by audio and face images to the original feature dimension.
4. The audio-visual matching method based on relation-aware corrected attention network according to claim 1, characterized in that: The working content of the adaptive intra-modality correction attention module RIRA is as follows: First get the attention matrix after intra-modal correction and The expression is as follows: Among them, conv p is a dimensionality reduction convolution operation, ε is a positive threshold for online learning; for the matrix and The positive elements less than ε are set to zero by the clamp function, and the negative elements greater than -ε are set to zero by the clamp function; Then, decode them separately and get the following features: Among them, δ is a constant term, decode p It is a dimensionality-enhancing convolutional layer.
5. The audio-visual matching method based on relation-aware corrected attention network according to claim 1, characterized in that: The step 4 generates modality-independent features {h i0 ,…,h ik }, the nonlinear discrimination ability of the multilayer perceptron is used to identify potential matches, and the cross entropy method is used to calculate the loss value; Where Cm represents the matching category; l i is the matching identity tag; And use the relative distance stretching metric loss to expand the distance between identity-related features and narrow the interval between non-identity-related features. The cross-modal relative distance stretching metric loss is as follows: The hyperparameters μ1 and μ2 are set to 10 and 4 respectively, and the parameter θ is set to 1.2; Represents anchor samples With positive samples The Euclidean distance between is the anchor sample With negative samples The Euclidean distance between It is a positive sample With negative samples The Euclidean distance between Indicates whether the identities match, if not 1 if it matches is 0.
6. The audio-visual matching method based on relation-aware rectified attention network according to claim 1, characterized in that: The total loss of the adversarial network is: L=L cls +L dis +L gen +L RDSM , L cls Calculate the loss for cross entropy, L RDSM Stretch metric loss for cross-modal relative distance.