Audio positioning model training method and device, storage medium and program product

By introducing a lightweight audio adapter into the CLAP model, extracting frame-level audio features and performing similarity training, the problem of insufficient performance of CLAP in frame-level tasks is solved, and more accurate audio signal understanding and time boundary prediction are achieved.

CN120260602APending Publication Date: 2025-07-04SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510430717.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing contrast language-audio pretrained models (CLAPs) lack fine-grained timing feature modeling capabilities in frame-level audio comprehension tasks, resulting in insufficient performance accuracy for tasks such as sound event detection and text-to-audio positioning.

Method used

By introducing a lightweight audio adapter into the CLAP model, the frame-level audio features are extracted and trained in combination with frame-level audio-phrase similarity and sound event tags, the model's frame-level audio comprehension capabilities are optimized.

Benefits of technology

The performance of the audio positioning model in frame-level audio understanding tasks is significantly improved, and more accurate audio signal understanding and time boundary prediction are achieved, suitable for sound event detection and text-to-audio positioning tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260602A_ABST
    Figure CN120260602A_ABST
Patent Text Reader

Abstract

The invention discloses an audio positioning model training method and device, a storage medium and a program product, and relates to the technical field of audio processing, and the method comprises the steps: obtaining an audio-subtitle sample which comprises an audio clip and a subtitle clip which are aligned at a time axis; based on the audio-subtitle sample and the contrast loss function, performing CLAP training on the audio positioning model; extracting frame-level audio features of the audio clip based on an audio adapter; calculating the frame-level audio-phrase similarity between the frame-level audio feature of each frame and the corresponding phrase embedding; according to the frame-level audio-phrase similarity and the sound event label corresponding to each frame, sound event classification training is carried out on the audio positioning model, and the sound event label is used for indicating whether the audio frame is matched with a real sound event described by phrase embedding or not. Therefore, the performance of the audio positioning model in the frame-level audio understanding task is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio detection, and in particular, to a training method, device, storage medium and program product of an audio localization model. Background Art

[0002] In recent years, significant progress has been made in the Contrastive Language-Audio Pre-training (CLAP) method. By aligning the contrastive losses of the audio encoder and the text encoder, these models can learn rich multimodal joint representations, thus promoting the implementation of a series of downstream tasks, including audio-text retrieval, zero-shot audio classification, automatic audio generation description, and text-to-audio generation.

[0003] Although researchers have conducted extensive exploration in improving the retrieval and classification capabilities of CLAP, relatively few studies have been carried out on how to utilize its semantically rich multimodal space for frame-level audio understanding tasks (such as Sound Event Detection (SED) and Text-to-Audio Grounding (TAG)). Different from the segment-level tasks that only require identifying the segments of sound events in the audio, frame-level tasks not only require the model to be able to complete event classification but also need to determine the corresponding time boundaries, thus posing higher requirements for the fine-grained understanding of audio signals. Although CLAP models perform well in segment-level tasks, they still lack the frame-level audio understanding ability and perform poorly in tasks such as text-to-audio grounding and sound event detection.

[0004] In response to the above problems, the industry has not yet proposed a better solution. Summary of the Invention

[0005] The present application provides a training method, an electronic device, a storage medium and a program product of an audio localization model, so as to at least solve the problem that the CLAP method lacks the ability to model the fine-grained temporal features of audio signals, resulting in insufficient performance accuracy of frame-level tasks such as sound event detection and text-to-audio grounding.

[0006] In a first aspect, an embodiment of the present application provides a method for training an audio localization model, including: obtaining an audio-caption sample, where the audio-caption sample includes an audio segment and a caption segment aligned on a timeline; performing CLAP training on the audio localization model based on the audio-caption sample and a contrastive loss function; extracting frame-level audio features of the audio segment based on an audio adapter; calculating frame-level audio-phrase similarity between the frame-level audio features of each frame and the corresponding phrase embedding; performing sound event classification training on the audio localization model according to the frame-level audio-phrase similarity corresponding to each frame and a sound event label, where the sound event label is used to indicate whether an audio frame matches the real sound event described by the phrase embedding.

[0007] In a second aspect, an electronic device is provided, which includes: at least one processor, and a memory communicatively connected to the at least one processor, where the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the steps of the method for training an audio localization model according to any embodiment of the present application.

[0008] In a third aspect, an embodiment of the present application provides a storage medium, on which a computer program is stored, and characterized in that when the program is executed by a processor, the steps of the method for training an audio localization model according to any embodiment of the present application are implemented.

[0009] In a fourth aspect, an embodiment of the present application provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method for training an audio localization model according to any embodiment of the present application are implemented.

[0010] The beneficial effects of the embodiments of the present application are as follows:

[0011] Based on the CLAP model training, by introducing an audio adapter to extract frame-level audio features and calculating frame-level audio-phrase similarity, and using the alignment joint constraint of frame-level similarity and sound event labels in model training, a more refined audio signal understanding ability is achieved, and the alignment accuracy of the multimodal feature space is optimized. As a result, the performance of the audio localization model in the frame-level audio understanding task is significantly improved, providing a more accurate and robust solution for task applications such as sound event detection and text-to-audio localization. Description of the Drawings

[0012] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0013] Figure 1 Fig. shows a comparison simulation effect diagram of an example of the native CLAP and the F-CLAP proposed in the embodiment of the present application in the sound event audio localization task;

[0014] Figure 2 Fig. shows an operation flowchart of an example of the training method of the audio localization model according to the embodiment of the present application;

[0015] Figure 3 Fig. shows a structural block diagram of an example of the audio adapter according to the embodiment of the present application;

[0016] Figure 4 Fig. shows a schematic diagram of the architecture comparison between the CLAP training paradigm and the F-CLAP training in an example

[0017] Figure 5 Fig. is a structural schematic diagram of an embodiment of the electronic device of the present application. Detailed implementation manners

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present application belong to the scope of protection of the present application.

[0019] It should be noted that since the frame-level task still essentially requires segment-level audio understanding, it can also benefit from the audio-text multimodal alignment learned by CLAP. In addition, due to the high cost of fine-grained time annotation, the scale of the frame-level annotated sound dataset is much smaller than the segment-level annotated dataset such as AudioSet. Introducing CLAP not only helps alleviate the problem of data scarcity but also enhances the generalization ability of the model.

[0020] However, the native contrast learning paradigm in CLAP mainly focuses on the coarse-grained alignment between the entire audio segment and the whole sentence caption, which limits the model's ability to perform temporal reasoning and locate the start and end times of sound events.

[0021] Figure 1The figure shows a schematic diagram of the comparative simulation effect of the original CLAP and the improved CLAP proposed in this paper in the audio localization task of sound events.

[0022] Sound Event Detection (SED) or sound event audio localization is a task of identifying sound events and their start and end times in a recording. As Figure 1 shown, the similarity graph of the frame embeddings and phrase representations of the original CLAP indicates that although the model can successfully identify the sounds of vehicles and bells, it lacks the temporal boundary information of these events.

[0023] In this paper, a simple and effective method is proposed to endow CLAP with frame-level audio understanding ability while retaining its original ability in segment-level tasks. As Figure 1 , it shows the similarity graph between the frame embeddings and phrase representations of the original CLAP and the F-CLAP proposed in this paper, as well as the ground truth labels indicating the corresponding sound events. In F-CLAP, the original coarse-grained global alignment in the CLAP multimodal space is refined to achieve fine-grained phrase-frame correspondence, thus providing more detailed temporal information and enabling the audio-language model to accurately locate the heard sound events. The original CLAP model is insufficient in the frame-level audio understanding task, while the method in this paper successfully improves the phrase-frame correspondence, enabling the audio-language model to accurately locate the time of sound events.

[0024] It should also be noted that in the current related technologies, some experts and scholars have proposed their respective technical solutions for cross-modal localization from text to audio, mainly including UACA (Unsupervised Audio-Caption Aligning), WS-TAG (Weakly Supervised Text-to-Audio Grounding), and MGA-CLAP (Multi-Grained Alignment for Contrastive Language-Audio Pre-training).

[0025] In UACA, the frame-word alignment matrix is aggregated by mean and max pooling to learn the detailed alignment between audio and text. UACA is a weakly-supervised text-to-audio localization (TAG) model, aiming to learn the alignment relationship between audio and text in an unsupervised manner. It learns the detailed alignment between audio and text by aggregating the frame-word alignment matrix (using mean and max pooling). This method does not require frame-level annotation data, but is trained with a large number of audio-text pairs. However, although UACA can learn the alignment relationship between audio and text by aggregating the frame-word alignment matrix through mean and max pooling, the pooling operation will lose fine-grained temporal information, resulting in insufficient alignment accuracy.

[0026] In WS-TAG, negative sampling and soft labels are used to build a powerful audio localization system. WS-TAG is another weakly-supervised text-to-audio localization model, which builds a powerful audio localization system through negative sampling and soft labels. Negative sampling helps the model distinguish positive and negative samples, while soft labels provide more fine-grained alignment information. The goal of WS-TAG is to improve the performance of the model in the audio localization task in a weakly-supervised manner. However, although the weakly-supervised model WS-TAG can provide certain alignment information by using negative sampling and soft labels to build the audio localization system, the weak supervision signal itself lacks accurate frame-level annotation, making it difficult for the model to achieve high-precision alignment.

[0027] In MGA-CLAP, a modality-shared codebook and local perception Transformer blocks are adopted to improve contrastive pre-training. MGA-CLAP is an improved contrastive language-audio pre-training model, which enhances the effect of contrastive pre-training by introducing a modality-shared codebook and local perception Transformer blocks. The goal of MGA-CLAP is to improve the performance of the audio-language model through multi-granularity alignment (including frame-level and word-level alignment). MGA-CLAP learns the global alignment between audio and text through contrastive pre-training, but its alignment granularity is still limited by the framework and cannot achieve frame-level precise alignment.

[0028] However, although these techniques (such as UACA, WS-TAG, and MGA-CLAP) can learn the alignment relationship between audio and text, their alignment methods are relatively rough and it is difficult to accurately capture the fine-grained correspondence between audio frames and text phrases. This results in limited performance in tasks that require high-precision localization. In addition, some techniques (such as MGA-CLAP and CRNN-BERT) improve the model performance by introducing additional modules (such as modality-sharing codebooks, local perception Transformer blocks, and BERT encoders), but these modules increase the computational complexity of the model, leading to higher training and inference costs. Some techniques (such as TAG-Baseline) use relatively simple text embedding methods (such as Word2Vec), which can capture certain semantic information, but their representation ability is relatively weak and it is difficult to handle complex text descriptions.

[0029] Figure 2 FIG. shows an operation flowchart of an example of a training method of an audio localization model according to an embodiment of the present application.

[0030] As Figure 2 shown, in step S210, an audio-caption sample is obtained, and the audio-caption sample includes an audio segment and a caption segment that are aligned on the time axis.

[0031] In some embodiments, the audio-caption sample can be extracted from a multi-modal dataset. The audio sample can come from various audio scenarios (such as ambient sound, music, speech, etc.), and the caption sample is a text description of the audio segment. The two need to be strictly aligned on the time axis to ensure that each audio segment can accurately correspond to the corresponding caption segment.

[0032] It should be understood that in CLAP training, the core idea is to enable the model to learn the association between audio and text through contrastive learning of positive and negative samples. In some examples, the audio-caption sample includes positive sample pairs and negative sample pairs. In the positive sample pairs, the audio segment and the caption segment are semantically matched, and in the negative sample pairs, the audio segment and the caption segment are semantically unmatched.

[0033] Specifically, each positive sample pair consists of an audio segment and its caption segment that is aligned on the time axis and semantically matched. The caption describes the sound events, environmental features, or speech content contained in the audio segment, ensuring that the positive sample pairs are highly relevant semantically. In addition, to enhance the discriminative ability of the model, a certain number of negative sample pairs need to be constructed. For example, the negative sample pairs can be constructed by random pairing, similar interference, or time misalignment, etc.

[0034] In step S220, based on the audio-caption sample and the contrastive loss function, the audio localization model is trained by CLAP.

[0035] The audio encoder and the text encoder are jointly trained using the Contrastive Loss function, aiming to map semantically related audio and text embeddings into adjacent feature spaces.

[0036] Exemplarily, the audio localization model is CLAP-trained by minimizing the contrastive loss to shorten the distance between the audio features and text embeddings in the positive sample pairs and widen the distance between the audio features and text embeddings in the negative sample pairs. Specifically, the feature similarity of the positive sample pairs can be maximized and the feature similarity of the negative sample pairs can be minimized through the contrastive loss function, and the model gradually learns the cross-modal semantic association between audio and text. Thus, through CLAP training, better segment-level alignment can be achieved.

[0037] In step S230, frame-level audio features of the audio segment are extracted based on the audio adapter.

[0038] Here, the Audio Adapter, as a lightweight module, can be embedded into the audio encoder to enhance its ability to perceive temporal details. In addition, the structural type of the audio adapter can be diversified, such as based on deep learning structures like Convolutional Neural Network (CNN), Temporal Convolutional Network (TCN), or Transformer, etc., to extract frame-by-frame features of the audio signal. By introducing the audio adapter, the model can further capture the detailed information of the audio signal in a finer-grained time dimension on the basis of the original segment-level feature extraction ability of the audio encoder, thereby effectively improving the model's perception accuracy of the start and end times of sound events.

[0039] In step S240, the frame-level audio-phrase similarity between the frame-level audio features of each frame and the corresponding phrase embeddings is calculated.

[0040] Here, the frame-level audio features extracted by the trained audio encoder are used to calculate the frame-by-frame similarity with the phrase embeddings generated by the text encoder. In addition, the similarity measurement method can be diversified, such as cosine similarity, dot product similarity, etc. In some embodiments, in order to further improve the accuracy of similarity matching, a sliding window mechanism can also be introduced to establish a temporal smoothing relationship between consecutive audio frames to mitigate the error caused by frame-level feature fluctuations.

[0041] By calculating the similarity frame by frame, each time point in the audio signal is aligned and matched with the specific events described in the caption text, significantly improving the model's prediction accuracy of the time boundaries of sound events.

[0042] In step S250, based on the frame-level audio-phrase similarity and the sound event label corresponding to each frame, the audio localization model is trained for sound event classification. The sound event label is used to indicate whether the audio frame matches the real sound event described by the phrase embedding.

[0043] Here, the frame-level audio-phrase similarity is used as the key feature index of the model, and at the same time, the sound event label is introduced as a supervision signal to further optimize the audio localization model. The sound event label can be in the form of binary classification (match / mismatch) or multi-classification (multiple event types). During the training process, the Cross-Entropy Loss or Binary Logarithmic Loss function can be used to maximize the matching probability of the model for real event frames and minimize the misjudgment probability for irrelevant event frames. Thus, combined with the supervised learning strategy of the sound event label, the model can more accurately identify the specific start and end times of sound events in complex and noisy environments, significantly improving the performance of the audio localization model in identifying sound events in audio and the corresponding event start and end times in the SED task application, as well as predicting the event start and end times of the corresponding sound events in the audio according to the natural language description in the TAG task application. Details of more task applications and performance experiment verifications will be presented in combination with other examples below.

[0044] In some examples of the embodiments of the present application, the audio localization model is trained by minimizing the binary cross-entropy loss between the frame-level audio-phrase similarity and the corresponding sound event label.

[0045] Thus, by introducing the Binary Cross-Entropy (BCE) loss into the frame-level optimization strategy, the time boundary recognition accuracy of the model in short-term sudden events (such as gunshots, glass breaking sounds) and continuous events (such as background music, rain sounds) is effectively improved. In diverse scenarios of audio data (such as far-field sound pickup, noise interference, multi-source mixing, etc.), the BCE loss guides the model to focus on the key sound events indicated in the phrase description, avoiding performance degradation due to non-target sound source interference. In addition, due to the large number of event-free frames in the audio, the weighted strategy of the BCE loss is particularly effective in sparse label scenarios, significantly reducing the false detection rate of the model in negative samples.

[0046] It should be noted that the traditional native CLAP model mainly focuses on segment-level tasks and is difficult to capture the fine-grained temporal information in audio signals. Through the F-CLAP model provided by the embodiments of the present application, an audio adapter is introduced to extract the frame-level features of audio segments. On the basis of CLAP training, frame-level similarity calculation is introduced, and the original segment-level representation of the CLAP model is further refined to the frame level, realizing more accurate audio-text feature alignment, effectively alleviating problems such as feature ambiguity and temporal boundary ambiguity existing in the CLAP model in frame-level audio understanding tasks, and improving the event localization ability and temporal boundary prediction accuracy of the model. Thereby, the task performance of the audio localization model in SED task applications and TAG task applications is effectively improved.

[0047] In addition, according to the research directions of current related technologies, when solving defects such as limited alignment accuracy, high computational complexity, and limited text representation ability, the industry generally adopts the following methods: for the alignment accuracy problem, more complex alignment mechanisms (such as multi-modal attention mechanisms) may be tried or more supervision signals (such as using stronger annotation data) may be added; for the computational complexity problem, the computational cost may be reduced by model compression, knowledge distillation, or using more efficient hardware; for the text representation ability problem, more powerful pre-trained language models (such as BERT or GPT) may be adopted to enhance the representation ability of text features. However, these current research directions often make improvements within the existing framework and are difficult to break out of the traditional contrastive pre-training or strong supervised learning paradigm.

[0048] Through the F-CLAP provided by the embodiments of the present application, a brand-new idea is adopted: by introducing a lightweight audio adapter into the pre-trained CLAP model, both the global alignment ability of the model is retained and the frame-level fine-grained alignment is achieved. It not only avoids complex model modifications, but also significantly reduces the computational cost, while making full use of the powerful representation ability of the pre-trained model.

[0049] Regarding the implementation details of step S230, in some embodiments, at least one model component in the audio localization model is frozen, and frame-level audio features of the audio segment are extracted based on the audio adapter. The model component includes any one of the following: audio encoder, audio global projector, text encoder, and text global projector.

[0050] In this way, after the CLAP pre-training is completed, the core components in the audio localization model are frozen to avoid damaging the learned multi-modal alignment information in the subsequent frame-level feature extraction tasks. In some embodiments, to ensure that the model obtains both segment-level and frame-level features in a single forward pass, the global features output by the audio encoder are used as the segment-level representation and directly applied to coarse-grained tasks (such as audio classification, audio-text retrieval, etc.), while the frame-level features extracted by the audio adapter are used as the fine-grained representation and further applied to frame-level tasks (such as sound event detection, text-to-audio localization, etc.). Through this joint output mechanism, the additional overhead of separately training the frame-level branch for the model is avoided, thereby improving the inference efficiency of the model.

[0051] Preferably, after the CLAP pre-training, all model components in the audio localization model are frozen, and only the lightweight audio adapter is updated, thus significantly improving the training efficiency. In addition, the original global alignment of CLAP is retained, enabling the model to generate both segment-level and frame-level features in a single forward pass, thereby taking into account both coarse-grained and fine-grained tasks.

[0052] Figure 3 A structural block diagram showing an example of an audio adapter according to an embodiment of the present application is shown.

[0053] As Figure 3 shown, the audio adapter 300 includes a cascaded upsampling module 310 and a context analysis module 320.

[0054] The upsampling module 310 is used to upsample the low-resolution frame-level features output by the audio encoder of the audio localization model to generate high-resolution frame-level audio features, so as to more finely capture the temporal details of the audio signal.

[0055] Specifically, in the CLAP model, the audio encoder usually outputs low-resolution frame-level features (for example, one feature vector corresponds to every 10 ms or longer time step), which may make it difficult to accurately locate the time boundaries of sound events in frame-level tasks. To solve this problem, the upsampling module in the audio adapter can use methods such as transposed convolution and interpolation to improve the temporal resolution of the features.

[0056] The context analysis module 320 is used to capture and model the temporal correlations between consecutive frame-level audio features, so that the model can understand the temporal structure and event boundaries in the audio signal.

[0057] It should be noted that audio signals usually have strong context correlations in the time dimension. For example, some sound events may be continuous (such as wind sounds, rain sounds, background noise) or have clear start and end boundaries (such as gunshots, clapping sounds). To effectively model this time dependence, the context analysis module can model the long-term or short-term dependence relationships in the audio signal in various ways, such as through Transformer or bidirectional gated recurrent units, etc.

[0058] In some examples of the embodiments of the present application, the context analysis module adopts a bidirectional GRU or a multi-layer perceptron, and the multi-layer perceptron includes a cascaded first linear layer, a ReLU activation layer, and a second linear layer. By introducing the bidirectional GRU, it aims to capture the bidirectional dependence relationships of the audio signal in the time dimension, thereby enhancing the model's understanding of the start and end time boundaries of sound events and complex temporal patterns. In addition, by introducing the multi-layer perceptron, the model's non-linear mapping ability for complex audio features is enhanced, so as to show better feature separation ability in complex scenarios such as mixed audio and noise interference.

[0059] In the embodiments of the present application, by introducing an audio adapter composed of an upsampling module and a context analysis module, the model's frame-level understanding ability of audio signals is significantly enhanced. The upsampling module effectively improves the time resolution, making the model perform better when capturing short-term and sudden sound events (such as knocking sounds, gunshots); the context analysis module further strengthens the model's ability to model long-term dependence information in complex audio scenarios, and improves the model's understanding accuracy of continuous sound events (such as rain sounds, wind sounds).

[0060] Furthermore, since the audio adapter adopts a lightweight structure and combines the CLAP model component freezing strategy, it greatly reduces the scale of parameter updates during the training process, significantly improves the training efficiency and resource utilization rate of the model. In addition, it also ensures that the original segment-level feature alignment ability of the CLAP model is retained, enabling the model to complete segment-level tasks (such as audio-text retrieval) and frame-level tasks (such as sound event detection, text-to-audio localization) in a single forward pass, improving the application breadth and practicality of the audio alignment model.

[0061] With the F-CLAP provided by the embodiments of the present application, on the basis of freezing the pre-trained CLAP model, a lightweight audio adapter is introduced to extract high-resolution audio embeddings, and these embeddings are aligned with text phrase representations through frame-level binary cross-entropy loss. This method achieves frame-level fine-grained alignment without modifying the original CLAP model, significantly improving the alignment accuracy. At the same time, since only the lightweight adapter is trained, the computational complexity is greatly reduced, avoiding the high computational costs brought by introducing additional modules or large-scale pre-trained models in traditional methods. In addition, F-CLAP makes full use of the powerful text representation ability of the CLAP model, avoiding the limitations of simple text embedding methods (such as Word2Vec), and being able to better handle complex text descriptions. Through this design, F-CLAP achieves precise frame-level audio understanding while maintaining high computational efficiency and powerful text representation ability.

[0062] It should be noted that since the contrastive pre-training framework of CLAP mainly focuses on the coarse-grained alignment between the entire audio segment and the complete caption sentence, this embodiment proposes to use a fine-grained adapter on top of the frozen CLAP audio encoder to optimize the original global alignment and achieve the phrase-frame correspondence.

[0063] During the inventors' practice of the present application, other alternative solutions were also proposed and compared with the solution provided by the embodiments of the present application.

[0064] An alternative solution is a fine-grained alignment model based on the multimodal attention mechanism. The core idea of this solution is to directly establish a fine-grained alignment relationship between audio and text by introducing the multimodal attention mechanism. Specifically, the model calculates the attention weights between audio frames and text phrases to dynamically capture the correspondence between the two. The advantage of this method lies in its flexibility and powerful alignment ability, which can achieve frame-level fine-grained alignment without relying on additional labeled data. However, the disadvantages of this solution are also very obvious: First, the computational complexity of the multimodal attention mechanism is relatively high, especially when dealing with long audio and complex text, the computational resource requirements of the model will increase significantly; Second, since the attention mechanism depends on the self-learning ability of the model, its alignment accuracy may be limited by the quality and scale of the training data. Especially in the case of scarce data, the performance of the model may be greatly reduced.

[0065] Another alternative is an audio-text alignment model based on Generative Adversarial Networks (GAN). This solution utilizes the framework of GAN to achieve fine-grained alignment between audio and text through adversarial learning of a generator and a discriminator. The generator is responsible for generating the correspondence between audio frames and text phrases, while the discriminator is responsible for judging whether the generated correspondence is real. The advantage of this method lies in its powerful generation ability and the potential of adversarial learning, which can learn complex alignment relationships between audio and text under unsupervised or weakly supervised conditions. However, the disadvantage of this solution is that its training process is relatively complex, the training stability of the GAN is poor, and problems such as mode collapse or non-convergence of training are likely to occur; in addition, the training of the GAN requires a large amount of computing resources, and the inference speed of the model is slow, making it difficult to achieve efficient real-time processing in practical applications.

[0066] Finally, the F-CLAP solution was selected because it achieved precise frame-level audio understanding while maintaining efficient computing and powerful text representation capabilities, overcoming the deficiencies of the above alternatives. By introducing a lightweight audio adapter and frame-level binary cross-entropy loss, F-CLAP achieved frame-level fine-grained alignment without modifying the original CLAP model, significantly improving the alignment accuracy while reducing the computational complexity. This design not only retained the global alignment ability of the CLAP model but also made full use of the powerful representation ability of the pre-trained model, enabling F-CLAP to perform well in text-to-audio localization and sound event detection tasks.

[0067] In addition, during the inventors' practice of this application, multiple feasible beta versions were considered, and F-CLAP was finally confirmed and selected.

[0068] One beta version solution is a preliminary model based on frame-level contrastive learning. In this version, an attempt was made to directly introduce frame-level contrastive learning on the basis of the CLAP model to achieve fine-grained alignment by comparing the embeddings of audio frames and text phrases. Specifically, first, the audio is segmented at the frame level, then the frame-level features are extracted using the CLAP audio encoder, and then contrastive learning is performed with the embeddings of text phrases. The advantage of this solution is its simplicity and directness, which can achieve frame-level alignment to a certain extent and does not require the introduction of additional complex modules. However, the disadvantages of this beta version are also very obvious: First, the computational cost of frame-level contrastive learning is relatively high, especially when dealing with long audio, the computational resource requirements of the model increase significantly; second, since the original design of the CLAP model is for global alignment, directly introducing frame-level contrastive learning will make it difficult for the model to balance between global alignment and frame-level alignment, thus affecting the overall performance.

[0069] Another beta version solution is a frame-level alignment model based on Temporal Convolutional Network (TCN). In this version, a temporal convolutional network is introduced based on the CLAP model, and the TCN is used to model the temporal sequence of audio frames, so as to achieve fine-grained alignment at the frame level. Specifically, first, the audio encoder of CLAP is used to extract audio features, then the TCN is used to model the temporal sequence of these features, and finally, the output of the TCN is aligned with the embedding of the text phrase. The advantage of this solution is that the TCN can effectively capture the temporal relationship between audio frames, thus improving the alignment accuracy; in addition, the TCN has high computational efficiency and can achieve frame-level alignment at a relatively low computational cost. However, the disadvantage of this beta version is that its alignment accuracy is still limited by the original design of the CLAP model and cannot fully achieve precise frame-level alignment; in addition, the introduction of the TCN increases the complexity of the model and is prone to problems such as vanishing gradients or exploding gradients during training.

[0070] Finally, F-CLAP was selected because it achieved precise frame-level audio understanding while maintaining high computational efficiency and strong text representation capabilities, overcoming the defects of the above beta version solutions. By introducing a lightweight audio adapter and frame-level binary cross-entropy loss, F-CLAP achieved fine-grained alignment at the frame level without modifying the original CLAP model, significantly improving the alignment accuracy while reducing the computational complexity. This design not only retains the global alignment ability of the CLAP model but also makes full use of the powerful representation capabilities of the pre-trained model, enabling F-CLAP to perform well in text-to-audio localization and sound event detection tasks.

[0071] Through the F-CLAP solution provided by the embodiments of this application, it can not only directly improve the performance of audio-language models in frame-level tasks but also trigger a series of deep chain reactions and broader impacts. First, in terms of direct effects, by introducing a lightweight audio adapter and frame-level binary cross-entropy loss, F-CLAP significantly improved the model's performance in text-to-audio localization (TAG) and sound event detection (SED) tasks. Specifically, F-CLAP can accurately locate the start and end times of sound events in audio while maintaining the ability to understand the global audio content. This ability makes the model more robust in dealing with complex audio scenarios and can better meet the requirements of high-precision audio analysis in practical applications.

[0072] At a deeper level of impact, the success of F-CLAP has opened up new research directions for the fine-grained understanding of audio-language models. First, it has demonstrated the feasibility of achieving frame-level alignment through lightweight adapters without disrupting the global alignment ability of pre-trained models. This discovery provides new ideas for other multimodal tasks such as video-text alignment or image-text alignment, suggesting that fine-grained understanding can be achieved through similar adapter mechanisms without having to train complex models from scratch. Second, the efficient design of F-CLAP (training only lightweight adapters) reduces computational costs, making it easier to deploy fine-grained audio understanding technologies to resource-constrained devices such as mobile devices or embedded systems, thus promoting the wide application of audio analysis technologies in fields such as smart homes, autonomous driving, and healthcare.

[0073] In addition, the success of F-CLAP has also triggered a rethinking of the pre-training paradigm for audio-language models. Traditional contrastive pre-training mainly focuses on global alignment, while F-CLAP has shown how to achieve fine-grained alignment while maintaining global alignment. This breakthrough provides new directions for the design of future multimodal pre-training models and may give rise to more general models that can handle both global and local alignment tasks simultaneously. The potential of such models is not limited to the audio domain but may also extend to fine-grained understanding tasks in other modalities such as video and image.

[0074] Finally, the frame-level understanding ability of F-CLAP offers new possibilities for the generation and editing of audio content. For example, in audio generation tasks, F-CLAP can be used to generate audio segments that are precisely aligned with text descriptions; in audio editing tasks, it can be used to automatically identify and edit specific events in audio. These applications not only improve the efficiency of audio generation and editing but also bring new tools and methods to the creative industries such as music production and film dubbing.

[0075] In summary, F-CLAP not only directly improves the performance of audio-language models in frame-level tasks but also triggers a chain reaction in multiple fields such as multimodal research, computational efficiency improvement, pre-training paradigm innovation, and audio generation and editing through its innovative design, demonstrating its profound impact in academic research and practical applications.

[0076] The following will expand on more specific implementation details of F-CLAP, which endows audio-language models with frame-level understanding ability, in combination with other examples.

[0077] Based on F-CLAP, a simple and effective method is proposed to enable CLAP to locate the start and end times of sound events while retaining its original capabilities in segment-level audio understanding tasks. A lightweight adapter is trained on top of the frozen CLAP encoder to extract high-resolution audio embeddings and align them with phrase representations using frame-level binary cross-entropy loss. Experimental results show that the method in this paper significantly improves the performance of CLAP in sound event detection and achieves state-of-the-art results in audio localization tasks, demonstrating the potential of using CLAP for fine-grained audio understanding.

[0078] 1. Introduction

[0079] Specifically, a lightweight audio adapter is trained on top of the frozen CLAP encoder to extract high-resolution audio embeddings. These fine-grained audio embeddings are then aligned with phrase representations from the initial text branch using frame-level binary cross-entropy (BCE) loss. Additionally, since the weights of the backbone CLAP encoder are not modified, the model is able to retain its original global alignment and generate segment-level and frame-level features in a single forward pass. Experimental results show that the method in this paper significantly improves the performance of CLAP in sound event detection and reaches the state-of-the-art level in text-to-audio localization tasks.

[0080] 2. Frame-Level Audio Understanding

[0081] Most SED systems mainly consist of two core components: a pre-trained audio encoder for extracting high-dimensional audio features and a context network for generating frame-by-frame predictions. For example, the ATST encoder is jointly fine-tuned with a bidirectional gated recurrent unit (BiGRU) to predict frame-level labels; in addition, the PaSST encoder and a Transformer network are used for sound event detection. Although significant progress has been made in the SED field, current models are still limited to specific predefined categories, which poses significant challenges in domain transfer and thus limits their practical application value.

[0082] Text-to-Audio Grounding (TAG) aims to predict the start and end times of sound events based on natural language descriptions and can be regarded as an open-vocabulary version of the sound event detection task. Existing TAG systems usually adopt a dual-encoder structure to extract audio and text representations for subsequent alignment. However, due to the limited amount of strongly annotated data and the modality gap between audio and text, the performance of these models is still not satisfactory.

[0083] In contrast, the method in this paper processes frame-level audio understanding tasks by learning fine-grained audio-text alignment and leveraging the CLAP model. Compared with closed-set methods, it offers greater flexibility and can benefit from the extensive multimodal knowledge obtained through large-scale contrastive pre-training.

[0084] 3. Method

[0085] Figure 4 Figure 1 shows a schematic diagram of the architecture comparison between the CLAP training paradigm and F-CLAP training.

[0086] As Figure 4 shown on the left side of Figure 1, it shows an overview of the CLAP training paradigm. As Figure 4 shown on the right side of Figure 1, it shows an overview of F-CLAP, demonstrating how to endow CLAP with fine-grained audio understanding capabilities. Since the contrastive pre-training framework of CLAP mainly focuses on the coarse-grained alignment between the entire audio segment and the entire sentence caption, a fine-grained adapter is proposed to be added above the frozen CLAP audio encoder to refine the original global alignment and achieve phrase-frame correspondence.

[0087] 3.1. Contrastive Language-Audio Pre-training

[0088] Consider a set of B audio-caption samples where A i represents the i-th audio segment, and its corresponding caption is The audio encoder f of CLAP a and the text encoder f t first extract frame-level embeddings and word-level representations respectively:

[0089]

[0090] where L, N, and D represent the number of frames, the number of words, and the feature dimension respectively. Subsequently, these local representations are aggregated through global audio and text projectors to obtain segment-level and sentence-level embeddings:

[0091]

[0092] Finally, by minimizing the contrastive loss, the paired audio and text embeddings are made as close as possible, while the unpaired embeddings are pulled apart:

[0093]

[0094] where, represents the cosine similarity between the global embeddings of the i-th audio segment and the j-th caption segment. As described in Section 1, although this process has achieved segment-level alignment, the model still lacks the ability to associate corresponding frames with text descriptions.

[0095] 3.2. Learning Phrase-Frame Correspondence

[0096] Furthermore, a lightweight audio adapter φ is trained on top of the frozen CLAP audio encoder a :R L×D →R L×D to achieve the correspondence between phrases and frames. Let

[0097]

[0098] denote the extracted high-resolution audio embeddings (the subscript i is omitted here for simplicity). Then, the similarity between E a and the original phrase embeddings in the CLAP text branch is obtained by calculation:

[0099]

[0100] where is the Sigmoid activation function, denotes the transpose of E a . The model is trained by minimizing the binary cross-entropy loss between the frame-level audio-phrase similarity and the ground truth labels:

[0101]

[0102] where represents the similarity between the l-th frame embedding of audio A and the text features of phrase , and y l ∈{0,1} is the ground truth label of the sound event, indicating whether the sound event described by exists in this frame.

[0103] With this objective, the embeddings of the audio at the corresponding timestamps are aligned with the paired phrase representations, enabling the model to detect the start and end times of the sound event. In fact, the audio adapter φ a contains an upsampling layer to achieve the required resolution for the task, and is equipped with a context network to model the temporal relationships between frames.

[0104] During training, the audio encoder f a , text encoder f t in CLAP, as well as the original global projectors g a and g t are all kept frozen, and only the lightweight audio adapter φ a is updated, thus significantly improving the training efficiency. In addition, the original global alignment in CLAP is retained, enabling the model to generate both segment-level and frame-level features in a single forward pass, thus taking into account both coarse-grained and fine-grained tasks.

[0105] 4. Experimental Setup

[0106] 4.1. Datasets

[0107] To evaluate the effectiveness of the proposed method, experiments were conducted on two widely used benchmark datasets for frame-level audio understanding tasks.

[0108] In the Text-to-Audio Grounding (TAG) task, the AudioGrounding dataset was used. This dataset contains 4,657 audio clips obtained from AudioCaps and 13,958 sound event phrases extracted from the captions. Each phrase was annotated with precise start and end timestamps by experienced human annotators. The standard data split of Version-2.21 was adopted to train and evaluate the model.

[0109] In the Sound Event Detection (SED) task, the DESED dataset was used. Only the strongly annotated subset of DESED was used, which contains 3,470 audio clips, each with frame-level labels for 10 indoor sound events. The validation subset of DESED contains 1,168 audio clips for model evaluation. To avoid excessive prompt engineering, the sound category names were directly used as phrase inputs.

[0110] 4.2. Evaluation Metrics

[0111] In the TAG task, PSDS and Th-AUC were used to evaluate the performance of the proposed method. The Multi-Channel Sound Detection Score (PSDS) reflects the relationship between the True Positive Rate (TPR) and the False Positive Rate (FPR) by quantifying the area under the PSD-ROC curve, while Th-AUC captures the average performance at all possible thresholds by measuring the area under the F1-threshold curve.

[0112] Similarly, in the SED task, PSDS1 and PSDS2 were adopted as evaluation metrics. PSDS1 focuses more on the accuracy of event localization, while PSDS2 emphasizes more on the discriminability between classes and less on event localization. All scores were calculated using the toolset sed scores eval2.

[0113] 4.3. Implementation Details

[0114] Use the CLAP model, where the audio encoder is HTS-AT and the text encoder is RoBERTa. This CLAP model was pre-trained on a total of 450k audio-caption pairs, which were sourced from WavCaps, AudioCaps, and Clotho. For the fine-grained audio adapter, two BiGRU layers were selected as the context network because they performed well in previous frame-level audio understanding models.

[0115] Train the fine-grained audio adapter for 30 epochs on the AudioGrounding dataset and 15 epochs on the DESED dataset. Use the Adam optimizer with a batch size of 32. During training, a cosine annealing schedule is adopted, with a maximum learning rate of 5e-5 and a warm-up phase for the first 10% of the epochs. The model is evaluated using the weights of the last epoch.

[0116] Table 1: Performance comparison of weakly-supervised and strongly-supervised TAG models on the test set (TAG-Eval) of the AudioGrounding dataset. Higher values indicate better performance. The optimal scores are shown in bold.

[0117]

[0118] Table 2: Performance comparison of closed-vocabulary SED systems and open-vocabulary SED models on the validation set (DESED-Val) of the DESED dataset. Higher values indicate better performance, and the methods using additional data for fine-tuning are shown in gray.

[0119]

[0120] 4.4. Baselines

[0121] In the text-to-audio localization task, the system in this paper was compared with weakly-supervised TAG models and strongly-supervised TAG models. The weakly-supervised TAG models were trained on large-scale audio-caption pairs without using frame-level annotation data. Specifically, UACA proposed aggregating the frame-word alignment matrix through mean and max pooling to learn the detailed alignment between audio and text; WS-TAG used negative sampling and soft labels to build a powerful audio localization system; MGA-CLAP adopted a modality-shared codebook and a local perception Transformer module to improve contrastive pre-training. In the strongly-supervised methods, TAG-Baseline used a convolutional recurrent neural network (CRNN) to associate audio embeddings with Word2Vec text embeddings, and CRNN-BERT further improved the model performance using a frozen BERT encoder.

[0122] In the sound event detection task, the closed-set method refers to systems that are specifically trained for the fixed classes of the DESED dataset, while the open-vocabulary method supports natural language as queries. For example, BEATs-CRNN is the baseline system for task4 of the DCASE2023 Challenge, which combines the state-of-the-art audio encoder BEATs with CRNN to generate frame-by-frame labels. ATST-SED proposed an effective fine-tuning paradigm to adapt the ATST encoder to the SED task and achieved state-of-the-art performance on the DESED dataset. Since these methods undergo carefully designed multi-stage training and involve additional data in fine-tuning, we only show their results in grey for reference.

[0123] Since there is little research directly addressing the open-vocabulary SED problem, the performance of the TAG system is also reported for comparison.

[0124] Table 3: Comparison of results when different components of the frozen model. EncT, EncA, and ProjT represent the audio encoder, text encoder, and text projector respectively. "Flame" indicates that the module is trainable, while "Snowflake" indicates that the module is frozen.

[0125]

[0126] 5. Experimental Results

[0127] 5.1. Main Results

[0128] CLAP with fine-grained adapters is called F-CLAP. As shown in Table 1 and Table 2, F-CLAP achieved the latest state-of-the-art results in the text-to-audio localization task and significantly improved the performance of the original CLAP model in the SED task, although it only requires very few parameters for training. This indicates that using CLAP for fine-grained audio understanding has potential and demonstrates the advantages of the multi-modal latent space in frame-level tasks. However, the system in this paper still lags behind the state-of-the-art closed-set SED models in the SED task, which may be due to the limitation of the labeled training data volume. In addition, different from the text descriptions provided in the TAG task, the simple sound category names in the SED task may not provide sufficient semantic information for the CLAP encoder.

[0129] 5.2. Ablation Experiments

[0130] Here, comprehensive ablation experiments were conducted to evaluate the impact of each design in the method.

[0131] Parameter Freezing

[0132] First, the impact of fine-tuning the CLAP backbone was studied. As shown in Table 3, setting the audio encoder to be trainable can significantly improve the model's performance in audio localization tasks, while fine-tuning the text branch only brings a slight gain to the Th-AUC score. However, for sound event detection, fine-tuning other components except the audio adapter will reduce the model's performance. A possible reason is that the scarcity of labeled data can lead to overfitting, so increasing the amount of training data or designing a better fine-tuning paradigm may be helpful. The method in this paper only needs to train the audio adapter, significantly improving the training efficiency, retaining the original capabilities of CLAP, while the performance can be comparable to the case of full fine-tuning.

[0133] Table 4: Ablation study of model structure.

[0134]

[0135] A system using a linear layer as the audio adapter.

[0136] A system using a stronger CLAP model pre-trained on 2 million pairs of audio-caption.

[0137] *: Systems where the audio encoder and text encoder are not contrastively pre-trained.

[0138] Fine-grained audio adapter

[0139] Here, the BiGRU was replaced by two linear layers and an intermediate ReLU activation to study its effect. As shown in Table 4, the recurrent neural network shows better results than the linear structure in all metrics, especially in the SED task, demonstrating its strong ability in modeling temporal dependencies.

[0140] Contrastive pre-training

[0141] Here, the impact of contrastive language-audio pre-training on fine-grained audio understanding tasks was also explored. First, an evaluation was conducted based on the CLAP model, which was pre-trained on 2 million pairs of audio-caption. This model is called CLAP-2M and achieved the best results in audio-text retrieval tasks. The fine-grained audio adapter of CLAP-2M was re-trained on the above benchmark datasets. As shown in Table 4, the system using CLAP-2M shows consistent improvements in most metrics, highlighting the advantages of a stronger audio-language model for frame-level downstream tasks.

[0142] In addition, the performance of using the method on encoders not aligned by contrastive learning was also evaluated. Specifically, an audio adapter and a text projector were trained on top of the frozen HTS-AT and RoBERTa encoders. As shown in Table 4, the system without large-scale audio-text pre-training suffered a significant performance drop on all metrics, indicating the importance and advantages of using the latent space of audio-language models for fine-grained audio understanding.

[0143] 6. Conclusions and Future Work

[0144] This paper proposes a simple yet effective method to endow CLAP with frame-level audio understanding capabilities. A lightweight audio adapter is trained on top of the CLAP encoder to extract fine-grained audio embeddings, and these embeddings are aligned with phrase representations via frame-level binary cross-entropy loss. By freezing the CLAP backbone model, not only is efficiency achieved, but the original capabilities of the original audio-language model are also retained.

[0145] The experimental results show that the method in this paper achieves state-of-the-art performance in the text-to-audio localization task and significantly improves the sound event detection performance of CLAP, demonstrating the potential of leveraging CLAP for fine-grained audio understanding. Future research will focus on exploring improved audio adapter architectures, training paradigms, and advanced negative sampling techniques to further enhance the audio-language model's ability to understand temporal characteristics.

[0146] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of actions combined. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application. In the above embodiments, each embodiment is described with emphasis. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0147] In some embodiments, the embodiments of this application provide a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored. The execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to be used for executing any one of the above audio localization model training methods of this application.

[0148] In some embodiments, the embodiments of the present application further provide a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions that, when executed by a computer, cause the computer to execute the training method of any one of the above audio localization models.

[0149] In some embodiments, the embodiments of the present application further provide an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor. Wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the training method of the audio localization model.

[0150] Figure 5 is a schematic hardware structure diagram of an electronic device for executing the training method of the audio localization model provided by another embodiment of the present application, as Figure 5 shown, the device includes:

[0151] One or more processors 510 and a memory 520, Figure 5 Taking one processor 510 as an example.

[0152] The device for executing the training method of the audio localization model may further include: an input device 530 and an output device 540.

[0153] The processor 510, the memory 520, the input device 530 and the output device 540 may be connected through a bus or other means, Figure 5 Taking the connection through a bus as an example.

[0154] The memory 520, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs and modules, such as the program instructions / modules corresponding to the training method of the audio localization model in the embodiments of the present application. The processor 510 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions and modules stored in the memory 520, that is, implements the training method of the audio localization model in the above method embodiments.

[0155] The memory 520 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created according to the use of the electronic device and the like. In addition, the memory 520 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 520 may optionally include a memory remotely provided with respect to the processor 510, and these remote memories may be connected to the electronic device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0156] The input device 530 may receive input digital or character information and generate signals related to user settings and function control of the electronic device. The output device 540 may include a display device such as a display screen.

[0157] The one or more modules are stored in the memory 520 and, when executed by the one or more processors 510, execute the training method of the audio localization model in any of the above method embodiments.

[0158] The above product may execute the method provided in the embodiments of the present application and has functional modules and beneficial effects corresponding to the executed method. Technical details not described in detail in this embodiment may be referred to the method provided in the embodiments of the present application.

[0159] The electronic device in the embodiments of the present application exists in various forms, including but not limited to:

[0160] (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.

[0161] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, etc.

[0162] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and smart toys and portable vehicle navigation devices.

[0163] (4) Other airborne electronic devices with data interaction functions, such as in-vehicle device installed in a vehicle.

[0164] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0165] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0166] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A training method for an audio localization model, comprising: Obtaining audio-caption samples, where the audio-caption samples include audio segments and caption segments aligned on the timeline; Performing CLAP training on the audio localization model based on the audio-caption samples and a contrast loss function; Extracting frame-level audio features of the audio segments based on an audio adapter; Calculating the frame-level audio-phrase similarity between the frame-level audio features of each frame and the corresponding phrase embeddings; Performing sound event classification training on the audio localization model according to the frame-level audio-phrase similarity corresponding to each frame and a sound event label; the sound event label is used to indicate whether an audio frame matches the real sound event described by the phrase embedding.

2. The method according to claim 1, wherein, The extracting the frame-level audio features of the audio segments based on the audio adapter includes: Freezing at least one model component in the audio localization model and extracting the frame-level audio features of the audio segments based on the audio adapter; the model component includes any one of the following: an audio encoder, an audio global projector, a text encoder, and a text global projector.

3. The method according to claim 1, wherein The performing sound event classification training on the audio localization model according to the frame-level audio-phrase similarity corresponding to each frame and the sound event label includes: Training the audio localization model by minimizing the binary cross-entropy loss between the frame-level audio-phrase similarity and the corresponding sound event label.

4. The method according to claim 1, wherein The audio adapter includes an upsampling module and a context analysis module; The upsampling module is used to upsample the low-resolution frame-level features output by the audio encoder of the audio localization model to generate high-resolution frame-level audio features; The context analysis module is used to capture and model the temporal correlation between consecutive frame-level audio features.

5. The method according to claim 4, wherein, The context analysis module employs a bidirectional GRU or a multi-layer perceptron, and the multi-layer perceptron includes a cascaded first linear layer, a ReLU activation layer, and a second linear layer.

6. The method according to claim 1, wherein The audio-caption samples include positive sample pairs and negative sample pairs, where the audio segments and caption segments in the positive sample pairs are semantically matched, and the audio segments and caption segments in the negative sample pairs are not semantically matched; The performing CLAP training on the audio localization model based on the audio-caption samples and the contrast loss function includes: Performing CLAP training on the audio localization model by minimizing the contrast loss to shorten the distance between the audio features and the text embeddings in the positive sample pairs and widen the distance between the audio features and the text embeddings in the negative sample pairs.

7. The method according to any one of claims 1-6, wherein, The audio localization model is used to identify sound events and corresponding event start and end times in audio, or predict the event start and end times of corresponding sound events in audio according to a natural language description.

8. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1-7.

9. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the method according to any one of claims 1-7.

10. A computer program product comprising computer programs / instructions which, when executed by a processor, perform the steps of the method according to any one of claims 1-7.