Data processing method and apparatus, storage medium, device, and program product

CN122777752APending Publication Date: 2026-09-18BEIJING SOGOU TECHNOLOGY DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610913726.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0003]然而,传统的视觉-语言嵌入模型主要关注文本和图像的跨模态检索,对于音频模态的支持相对有限

Benefits of technology

[0011] This application embodiment receives query data; when the query data includes audio data, it uses an external audio encoder extracted from a full-modal large language model to convert the audio data into an audio feature sequence; it then maps the audio feature sequence to the latent space of the language backbone network using an audio projector to obtain an audio projection feature sequence; it determines an input sequence based on the query data, the input sequence containing at least one audio placeholder; it replaces the embedding of the audio placeholder in the input sequence with the corresponding feature vector in the audio projection feature sequence, inserts a start marker before the start position and an end marker after the end position of the audio projection feature sequence to obtain an extended sequence; it inputs the extended sequence into the language backbone network for encoding processing to obtain a query embedding vector corresponding to the query data; and it outputs the retrieval result based on the query embedding vector. This application embodiment, by utilizing an external audio encoder in a full-modal large language model, in conjunction with an audio projector, effectively converts audio data into an audio projection feature sequence that can be processed by the language backbone network, and by combining the replacement and insertion of audio placeholders, achieves efficient encoding processing of multimodal query data, obtains accurate query embedding vectors to output retrieval results, and significantly improves the efficiency and accuracy of multimodal data retrieval containing audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122777752A_ABST
    Figure CN122777752A_ABST
Patent Text Reader

Abstract

The application discloses a data processing method and device, a storage medium, equipment and a program product. The method comprises the following steps: when the query data contains audio data, converting the audio data into an audio feature sequence by using an external audio encoder; mapping the audio feature sequence to a hidden space of a language backbone network by using an audio projector to obtain an audio projection feature sequence; determining an input sequence according to the query data, wherein the input sequence contains at least one audio placeholder mark; replacing the embedding of the audio placeholder mark in the input sequence with a corresponding feature vector in the audio projection feature sequence, and inserting a start mark before the start position of the audio projection feature sequence and inserting an end mark after the end position to obtain an extended sequence; inputting the extended sequence into the language backbone network for encoding processing to obtain a query embedding vector corresponding to the query data; and outputting a retrieval result based on the query embedding vector, thereby improving the efficiency and accuracy of multi-modal data retrieval containing audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to a data processing method, apparatus, storage medium, device, and program product. Background Technology

[0002] With the rapid development of deep learning technology, cross-modal retrieval has become an important research direction in the field of information retrieval. Traditional information retrieval systems often only support single-modal queries, such as text retrieval or image retrieval, which cannot meet the increasingly diverse query needs of users. In order to achieve more efficient and comprehensive information retrieval, cross-modal retrieval systems have emerged. Their core objective is to map information from multiple modalities, such as text, images, videos, and audio, onto a unified semantic vector space to achieve joint retrieval of multimodal information.

[0003] However, traditional visual-language embedding models primarily focus on cross-modal retrieval of text and images, with relatively limited support for audio modalities. Some models attempt to achieve audio retrieval functionality by training audio encoders from scratch, but this approach requires not only a large amount of audio-text pairing data but also significant computational and time costs. Furthermore, audio encoders trained from scratch often struggle to reuse the capabilities of existing visual-language models, resulting in limited model performance.

[0004] Therefore, how to extend the capabilities of the original visual-language embedding model to support audio retrieval has become an urgent problem to be solved in the current cross-modal retrieval field. Summary of the Invention

[0005] This application provides a data processing method, apparatus, storage medium, device, and program product. By reusing the external audio encoder in the full-modal large language model and cooperating with the audio projector, combined with placeholder replacement and start / end marker insertion, it achieves efficient encoding processing of multimodal query data, significantly improving the efficiency and accuracy of multimodal data retrieval containing audio.

[0006] On one hand, embodiments of this application provide a data processing method, the method comprising: receiving query data containing audio data; converting the audio data into an audio feature sequence using an external audio encoder extracted from a full-modal large language model; mapping the audio feature sequence to the latent space of a language backbone network using an audio projector to obtain an audio projection feature sequence; determining an input sequence based on the query data, the input sequence containing at least one audio placeholder; replacing the embedding of the audio placeholder in the input sequence with the corresponding feature vector in the audio projection feature sequence, and inserting a start marker before the start position and an end marker after the end position of the audio projection feature sequence to obtain an extended sequence; inputting the extended sequence into the language backbone network for encoding processing to obtain a query embedding vector corresponding to the query data; and outputting retrieval results based on the query embedding vector.

[0007] On the other hand, embodiments of this application provide a data processing apparatus, the apparatus comprising: The receiving unit is used to receive query data; A conversion unit is used to convert audio into an audio feature sequence using an external audio encoder extracted from a full-modal large language model when the query data contains audio data. The mapping unit is used to map the audio feature sequence to the latent space of the language backbone network through the audio projector to obtain the audio projection feature sequence; A determining unit is configured to determine an input sequence based on the query data, the input sequence containing at least one audio placeholder; The processing unit is configured to replace the embedding of the audio placeholder in the input sequence with the corresponding feature vector in the audio projection feature sequence, and insert a start marker before the start position and an end marker after the end position of the audio projection feature sequence to obtain an extended sequence; An encoding unit is used to input the extended sequence into the language backbone network for encoding processing to obtain the query embedding vector corresponding to the query data. The retrieval unit is used to output retrieval results based on the query embedding vector.

[0008] On the other hand, an embodiment of this application provides a computer-readable storage medium storing a computer program adapted for loading by a processor to perform the data processing method as described in any of the above embodiments.

[0009] On the other hand, an embodiment of this application provides a computer device, which includes a processor and a memory. The memory stores a computer program, and the processor executes the data processing method described in any of the above embodiments by calling the computer program stored in the memory.

[0010] On the other hand, an embodiment of this application provides a computer program product, including computer instructions, which, when executed by a processor, implement the data processing method as described in any of the above embodiments.

[0011] This application embodiment receives query data; when the query data includes audio data, it uses an external audio encoder extracted from a full-modal large language model to convert the audio data into an audio feature sequence; it then maps the audio feature sequence to the latent space of the language backbone network using an audio projector to obtain an audio projection feature sequence; it determines an input sequence based on the query data, the input sequence containing at least one audio placeholder; it replaces the embedding of the audio placeholder in the input sequence with the corresponding feature vector in the audio projection feature sequence, inserts a start marker before the start position and an end marker after the end position of the audio projection feature sequence to obtain an extended sequence; it inputs the extended sequence into the language backbone network for encoding processing to obtain a query embedding vector corresponding to the query data; and it outputs the retrieval result based on the query embedding vector. This application embodiment, by utilizing an external audio encoder in a full-modal large language model, in conjunction with an audio projector, effectively converts audio data into an audio projection feature sequence that can be processed by the language backbone network, and by combining the replacement and insertion of audio placeholders, achieves efficient encoding processing of multimodal query data, obtains accurate query embedding vectors to output retrieval results, and significantly improves the efficiency and accuracy of multimodal data retrieval containing audio. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a schematic diagram illustrating an application scenario of the data processing system provided in an embodiment of this application.

[0014] Figure 2 A schematic diagram of the framework of the full-modal embedding model provided in the embodiments of this application.

[0015] Figure 3 This is a flowchart illustrating the data processing method provided in an embodiment of this application.

[0016] Figure 4 This is a schematic diagram of the inference stage of the data processing method provided in the embodiments of this application.

[0017] Figure 5 This is a schematic diagram illustrating an application scenario of the data processing method provided in the embodiments of this application.

[0018] Figure 6 This is a schematic diagram of the structure of the data processing apparatus provided in the embodiments of this application.

[0019] Figure 7 A schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0021] This application provides a data processing method, apparatus, storage medium, device, and program product. Exemplarily, the data processing method of this application can be executed by a computer device, which can be a terminal or a server, etc. The terminal can be a smartphone, tablet, laptop, personal computer, desktop computer, smart TV, wearable smart device, smart vehicle terminal, etc. The terminal may also include a client, which can be a browser client, instant messaging client, or mini-program, etc. The server can be an independent physical server, a server cluster composed of multiple physical servers, or a distributed system. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0022] The embodiments of this application can be applied to scenarios such as retrieval, search, recommendation, intelligent customer service, and data processing.

[0023] First, the following explanations are given for some of the nouns or terms that appear in the embodiments of this application: Visual-language embedding models: These models use a large visual-language model as the language backbone network, mapping inputs such as text, images, and videos to a unified semantic vector space for cross-modal retrieval. They are typically based on a Transformer architecture and trained through contrastive learning. The model input is text or an image, and the output is an L2-normalized high-dimensional vector. Retrieval is performed by calculating similarity through the vector inner product.

[0024] Audio encoder: A neural network that encodes raw audio signals into high-dimensional feature vectors, typically based on a transformer architecture. An audio encoder takes a Mel spectrogram as input and outputs a fixed-dimensional sequence of audio features. The audio encoder is a core component of audio understanding, and its performance directly determines the effectiveness of downstream retrieval tasks.

[0025] Audio projector: An adaptation module that maps the output features of the audio encoder to the latent space of the language backbone network. The role of the audio projector is to realize feature transformation between different dimensional spaces, enabling the audio features to be correctly processed by the language backbone network. In this application, the audio projector can be a two-layer multilayer perceptron (MLP) structure.

[0026] Full-modal embedding model: A unified embedding model that supports multiple modal inputs such as text, images, videos, visual documents, and audio. The embeddings of all modalities reside in the same vector space, enabling direct cross-modal retrieval without additional alignment steps.

[0027] The language backbone network is the core transformer network in the visual-language embedding model. It receives text tag sequences, image feature sequences, and expanded audio feature sequences, and jointly encodes multimodal information through a self-attention mechanism, outputting semantically aligned embedding vectors. The parameters of the language backbone network in this application can be fine-tuned during the training phase using a low-rank adapter.

[0028] Audio placeholders: A special type of marker (e.g., <|AUDIO|>) used to reserve positions for audio features in the input sequence. During inference or training, the embedding of this marker is replaced by the corresponding feature vector in the actual audio projection feature sequence, enabling the language backbone network to process audio features as part of the sequence.

[0029] Start marker: A special marker (e.g., <|audio_start|>) used to mark the starting boundary of the audio projection feature sequence in the extended sequence, helping the language backbone network to identify the starting position of the audio region.

[0030] End marker: A special marker (e.g., <|audio_end|>) used to mark the end boundary of the audio projection feature sequence in the extended sequence, helping the language backbone network to identify the end position of the audio region.

[0031] Padding markers: A special marker (e.g., <|audio_pad|>) is used to pad the extended sequence when its length is less than a preset length, ensuring that multiple extended sequences in the same batch have the same length, thus meeting the requirement of consistent batch input sequence length in the language backbone network. During self-attention computation, the padding marker positions are typically ignored by the mask and do not participate in feature fusion.

[0032] Extended sequence: The sequence obtained by replacing the embeddings of audio placeholders (and visual placeholders, etc.) in the input sequence with the corresponding actual feature vectors, and inserting start and end markers at the start and end positions of the audio (or visual) feature sequence, is directly input into the language backbone network for encoding processing.

[0033] Query embedding vector: After the language backbone network encodes the extended sequence, it extracts the feature vector of the last non-padded marker position from the encoded output sequence and normalizes it (e.g., L2 normalization) to obtain a unified embedding output vector representation. This vector is used to calculate similarity with pre-stored embedding vectors in the candidate library to achieve cross-modal retrieval.

[0034] Positional encoding: Positional information superimposed on each position vector of the extended sequence, enabling the language backbone network to distinguish vectors at different positions within the same sequence, thereby preserving the temporal information of the audio feature sequence. Positional encoding can be implemented using sine and cosine functions, learnable position embeddings, or Rotated Positional Encoding (RoPE), among other methods.

[0035] Self-attention mechanism: An attention mechanism in Transformer networks that allows each position in a sequence to pay attention to all other positions in the sequence, thereby capturing long-distance dependencies. In this application, the language backbone network uses a self-attention mechanism to jointly encode audio projection features, text tag embeddings, visual features, etc., in the extended sequence, achieving the fusion of multimodal features and semantic alignment.

[0036] Candidate library: A pre-constructed collection containing multiple candidate data (such as images, videos, documents, audio, text, etc.), where each candidate data has a corresponding embedding vector pre-generated and stored using the same method as in this application. During retrieval, the similarity between the query embedding vector and each candidate embedding vector in the candidate library is calculated, and the retrieval results are output in order of similarity.

[0037] Contrastive learning: a self-supervised or supervised learning method that enables the model to learn discriminative feature representations by narrowing the distance between matching positive sample pairs and widening the distance between mismatched negative sample pairs. This application uses Information Noise Contrastive Estimation (InfoNCE) loss as the training objective.

[0038] Information Noise Contrastive Estimation (InfoNCE) Loss: A contrastive learning loss function used to measure the proportion of similarity between a query embedding vector and its corresponding positive sample embedding vector relative to all sample embedding vectors. The core idea of ​​InfoNCE loss is to bring positive sample pairs closer together while pushing negative sample pairs further apart, thereby achieving cross-modal semantic alignment in a unified vector space.

[0039] Positive samples: In contrastive learning, samples that semantically match the query sample. For example, in audio-text pairing training samples, for a given audio query sample, the text description sample that matches it is a positive sample of that audio query sample.

[0040] Negative samples: In contrastive learning, samples that do not semantically match the query sample. These are typically other samples in the current batch besides the positive samples.

[0041] Low-Rank Adaptation (LoRA): A parameter-efficient fine-tuning method that adds a low-rank decomposition bypass (downward projection + upward projection) next to the original model parameter matrix. During training, only the parameters of the low-rank adapter are updated, while the original model parameters are frozen, thus achieving model adaptation to downstream tasks with minimal trainable parameters. In this application, the low-rank adapter has a rank of 32, a scaling factor (alpha) of 64, and a dropout rate of 0.05, i.e., rank=32, alpha=64, dropout=0.05.

[0042] Visual encoder: A pre-trained transformer network in a visual-language embedding model, used to convert visual inputs such as images and videos into sequences of visual features. In this application, the visual encoder maintains the original model parameters or is jointly fine-tuned through contrastive learning, and its output is used to replace visual placeholders in the input sequence.

[0043] Standardized depreciated cumulative gain: an evaluation metric for ranking quality that comprehensively considers the relevance of search results and the position of relevant results in the ranking list. In calculation, the gain values ​​of the top K results in the ranking list are depreciated according to their position and then summed to obtain the depreciated cumulative gain. This summation is then divided by the maximum depreciated cumulative gain under ideal ranking to obtain the standardized depreciated cumulative gain. For example, this application uses nDCG@10 as the evaluation metric.

[0044] Text-to-audio retrieval task: This refers to a retrieval task that uses text as the query and audio as the retrieval target. The query data is text, and the candidate library stores pre-stored embedding vectors for each candidate audio. By calculating the similarity between the text evaluation embedding vector of the query text and the embedding vectors of the candidate audio, the audio results that are most semantically relevant to the query text are retrieved from the candidate audio.

[0045] Audio-to-text retrieval task: This refers to a retrieval task that uses audio as the query and text as the retrieval target. The query data is audio, and the candidate library stores pre-stored embedding vectors corresponding to each candidate text. By calculating the similarity between the audio evaluation embedding vector of the query audio and the embedding vectors of the candidate texts, the text results that are most semantically relevant to the query audio are retrieved from the candidate texts.

[0046] Image retrieval task: This refers to a retrieval task that uses images as queries and images as retrieval targets. The query data is images, and the candidate database stores pre-stored embedding vectors corresponding to each candidate image. By calculating the similarity between the visual evaluation embedding vector of the query image and the embedding vectors of the candidate images, the image results that are most semantically relevant to the query image are retrieved from the candidate images.

[0047] Visual document retrieval task: This refers to a retrieval task that uses images as queries and documents containing visual information as retrieval targets. The query data is images, and the candidate library stores pre-stored embedding vectors corresponding to each candidate document (such as web pages containing images, documents with mixed text and images, etc.). By calculating the similarity between the visual evaluation embedding vector of the query image and the embedding vector of the candidate documents, the document results that are most semantically relevant to the query image are retrieved from the candidate documents.

[0048] The solutions provided in this application involve technologies such as data processing, which are specifically illustrated in the following embodiments. These embodiments are described in detail below. It should be noted that the order of description of the following embodiments is not intended to limit the priority of the embodiments.

[0049] It is understood that in the specific implementation of this application, user audio data, image data, video data and other related data are involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0050] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an application scenario of the data processing system provided in this application embodiment. The data processing system includes a terminal 10 and a server 20, etc.; the terminal 10 and the server 20 are connected via a network, such as a wired or wireless network.

[0051] Terminal 10 can be used to display a graphical user interface. Terminal 10 is used to interact with the user through the graphical user interface, such as downloading and installing a corresponding client and running it, calling and running a corresponding app, or presenting the corresponding graphical user interface by logging into the client.

[0052] In this embodiment of the application, terminal 10 can obtain query data containing audio data input by the user and send the query data to the server.

[0053] In this embodiment, server 20 can receive query data; when the query data contains audio data, it uses an external audio encoder extracted from a full-modal large language model to convert the audio data into an audio feature sequence; it maps the audio feature sequence to the latent space of the language backbone network through an audio projector to obtain an audio projection feature sequence; it determines an input sequence based on the query data, the input sequence containing at least one audio placeholder; it replaces the embedding of the audio placeholder in the input sequence with the corresponding feature vector in the audio projection feature sequence, and inserts a start marker before the start position and an end marker after the end position of the audio projection feature sequence to obtain an extended sequence; it inputs the extended sequence into the language backbone network for encoding processing to obtain a query embedding vector corresponding to the query data; it outputs the retrieval result based on the query embedding vector and sends the retrieval result to terminal 10 for display.

[0054] Please see Figure 2 , Figure 2 This is a schematic diagram of the framework of the full-modal embedding model provided in this application embodiment. The full-modal embedding model includes the following core components: a language backbone network (8B-parameter vision-language embedding model), responsible for processing text and visual input and outputting semantic embedding vectors; an external audio encoder, extracted from the full-modal large language model, for example, the architecture of this external audio encoder is a Whisper style transformer: 32 transformer layers, a hidden dimension of 1280, 20 attention heads, and a feedforward network (FFN) with a hidden dimension of 5120. This encoder receives a 128-dimensional Mel bins spectrum as input and outputs an audio feature sequence with a dimension of 2048. The number of parameters of this external audio encoder can be 647.9M; an audio projector, using a two-layer multilayer perceptron (MLP) structure, maps the 2048 dimensions to 4096 dimensions, with a parameter count of 18.9M.

[0055] The core design principle of this full-modal embedding model framework is "non-intrusive modal extension": while maintaining the original visual-language embedding model structure and parameters, audio input processing is achieved through a newly added audio pathway (external audio encoder + audio projector + four special audio tags). The external audio encoder extracts data from an external full-modal large language model rather than training it from scratch, leveraging its existing audio understanding capabilities. The audio projector plays a dual role in dimensionality adaptation and semantic transformation, bridging the output space of the external audio encoder with the latent space of the language backbone. The special tagging mechanism enables the language backbone to recognize and process audio input regions without modifying its self-attention calculation logic. This design minimizes the impact of audio access on the original visual retrieval capabilities while achieving alignment between audio and visual embeddings in the same vector space. The final model output is an L2-normalized high-dimensional vector, i.e., the query embedding vector.

[0056] This application provides a data processing method, which can be executed by a terminal or a server, or by both a terminal and a server. This application uses the example of a data processing method executed by a server to illustrate the method.

[0057] Please see Figures 3 to 5 , Figure 3 This is a flowchart illustrating the data processing method provided in an embodiment of this application. Figure 4 This is a schematic diagram of the inference stage of the data processing method provided in the embodiments of this application. Figure 5 This is a schematic diagram illustrating an application scenario of the data processing method provided in this application embodiment. The method may include the following steps: Step 110: Receive the query data.

[0058] The query data may include at least one of the following: audio data, text data, image data, and video data.

[0059] When the query data includes audio data, the query data can be monomodal audio data (containing only audio data) or multimodal data containing audio data, including audio data, and simultaneously containing one or more of text, image, or video data. Monomodal audio data includes, for example, voice commands captured in real-time by the user through a microphone, uploaded music clips, ambient audio, or audio tracks extracted from video. Methods for receiving audio data include, but are not limited to: receiving audio files uploaded by the user from the client, capturing audio streams in real-time from the microphone, and retrieving pre-stored audio query data from the server. If the query data also contains text, the original text string is retained; if it contains images or video, it is decoded into pixel tensors. After receiving the query data, the system can preprocess it, for example, by uniformly resampling the audio data to a 16kHz mono format for subsequent Mel-spectrum extraction.

[0060] If the query data also contains other modalities such as text, images, or videos (for example, a user enters the text "dog barking" and uploads an audio clip at the same time), this step will also receive them so that an input sequence with multimodal tags can be constructed later.

[0061] Step 120: When the query data contains audio data, the audio data is converted into an audio feature sequence using an external audio encoder extracted from a full-modal large language model.

[0062] The external audio encoder is an audio encoding component pre-extracted and isolated from a pre-trained full-modal large language model. This full-modal large language model natively supports audio input, and its audio encoder has been fully trained with audio-text alignment, possessing strong audio understanding capabilities. When extracting the external audio encoder, all weights and configurations of the external audio encoder are retained, including 32 Transformer layers, a 128-dimensional Mel-bins feature extraction front-end, and a projection layer with an output dimension of 2048.

[0063] For example, the architecture of this external audio encoder is a Whisper style transformer: 32 transformer layers, 1280 hidden dimensions, 20 attention heads, and a feedforward network (FFN) with 5120 hidden dimensions. This encoder receives a 128-dimensional Mel-bin spectrum as input and outputs a sequence of audio features with a dimension of 2048. The number of parameters in this external audio encoder can be 647.9M.

[0064] This external audio encoder, based on the Whisper architecture (a transformer-based end-to-end speech recognition framework), has been pre-trained on tens of thousands of hours of audio data and can be used directly without any additional pre-training steps. This "extraction rather than training" strategy is one of the key innovations of this application: training traditional audio encoders typically requires millions of audio-text pairing data and thousands of GPU hours of computing resources, while this application directly reuses the pre-trained audio encoder from the full-modal large language model, reducing the cost of acquiring the audio encoder to near zero.

[0065] It should be noted that the parameters of the external audio encoder can be kept fixed (frozen) after extraction, or they can be jointly updated during subsequent comparative learning training. This application does not limit this approach, and both methods fall within the scope of protection of this application. During the inference phase, the external audio encoder always uses the trained parameters to encode the input audio.

[0066] Step 130: The audio feature sequence is mapped to the latent space of the language backbone network through an audio projector to obtain the audio projection feature sequence.

[0067] Among them, the audio projector is a newly designed adapter module in this application, which is used to map the output dimension (first preset dimension, such as 2048) of the external audio encoder to the latent space dimension (second preset dimension, such as 4096) of the language backbone network.

[0068] In some embodiments, the audio projector is used to transform the dimension of the audio feature sequence from the output dimension of the external audio encoder to the latent space dimension of the language backbone network; the audio projector includes a first projection layer and a second projection layer, the first projection layer includes a first linear transformation layer and a Gaussian error linear unit, and the second projection layer includes a second linear transformation layer; the first projection layer is used to linearly transform each feature vector of the audio feature sequence from the output dimension of the external audio encoder to the latent space dimension of the language backbone network through the first linear transformation layer, and then process it through an activation function layer to output an intermediate feature sequence; the second projection layer is used to linearly transform each feature vector in the intermediate feature sequence through the second linear transformation layer to output the audio projected feature sequence, wherein the input dimension and output dimension of the second linear transformation layer are both the latent space dimension of the language backbone network. Specifically, the audio projector in this application adopts a two-layer multilayer perceptron (MLP) structure: the first projection layer (Layer 1 of the first MLP structure) is a first linear transformation layer (Linear(2048, 4096)) (performing linear transformation processing), followed by a Gaussian error linear unit (GELU) activation function layer (performing nonlinear transformation processing); the second projection layer (Layer 2 of the second MLP structure) is a linear transformation layer (Linear(4096, 4096)) without additional activation functions. The input of the projector is a 2048-dimensional feature vector sequence output by an external audio encoder, and the output is a 4096-dimensional feature vector sequence with the same dimension as the latent space of the language backbone, enabling the audio features to be seamlessly embedded into the input sequence of the language backbone network.

[0069] By employing a two-layer multilayer perceptron (MLP) structure containing a first and second linear transformation layer, and introducing a Gaussian error linear unit (GELU) activation function layer, the output dimension of the external audio encoder (e.g., 2048 dimensions) can be non-linearly mapped to the latent space dimension of the language backbone network (e.g., 4096 dimensions). Compared to a single-layer linear mapping, this structure can learn more complex audio-language feature transformation relationships, achieving better dimensionality adaptation and semantic alignment, and improving the discriminative power of feature representation. Simultaneously, the number of parameters is only about 18.9M, less than 0.2% of the total model, significantly improving cross-modal retrieval accuracy while maintaining extremely low training overhead.

[0070] In some embodiments, the step of mapping the audio feature sequence to the latent space of the language backbone network through an audio projector to obtain an audio projection feature sequence includes: sequentially inputting each feature vector in the audio feature sequence into the first linear transformation layer for linear transformation, and processing it through an activation function layer to output the corresponding intermediate feature vector in the intermediate feature sequence; inputting the intermediate feature vector into the second linear transformation layer for linear transformation to output the corresponding audio projection feature vector in the audio projection feature sequence.

[0071] By explicitly processing each feature vector in the audio feature sequence sequentially through a first linear transformation layer (including GELU) and a second linear transformation layer, the complete preservation of the temporal dimension is ensured. This allows the audio features at each time step to be independently mapped to the target latent space, providing a one-to-one feature alignment basis for the subsequent language backbone network to capture the temporal dependencies between audio frames using a self-attention mechanism. This design enables the audio features to fully retain their original information during the mapping process, and by gradually adjusting the feature dimensions and semantic expression through a two-layer MLP structure, the final output is an audio projection feature sequence that matches the dimensions of the language backbone latent space. This is beneficial for the fusion and interaction of audio features with other modal features in subsequent cross-modal retrieval tasks.

[0072] It should be noted that the audio projector was learned from random initialization during the contrastive learning training process of this application, while pre-trained fixed weights were used during the inference phase. The design of this projector enables this application to extend the audio retrieval capability of the original visual-language embedding model with extremely low parameter increments (approximately 0.2%), without modifying any structure of the original model.

[0073] Step 140: Determine the input sequence based on the query data, wherein the input sequence contains at least one audio placeholder.

[0074] Specifically, this embodiment adds four special markers to the vocabulary of the language backbone network: <|audio_start|> (start marker, used to mark the beginning position of audio input), <|audio_end|> (end marker, used to mark the end position of audio input), <|audio_pad|> (pad marker, used for padding alignment of audio features), and <|AUDIO|> (audio placeholder marker, used to represent placeholders for audio features). The input sequence is a template sequence composed of these special markers (and possibly text markers and visual placeholders) in a certain order, used to reserve space for subsequent replacement operations. The embedding vectors of the above special markers are learned together with other parameters of the language backbone network during training, enabling the language backbone network to recognize the boundaries of the audio input region and fuse audio features with other modal features (such as text features, visual features, etc.) through a self-attention mechanism.

[0075] In some embodiments, determining the input sequence based on the query data further includes: when the query data also contains text data, adding text placeholders corresponding to the text data to the input sequence, wherein the embedding of the text placeholders is directly obtained from the vocabulary of the language backbone network; and / or when the query data also contains image data and / or video data, adding visual placeholders corresponding to the image data and / or video data to the input sequence, wherein the visual placeholders are used to replace the visual feature sequences extracted from the image data and / or video data.

[0076] The way the input sequence is constructed depends on the modal composition of the query data. For example... Figure 4As shown, when the query data is audio data, the original audio signal is extracted using Mel spectral features and then fed into an external audio encoder to obtain an audio feature sequence of shape [T, 2048]. Here, T is the number of time steps, which depends on the audio duration and downsampling rate; typically, a 10-second audio file generates approximately 150-300 time steps. The audio feature sequence is then mapped by an audio projector to an audio projection feature sequence of shape [T, 4096], where each time step corresponds to a feature vector. In the input sequence, the embedding of a <|AUDIO|> tag (audio placeholder) is replaced with the corresponding feature vector in the audio projection feature sequence; that is, the audio projection feature vector at each time step replaces the embedding of a <|AUDIO|> tag (audio placeholder), thus creating a unified "tag". The "replacement" framework embeds audio features into the input sequence.

[0077] When query data also contains text, images, or videos, text or visual placeholders are added to allow the same input sequence to accommodate placeholders for multiple modalities, such as audio, text, and visual. For example, for queries containing text, the built-in tokenizer of the language backbone network converts the text into a sequence of text tokens (each text token corresponds to a word in the vocabulary), and adds them to the input sequence in a preset order (e.g., text tokens first, audio placeholders later). For queries containing images or videos, visual placeholders (e.g., <|image|>) are added. The initial embeddings of these placeholders are replaced with visual feature sequences in subsequent steps. Through this design, the same input sequence can accommodate placeholders for multiple modalities, such as audio, text, and visual. The embeddings of text placeholders are directly obtained from the vocabulary (without replacement), while visual placeholders are used to replace visual feature sequences, thus achieving a unified "tag". The "Replace" framework supports query inputs with arbitrary modal combinations, providing a flexible extension interface for full-modal retrieval.

[0078] Step 150: Replace the embedding of the audio placeholder in the input sequence with the corresponding feature vector in the audio projection feature sequence, and insert a start marker before the start position and an end marker after the end position of the audio projection feature sequence to obtain the extended sequence.

[0079] like Figure 4As shown, after replacing the embedding of the <|AUDIO|> marker (audio placeholder) in the input sequence with the corresponding feature vector in the audio projection feature sequence, the embedding of the <|audio_start|> marker (start marker) is inserted before the start position of the audio projection feature sequence, and the embedding of the <|audio_end|> marker (end marker) is inserted after the end position, resulting in an expanded sequence. The embedding vectors of the <|audio_start|> marker (start marker) and <|audio_end|> marker (end marker) are also obtained from the vocabulary of the language backbone network. This design enables the language backbone network to clearly identify the start and end boundaries of audio feature regions through a self-attention mechanism, thereby accurately distinguishing the feature sources of different modalities when fusing audio features with other modal features (such as text features, visual features, etc.), and avoiding misalignment of cross-modal features.

[0080] In some embodiments, obtaining the extended sequence further includes: when the length of the extended sequence is less than a preset length, inserting one or more padding markers after the end marker and at the end of the extended sequence, so that the length of the extended sequence reaches the preset length.

[0081] Specifically, the padding markers are the `<|audio_pad|>` markers in the vocabulary. When the length of the extended sequence is less than the preset length, one or more `<|audio_end|>` markers are inserted after the `<|audio_end|>` marker and at the end of the extended sequence to make the length of the extended sequence reach the preset length. During the self-attention calculation, the positions of these padding markers are masked, so that they do not participate in the calculation of attention weights, nor do they affect the fusion of effective features (audio projection features, text marker embeddings, start / end markers, etc.). This achieves batch alignment while ensuring that the padding operation does not interfere with the extraction of semantic features. When the length of the extended sequence is less than the preset length, inserting padding markers after the end marker can unify the length of the extended sequences corresponding to audio queries of different durations in the same batch, thereby meeting the batch processing requirements of the language backbone network. The padding markers are placed at the end of the sequence and are usually masked in self-attention, so they do not interfere with the learning and fusion of effective features, achieving efficient and lossless batch alignment.

[0082] In some implementations, when the query data also includes image data and / or video data, the method further includes: converting the image data and / or video data into a visual feature sequence using a visual encoder; and when generating the extended sequence, replacing the embedding of the visual placeholder in the input sequence with the corresponding feature vector in the visual feature sequence; wherein the visual encoder is a pre-trained visual encoder in the visual-language embedding model corresponding to the language backbone network.

[0083] Specifically, for image or video input, the original visual... A pre-trained visual encoder (e.g., ViT) in the language embedding model converts the input sequence into a sequence of visual features. Then, similar to the processing of audio placeholders, the embeddings of visual placeholders in the input sequence are replaced with corresponding feature vectors from the visual feature sequence, and corresponding start / end markers (e.g., <|visual_start|> and <|visual_end|>) are inserted before and after the visual feature sequence. By using the pre-trained visual encoder to convert images / videos into visual feature sequences and replacing the visual placeholders in the input sequence, the language backbone network can process visual features in exactly the same way as audio features. This design reuses the original visual... The pre-trained visual encoding capabilities in the language embedding model eliminate the need for additional visual encoder training, enabling low-cost visual modality access.

[0084] Finally, the sequence obtained after the above replacement and insertion operations is called the extended sequence. This extended sequence contains audio projection features (and may also contain visual features, text tag embeddings, etc.), and the boundaries of each modal region are clearly defined by start / end markers. It has been aligned to a uniform length by padding markers and can be directly input into the language backbone network for encoding.

[0085] Step 160: Input the extended sequence into the language backbone network for encoding processing to obtain the query embedding vector corresponding to the query data.

[0086] In this step, the vectors at each position in the extended sequence are first stacked with positional encodings, enabling the model to distinguish audio projection features at different time steps and text / visual markers at different positions, thus preserving the temporal information of the audio feature sequence. Then, the positionally encoded extended sequence is input into the language backbone network and encoded through its self-attention mechanism. The self-attention mechanism allows audio projection feature vectors at different positions in the extended sequence to interact and fuse temporal contextual information. Simultaneously, it enables cross-modal information interaction between audio projection features, text placeholder embeddings, and visual feature sequences in the same attention calculation, naturally aligning features from different modalities to the same semantic space during the encoding process. Specifically, in the self-attention mechanism, the feature vector at each position is used to calculate attention weights with the feature vectors at all other positions in the sequence, and then weighted and aggregated. Because the extended sequence simultaneously includes audio projection feature sequences, text placeholder embeddings (from the vocabulary), and potential visual feature sequences, the self-attention mechanism allows these features from different modalities to "see" each other and interact: for example, the text tag "dog barking" can establish a strong attention connection with the time segment representing the dog barking in the audio projection features, and the dog image in the visual features can also be associated with the audio features. Through multi-layered stacked cross-attention, features from different modalities gradually merge and align to the same semantic space during the encoding process. This design achieves end-to-end unified semantic representation without the need for a separate cross-modal alignment module.

[0087] like Figure 4 As shown, after encoding by the language backbone network, an encoded output sequence is obtained, with a shape of [L, D], where L is the length of the extended sequence (containing all markers and features), and D is the latent space dimension of the language backbone network (e.g., 4096). Then, the feature vector of the last non-padding marker position is extracted from the encoded output sequence and subjected to L2 normalization to obtain the query embedding vector. This query embedding vector is the high-dimensional representation of the query data in the unified semantic space and can be directly used for subsequent retrieval. The "non-padding marker position" refers to the valid position after excluding all padding markers (<|audio_pad|>). Since the padding markers in the extended sequence are placed at the end of the sequence (after the end marker), the last non-padding marker position is usually the first valid position immediately following the end marker <|audio_end|>, or the last valid position of the extended sequence when there are no padding markers. Under the self-attention mechanism, the feature vector at this position gathers information from all positions before it (including all audio projection features, text markers, visual features, start / end markers, etc.), making it most suitable as the global semantic representation of the entire query data.

[0088] In some embodiments, the step of inputting the extended sequence into the language backbone network for encoding processing to obtain the query embedding vector corresponding to the query data includes: superimposing positional encoding on the vector of each position in the extended sequence to obtain a position-encoded extended sequence; inputting the position-encoded extended sequence into the language backbone network and encoding it through the self-attention mechanism of the language backbone network to obtain an encoded output sequence; extracting the feature vector of the last non-filled marker position from the encoded output sequence and performing normalization processing to obtain the query embedding vector.

[0089] The introduction of positional encoding aims to preserve the temporal information of each element in a sequence. In natural language processing and audio processing, the order of a sequence often carries important semantic information. Positional encoding typically uses a combination of sine and cosine functions to generate vectors related to the sequence positions. These vectors are added to the original input vector, enabling the model to perceive the position of each element in the sequence. In audio processing, positional encoding preserves the temporal information of audio feature sequences, ensuring that the language backbone network considers the order of audio frames when processing them, thereby more accurately capturing the semantic content in the audio.

[0090] The self-attention mechanism is a core component of the language backbone network, allowing the model to dynamically focus on different parts of a sequence when processing it. For each element in the input sequence, the self-attention mechanism calculates its relevance score to all other elements; these scores are then weighted and summed to generate a contextual representation of the current element. In extended sequences, features from different modalities (such as audio projection features, text embeddings, and visual features) can interact and fuse contextual information under the influence of the self-attention mechanism. For example, audio frames can adjust their attention based on text content, or visual features can corroborate audio features, thereby enhancing the model's semantic understanding capabilities.

[0091] By encoding the superimposed positions of the extended sequence, the temporal information of the audio feature sequence is preserved. Encoding using the self-attention mechanism of the language backbone network enables audio frames at different positions to interact and integrate context. Extracting and normalizing the feature vector from the last non-padded marker position yields a query embedding vector that aggregates global information. This vector is directly used for similarity calculation, ensuring both accuracy and efficiency in retrieval.

[0092] In some embodiments, the step of inputting the extended sequence into the language backbone network for encoding processing to obtain the query embedding vector corresponding to the query data includes: inputting the extended sequence into the language backbone network, wherein the extended sequence contains features of different modalities, and the features of different modalities include at least one of the audio projection feature sequence, the embedding of the text placeholder, and the visual feature sequence; jointly encoding the features of different modalities in the extended sequence through the self-attention mechanism of the language backbone network, so that the features of different modalities are fused together and aligned to the same semantic space during the encoding process to obtain an encoded output sequence; extracting the feature vector of the last non-filler mark position from the encoded output sequence and performing normalization processing to obtain the query embedding vector.

[0093] The extended sequences are designed to achieve unified processing of multimodal information. The audio projection feature sequence is obtained by converting the original audio signal using an audio encoder and projector, capturing key information within the audio. Embedded text placeholders represent the position and content of text information within the sequence. The visual feature sequence is extracted from images or videos using a visual encoder. These features from different modalities coexist in the extended sequences, providing a rich source of information for the language backbone network, enabling the model to consider audio, text, and visual information simultaneously, thus achieving a more comprehensive understanding of the query data.

[0094] In the joint encoding process, the self-attention mechanism not only focuses on feature interactions within the same modality but also promotes the fusion of features from different modalities. For example, when processing a query containing both audio and text, the self-attention mechanism might make audio features focus more on audio segments related to the text content, while making text features focus more on words that match the semantics of the audio. This cross-modal interaction and fusion allows features from different modalities to gradually align to the same semantic space during the encoding process, thus achieving true multimodal unified retrieval. This process does not require designing a separate cross-modal alignment module but fully utilizes the original self-attention capabilities of the language backbone network, achieving cross-modal alignment of audio, text, and vision at extremely low cost.

[0095] The encoded output sequence contains all the feature information processed by the self-attention mechanism. To obtain a holistic representation of the query data, the feature vector at the last non-padded marker position is typically extracted from the encoded output sequence. This feature vector encapsulates the global information in the sequence and can effectively represent the semantic content of the query data. Subsequently, this feature vector is normalized (e.g., L2 normalization) to ensure comparability of embedding vectors from different query data when calculating similarity, thereby guaranteeing the accuracy and efficiency of retrieval. The final query embedding vector can be directly used for similarity calculation, providing strong support for subsequent retrieval tasks.

[0096] The extended sequence simultaneously includes audio projection features, text placeholder embeddings, and / or visual feature sequences. These are jointly encoded using the self-attention mechanism of the language backbone network, allowing features from different modalities to naturally merge and align to the same semantic space during the encoding process. This eliminates the need for a separate cross-modal alignment module, achieving true end-to-end multimodal unified retrieval, significantly reducing system complexity and improving retrieval performance. This step fully utilizes the original self-attention capability of the language backbone network, achieving cross-modal alignment of audio, text, and vision at extremely low cost.

[0097] Step 170: Output the retrieval results based on the query embedding vector.

[0098] Before retrieval, a candidate library needs to be built in advance. The candidate library pre-stores pre-stored embedding vectors corresponding to at least one modality such as images, videos, documents, audio, and text. All pre-stored embedding vectors have been obtained through the same encoding process as described above and are in the same semantic space as the query embedding vector. Specifically, the process of constructing the candidate library is completely consistent with the process of generating the query embedding vector: For each candidate data in the candidate library, if it is audio data, the audio projection feature sequence is obtained through Mel feature extraction, audio encoder, and audio projector. After replacing the <|AUDIO|> marker and inserting the <|audio_start|> and <|audio_end|> markers, it is sent to the language backbone network for encoding. The feature vector is extracted at the last non-filled position and L2 normalized to obtain the audio pre-stored embedding vector. If it is text data, the embedding of its text placeholder markers is directly obtained from the vocabulary of the language backbone network. It is encoded in the same way and the feature vector at the last non-filled position is extracted and L2 normalized to obtain the text pre-stored embedding vector. If it is image or video data, it is converted into a visual feature sequence by the pre-trained visual encoder in the visual-language embedding model corresponding to the language backbone network. After replacing it with visual placeholder markers, it is sent to the language backbone network for encoding. Similarly, it is extracted at the last non-filled position and L2 normalized to obtain the visual pre-stored embedding vector. All pre-stored embedding vectors are 4096-dimensional and L2-normalized, residing in the same 4096-dimensional hyperspherical semantic space as the query embedding vector. The system calculates the cosine similarity between the query embedding vector and each candidate vector in the candidate library. Since both the query embedding vector and the candidate vectors in the candidate library have been L2-normalized, the similarity is expressed as a vector dot product (i.e., cosine similarity). A higher cosine similarity value indicates greater semantic relevance. The candidate results are sorted from highest to lowest cosine similarity, and the Top-K candidate vectors with the highest cosine similarity are selected (K is a preset positive integer, such as 10 or 100). K can be adjusted according to the application scenario: for search engine homepage display, K is typically 10-30; for candidate set retrieval in recommendation systems, K can be 100-1000. The candidate data (such as images, video clips, documents, audio clips, text paragraphs, etc.) corresponding to the selected candidate vectors are output as the search results. For example, the output format may include: displaying the results in a list or grid format on the user interface, along with similarity scores; or returning the result identifier and similarity score through an application programming interface (API).

[0099] For example, when a user queries "a bird call" using voice, the system extracts the query embedding vector of that voice and calculates its similarity with the pre-stored embedding vectors of all audio segments in the candidate library. Finally, it outputs the audio segments with the highest similarity (such as "thrush call.wav" and "magpie call.mp3") as the search results. Simultaneously, since the candidate library also includes modalities such as images and videos, the system can also return bird images or videos semantically related to "bird call," achieving cross-modal retrieval.

[0100] In some embodiments, the step of outputting retrieval results based on the query embedding vector includes: calculating the similarity between the query embedding vector and each candidate vector in the candidate library, wherein the candidate library contains pre-stored embedding vectors corresponding to at least one modality among images, videos, documents, audio, and text; selecting at least one candidate vector in descending order of similarity; and outputting the candidate data corresponding to the candidate vector as the retrieval result.

[0101] The system calculates the similarity between the query embedding vector and the pre-stored embedding vectors of various modalities (images, videos, documents, audio, and text) in the candidate database, and outputs the search results in order of similarity. It can directly reuse a unified embedding space for cross-modal comparisons. Since the embedding vectors of all modalities reside in the same semantic space, it can achieve various cross-modal retrieval tasks such as "text to audio," "audio to image," and "image to text," without requiring separate training of retrieval models for each modality. All modal retrievals can share a single model and index, avoiding the operational costs of deploying separate retrieval systems for each modality as in traditional methods.

[0102] like Figure 5 As shown, the data processing method provided in this application embodiment can be implemented through a full-modal embedding model. This full-modal embedding model supports users initiating queries in multiple modal forms. User query input can include any combination of one or more modalities such as text, audio, images, and videos. It can also retrieve multimodal search results (including at least one of image results, video results, document results, audio results, and text results) in a unified vector space and return cross-modal ranking results.

[0103] Single-modal query: Users can input only text (e.g., "dog barking"), upload only audio clips (e.g., a bird song recording), upload only images (e.g., a photo of a device indicator light), or upload only videos (e.g., a video recording of a malfunction). After processing by the method of this application, the above single-modal query data can all obtain corresponding query embedding vectors, and similarity calculations can be performed with the pre-stored embedding vectors of all modalities in the candidate library to achieve cross-modal retrieval.

[0104] Multimodal combined query: Users can simultaneously input query data in multiple modalities, such as "text + audio" (e.g., the text "Find songs similar to this audio" and upload an audio clip), "text + image" (e.g., "What animal is in this picture?" and upload an image), "audio + video" (e.g., hum a melody and upload a concert video clip), or even combinations of three or more modalities (e.g., text + audio + image). The method in this application can process query data in these modalities simultaneously. Its underlying principle is that the features of all modalities are ultimately mapped to the same unified vector space of a multimodal large language model. When constructing the extended sequence, the system sequentially concatenates the features of different modalities (embedding of text placeholders, audio projection feature sequences, and visual feature sequences), and uses the self-attention mechanism of the language backbone network to make the modal markers mutually attentive and fused, ultimately outputting a unified query embedding vector that simultaneously contains semantic information from multiple modalities, thereby achieving true multimodal joint retrieval and cross-modal alignment.

[0105] The embodiments of this application can cover application scenarios in multiple fields such as search, recommendation, and intelligent customer service.

[0106] For example, consider the application scenario of a unified multimodal retrieval system. A user enters text into the search box, and the system simultaneously returns matching images, videos, visual documents, and audio clips. Since all modalities share the same embedding space, search results can be sorted across modalities, achieving a truly multimodal search experience. Specifically, in this application scenario, a user enters a text query (e.g., "dog barking") through the search interface. The system executes the method of this application, specifically including: receiving query data containing text data (excluding audio). In this case, the method of this application can directly reuse existing visual... The language embedding model's text processing capabilities include: converting the text "dog barking" into a text tag sequence, inputting it into a language backbone network for encoding and normalization to obtain a text query embedding vector. The candidate library pre-stores pre-existing embedding vectors for various modalities such as images, videos, visual documents, and audio clips. All embedding vectors are generated through the encoding process described in this application (audio through an external audio encoder + audio projector + language backbone network, images / videos through a visual encoder + language backbone network, and text directly through the language backbone network) and reside in the same semantic space. The system calculates the cosine similarity between the text query embedding vector and the candidate vectors for each modality in the candidate library, sorting them across modalities from high to low similarity to obtain a list of search results. The final output can simultaneously include matched images (e.g., photos of dogs), videos (video clips of dogs running), visual documents (PDF documents introducing canines), and audio clips (recordings of dog barking). Because all modalities share the same embedding space, search results can be mixed according to semantic relevance, allowing users to obtain comprehensive search results without switching search categories. In another example, when a user enters "rain sound" as a text query, the system can simultaneously return audio clips of rain sounds, images of rainy scenes, and videos related to rainy days, achieving a truly multimodal search experience.

[0107] For example, consider an application scenario of audio content retrieval. A user uploads an audio clip, and the system can retrieve semantically matching text descriptions, related images, or videos. Specifically, in this application scenario, the user uploads an audio clip (e.g., birdsong recorded on a mobile phone). The system executes the method of this application, which specifically includes: receiving query data containing audio data; converting the audio into an audio feature sequence using an external audio encoder extracted from a full-modality large language model, and then mapping it to the latent space of the language backbone network through an audio projector to obtain an audio projection feature sequence; determining an input sequence (containing T audio placeholders) based on the audio data; replacing the embeddings of the placeholders with the corresponding feature vectors in the audio projection feature sequence; and inserting start and end markers at the start and end positions respectively to obtain an extended sequence; inputting the extended sequence into the language backbone network for encoding, and performing L2 normalization on the feature vector at the last non-filled marker position in the encoding output to obtain an audio query embedding vector; storing pre-stored embedding vectors for various modalities such as text descriptions, images, and videos in the candidate library; and calculating and ranking the similarity between the audio query embedding vector and each candidate vector. Output search results: For example, return the text description "the clear chirping of birds in the forest" that semantically matches the query data, as well as related bird images and bird ecology videos. Because audio embeddings reside in the same semantic space as text and visual embeddings, the system can achieve cross-modal retrieval such as "audio search for text" and "audio search for images".

[0108] For example, consider the application scenarios of multimodal recommendation systems. In content recommendation scenarios, the system can recommend text, images, videos, and audio content simultaneously based on the user's text query or audio input, providing a richer recommendation experience. Specifically, in this application scenario, content recommendations are made by the user humming a melody (e.g., humming the chorus of a popular song). The system executes the method of this application, specifically including: receiving query data containing audio data (user humming audio); extracting audio projection feature sequences according to the same process as in audio content retrieval application scenarios; constructing extended sequences; and obtaining audio query embedding vectors after encoding by a language backbone network. The candidate library pre-stores pre-stored embedding vectors of different modalities such as song audio, music videos, concert images, and lyrics. The system calculates the similarity between the humming query embedding vector and all candidate vectors in the candidate library, and sorts them from high to low similarity. The recommendation results are output: for example, the highest-ranked song audio (the original song that best matches the humming melody), related music videos (official MVs), concert images, and corresponding lyrics. This recommendation result integrates multiple modalities, providing users with a richer experience. In another example, users can also directly enter a text query (e.g., "XX singer's XX song"). The system uses a text processing flow to obtain the query embedding vector, and can also recommend song audio, video, images, and text lyrics at the same time, realizing text-based multimodal recommendation.

[0109] For example, consider the application scenario of audio understanding in intelligent customer service. In intelligent customer service scenarios, users may describe problems via voice, and the system needs to understand the voice content and retrieve relevant documents, images, or videos to respond. This application enables the customer service system to unify audio input with text and visual knowledge base retrieval, improving problem-solving efficiency. Specifically, in this application scenario, the intelligent customer service system receives fault problems described by users via voice, needs to understand the voice content, and retrieve relevant documents, images, or videos from the knowledge base to provide solutions. The system executes the method of this application, specifically including: the user describes the problem through the voice entry of the customer service system, such as "My device is beeping and cannot start." The system receives query data (user voice) containing audio data, and may also include text data (such as user ID or device model) as supplementary information. The user's voice is converted into an audio feature sequence using an external audio encoder extracted from a full-modal large language model, and then mapped to the latent space of the language backbone network through an audio projector to obtain an audio projection feature sequence. This step can extract acoustic features (such as the frequency, rhythm, and duration of the "beeping sound") and semantic content (textual intent after speech recognition) from the voice. The input sequence is determined based on the query data. If the query contains only audio, the input sequence includes audio placeholders corresponding to the number of audio time steps; if it also contains text data (such as the device model entered by the user), the input sequence also includes text placeholders. The embeddings of the audio placeholders are replaced with audio projection feature sequences, and start and end markers are inserted at the start and end positions to obtain an expanded sequence. The expanded sequence is input into the language backbone network, which jointly encodes the audio features (and text features, if any) through a self-attention mechanism. Since the language backbone network has aligned the audio embeddings with the text and visual embeddings to the same semantic space through contrastive learning, it can understand the fault type corresponding to the "beeping sound" (e.g., a buzzer alarm indicating a hardware fault). The feature vector at the last non-filled marker position is extracted from the encoded output sequence and L2 normalized to obtain the audio query embedding vector. The candidate library pre-stores various types of content from the customer service knowledge base and their corresponding embedding vectors, including: troubleshooting documents (text), equipment structure diagrams (images), fault indicator flashing videos (videos), and repair audio guides, etc. The system calculates the cosine similarity between the audio query embedding vector and these candidate vectors, sorts them by similarity from high to low, and obtains a list of cross-modal retrieval results. The system returns the most relevant results to the user. For example, it returns a troubleshooting document (text) matching "beeping sound," which reads, "Continuous short beeps indicate a memory fault; please check if the memory stick is loose." It also returns images of relevant alarm indicator lights, marking the location of the fault lights and comparing their normal / abnormal states. If a corresponding video is available, it can also play a repair guide video on "how to disassemble the device casing to check the memory."

[0110] In some embodiments, the method further includes: acquiring an audio query sample and a text description sample paired with the audio query sample; converting the audio query sample into an audio training feature sequence using the external audio encoder; mapping the audio training feature sequence to the latent space of the language backbone network using the audio projector to obtain an audio training projection feature sequence; determining a training input sequence based on the audio query sample, the training input sequence containing at least one audio placeholder; replacing the embedding of the audio placeholder in the training input sequence with the corresponding feature vector in the audio training projection feature sequence, and in the audio training projection feature sequence... A start marker is inserted before the start position and an end marker is inserted after the end position to obtain a training extension sequence. The training extension sequence is input into the language backbone network for encoding to obtain an audio training embedding vector. The text description sample is converted into a text tag sequence, and the text tag sequence is input into the language backbone network for encoding to obtain a text training embedding vector. Based on the audio training embedding vector and the text training embedding vector, the external audio encoder, the audio projector, and the language backbone network are trained through contrastive learning to align the audio training embedding vector and the text training embedding vector in the semantic space of the language backbone network.

[0111] During the training phase, this application uses audio. Text-to-audio pairing training data was used for comparative learning training. The training data included text-audio retrieval pairs (audio) from AudioCaps (approximately 50,000 audio description pairs) and AudioSetStrong (approximately 130,000 strongly labeled audio classification pairs). The training data consists of approximately 180,000 text-to-audio pairs, used during training to retrieve audio from text and text from audio. Audio was retrieved through this process. Text-paired data (including audio query samples and text description samples paired with the audio query samples) is used to generate training embeddings by performing the same encoding process as in the inference phase on both audio and text. Then, the parameters of the audio projector, language backbone network, and external audio encoder are jointly updated based on contrastive learning. This training only requires audio. With approximately 180,000 text-pairing data entries, audio embeddings and text embeddings can be aligned in semantic space without the need for image or video data. This approach offers low training costs and fast convergence.

[0112] During training, the parameters of the language backbone network are efficiently fine-tuned using a low-rank adapter (LoRA). Specifically, LoRA modules are connected in parallel to the attention and feedforward layers of the language backbone network, with a rank of 32, a scaling factor of alpha of 64, and a dropout rate of 0.05. During training, only the parameters of the LoRA modules and the audio projector are updated, while the original parameters of the language backbone network are frozen. This allows training only about 1-2% of the total parameters, significantly reducing memory requirements and training time. The parameters of the external audio encoder can be selectively updated (completely frozen or only fine-tuning higher layers). The training hyperparameters are set as follows: global batch size 256, maximum sequence length 1024, and learning rate 1e-4 (i.e., 1×10⁻⁴). 4 Cosine learning rate scheduling is used.

[0113] By acquiring audio The training data is paired with text, and the same encoding process as in the inference phase is performed on both audio and text to obtain training embeddings. Then, based on contrastive learning, the parameters of the audio projector, language backbone network, and external audio encoder are jointly updated. This training only requires audio. With approximately 180,000 text-pairing data entries, audio embeddings and text embeddings can be aligned in semantic space without the need for image or video data. This approach offers low training costs and fast convergence.

[0114] In some embodiments, training the external audio encoder, the audio projector, and the language backbone network through contrastive learning based on the audio training embedding vector and the text training embedding vector includes: for each audio training embedding vector in the current batch, calculating the similarity between the audio training embedding vector and all text training embedding vectors in the current batch, wherein all text training embedding vectors include positive sample embedding vectors of text description samples paired with the audio training embedding vector and negative sample embedding vectors corresponding to other audio training embedding vectors; dividing the similarity between the audio training embedding vector and its positive sample embedding vector by the sum of the similarities between the audio training embedding vector and all text training embedding vectors in the current batch to obtain a ratio; taking the logarithm and negative of the ratio to obtain the sub-loss corresponding to the audio training embedding vector; averaging the sub-losses corresponding to all audio training embedding vectors in the current batch to obtain the contrastive loss; updating the parameters of the audio projector, the language backbone network, and the external audio encoder according to the contrastive loss.

[0115] This application uses Information Noise Contrast Estimation (InfoNCE) loss (hereinafter referred to as contrast loss) as the training objective, which can be expressed as the following formula (1): (1); Among them, z i This is the i-th audio training embedding vector in the current batch (i.e., the query embedding vector during the training phase). This audio training embedding vector is obtained by encoding the audio query sample through an external audio encoder, audio projector, and language backbone network, extracting the feature vector at the last non-padding marker position, and then performing L2 normalization; z i + To train the embedding vector z with audio i The positive sample embedding vector of the paired text description samples is obtained by encoding the paired text description samples through a language backbone network, extracting the feature vector at the last non-padding marker position, and then performing L2 normalization; z j Let be the training embedding vector for the j-th text in the current batch. When j=i, it is the embedding vector for positive samples, and the rest are the embedding vectors for negative samples. τ is the temperature parameter (e.g., set to 0.02), and N is the global batch size (e.g., set to 256). The core idea of ​​InfoNCE loss is to bring the embedding distance of positive sample pairs closer together while pushing the embedding distance of negative sample pairs further apart, thereby achieving cross-modal semantic alignment in a unified vector space.

[0116] Specifically, the step-by-step implementation of the InfoNCE contrastive loss is as follows: for each audio training embedding vector z in the current batch... i First, calculate the audio training embedding vector z. i Compared with the training embedding vectors {z} of all N texts in the current batch j The similarity between the vectors is reduced to the vector dot product z since both have been L2 normalized. i ·z j Then, the audio training embedding vector z is trained. i Its positive sample embedding vector z i + The vector inner product z i ·z i + Divide by z i The sum of the inner product exponents of all N text training embedding vectors This yields the probability ratio of positive samples in the softmax distribution. The logarithm of this ratio is taken and then negative to obtain the sub-loss corresponding to the i-th audio training embedding vector. Finally, the average of the sub-losses corresponding to all N audio training embedding vectors in the current batch is calculated to obtain the contrast loss. The temperature parameter τ = 0.02 makes the inner product of positive sample pairs significantly higher than that of negative sample pairs after exponential operation, generating a stronger gradient signal, i.e., when z... i With z i +When the inner product ratio of positive samples is more than 0.02 higher than that of negative samples, the probability of positive samples will dominate, thereby driving the audio projector to learn to accurately map audio features to a semantic space consistent with the text embedding, while driving the language backbone network and external audio encoder to learn more discriminative cross-modal representations.

[0117] The InfoNCE contrastive loss is implemented stepwise. Within the current batch, the similarity between each audio embedding and all text embeddings is calculated, and the similarity of matched positive samples is compared with the similarity of all negative samples. This effectively brings positive sample pairs closer together and pushes negative sample pairs further apart. This loss function directly optimizes the discriminative power of the embedding space, significantly enhancing the trained model's ability to discriminate cross-modal semantic similarity.

[0118] In some embodiments, the method further includes: updating the parameters of the visual encoder according to the contrastive loss during training via contrastive learning; wherein the visual encoder is jointly optimized with the language backbone network, the audio projector, and the external audio encoder; the visual encoder is used to convert image data and / or video data into visual feature sequences.

[0119] During the contrastive learning training process, this application optimizes not only audio-related modules but also the parameters of the visual encoder. Specifically, for image or video samples included in the training data (e.g., obtained from a multimodal dataset), they are converted into visual feature sequences by the visual encoder and an extended sequence is constructed as input to the language backbone network in the same manner as the audio features to obtain visual training embedding vectors. Then, the audio training embedding vectors, text training embedding vectors, and visual training embedding vectors are incorporated into the contrastive learning framework and jointly optimized using multimodal contrastive loss (e.g., treating various pairings such as audio-text, audio-visual, and text-visual as positive samples). This joint optimization process enables the output features of the visual encoder to align with the audio and text features into the same semantic space, thereby ensuring that image / video queries can also perform cross-modal retrieval with audio and text during the inference stage. During the contrastive learning training process, the parameters of the visual encoder are updated according to the contrastive loss, realizing the joint optimization of the visual encoder, the language backbone network, the audio projector, and the external audio encoder. This joint optimization method can fully utilize the correlation information between multimodal data, improve the model's ability to process and fuse different modal data, and thus further enhance the performance of cross-modal retrieval. During training, the language backbone network, external audio encoder, audio projector, and visual encoder are optimized simultaneously. The optimization of the language backbone network is achieved through low-rank adaptation (LoRA), and the visual encoder can use pre-trained weights as initialization and participate in gradient updates.

[0120] In some embodiments, the method further includes: acquiring an audio retrieval evaluation dataset, the audio retrieval evaluation dataset comprising text-to-audio retrieval tasks and audio-to-text retrieval tasks, wherein each query data in each task is pre-configured with at least one correct retrieval result, and the evaluation dataset also pre-stores candidate embedding vectors corresponding to each candidate result in a candidate library; for each text query data in the text-to-audio retrieval task, converting the text query data into a text tag sequence, inputting the text tag sequence into the language backbone network for encoding processing to obtain a text evaluation embedding vector; for each audio query data in the audio-to-text retrieval task, using the external audio encoder to convert the audio query data into an audio feature sequence, and mapping the audio feature sequence through the audio projector. The audio projection feature sequence is obtained from the latent space of the language backbone network. An input sequence containing audio placeholders is determined. The embedding of the audio placeholders is replaced with the corresponding feature vector in the audio projection feature sequence. A start marker is inserted before the start position and an end marker is inserted after the end position to obtain an extended sequence. The extended sequence is input into the language backbone network for encoding processing to obtain an audio evaluation embedding vector. The similarity between the text evaluation embedding vector or the audio evaluation embedding vector and each candidate embedding vector in the candidate library is calculated respectively. The candidate results are sorted from high to low according to the similarity to obtain a first retrieval result list. The first retrieval result list is compared with the correct retrieval results corresponding to the corresponding query data, and the audio retrieval evaluation result is output according to the standardized loss cumulative gain index.

[0121] During the evaluation phase, this application uses the audio retrieval task from the Massive Text Embedding Benchmark (MTEB) as the standard evaluation dataset. This evaluation dataset can include AudioCaps (an audio caption dataset with multiple audio-text description pairs) and AudioSetStrong (a strongly labeled audio event classification dataset with tens of thousands of strongly labeled audio classification pairs), each supporting bidirectional text-to-audio (T2A) and audio-to-text (A2T) retrieval tasks. The candidate library consists of all candidate results (e.g., all audio segments or all text descriptions) from the evaluation dataset. The embedding vector of each candidate result is pre-computed and stored using the same encoding process as in this application, residing in the same semantic space as the query embedding vector. By obtaining the audio retrieval evaluation dataset and evaluating the model on text-to-audio and audio-to-text retrieval tasks, the model's performance in audio retrieval can be comprehensively assessed. The Standardized Decay Cumulative Gain (nDCG@10) metric is calculated as follows: For each query, based on the correctness of the top 10 results in the ranking list, the ratio of the Decay Cumulative Gain (DCG) to the Ideal Decay Cumulative Gain (IDCG) is calculated. This metric comprehensively considers the relevance of the search results and the position of the relevant results in the ranking list; a value closer to 1 indicates better search performance. The application of this metric can objectively and accurately measure the quality of the model's search results, providing strong support for model optimization and improvement.

[0122] In some embodiments, the method further includes: acquiring a visual retrieval evaluation dataset, the visual retrieval evaluation dataset including image retrieval tasks and visual document retrieval tasks, each visual query data in each task being pre-configured with at least one correct retrieval result, and the evaluation dataset also pre-storing candidate embedding vectors corresponding to each candidate result in a candidate library; for each visual query data in the image retrieval task or visual document retrieval task, using the visual encoder to convert the visual query data into a visual feature sequence, determining an input sequence containing visual placeholders, replacing the embedding of the visual placeholders with the corresponding feature vector in the visual feature sequence, inserting a start marker before the start position and an end marker after the end position to obtain an extended sequence, inputting the extended sequence into the language backbone network for encoding processing to obtain a visual evaluation embedding vector; calculating the similarity between the visual evaluation embedding vector and each candidate embedding vector in the candidate library, sorting the candidate results from high to low similarity to obtain a second retrieval result list; comparing the second retrieval result list with the correct retrieval results corresponding to the corresponding query data, and outputting the visual retrieval evaluation result according to the standardized loss cumulative gain index.

[0123] During the evaluation phase, this application may also use multiple (e.g., 60) visual retrieval tasks from the Massive Multimodal Embedding Benchmark (MMEB) as standard evaluation datasets. This dataset includes 36 image retrieval tasks (e.g., image classification, image-text retrieval, visual question answering, etc.) and 24 visual document retrieval tasks (e.g., information retrieval from document images, chart understanding, etc.), covering various task types such as classification, visual question answering (VQA), retrieval, and grounding. For example, classification tasks determine the category label of an image or video, such as identifying whether an image contains the word "dog"; visual question answering tasks are used to answer natural language questions based on image content, such as answering "red" to the question "What color car is in the picture?"; retrieval tasks are used to find semantically matching results from the candidate library based on a query (text, image, or audio), such as searching for images based on text or searching for audio based on images; and grounding tasks are used to accurately locate entities or phrases in text descriptions to their corresponding regions in images / videos (e.g., bounding boxes or pixel masks), such as outlining the location of a person in an image based on "the person wearing a hat on the left". The candidate library consists of the embedding vectors of all candidate images or visual documents in the evaluation dataset, which are pre-computed and stored using the visual encoder and language backbone network described in this application. For image queries, a visual encoder is used directly to extract visual feature sequences; for visual document queries, a visual encoder is also used to process document images. By obtaining a visual retrieval evaluation dataset and evaluating the model on image retrieval and visual document retrieval tasks, the model's performance in visual retrieval can be comprehensively assessed. This evaluation method can verify the model's ability to process visual data such as images and visual documents and its retrieval effectiveness, providing guidance for further optimization and improvement of the model. Meanwhile, the application of the standardized depreciation cumulative gain (nDCG@10) metric ensures the objectivity and accuracy of the evaluation results. This metric is also applicable to visual retrieval tasks, measuring the quality of the correct position in the ranking list returned by the model, and is a common metric in standard evaluation frameworks such as MTEB / MMEB.

[0124] In another embodiment, the external audio encoder includes at least two audio encoders, namely a first audio encoder and a second audio encoder, wherein the first audio encoder and the second audio encoder are audio encoders of different models; the method further includes: using the first audio encoder to convert the audio query data into a first audio feature sequence, and using the second audio encoder to convert the audio query data into a second audio feature sequence; mapping the first audio feature sequence to the latent space of the language backbone network through a first audio projector to obtain a first audio projection feature sequence, and mapping the second audio feature sequence to the latent space of the language backbone network through a second audio projector to obtain a second audio projection feature sequence; performing weighted fusion of the first audio projection feature sequence and the second audio projection feature sequence to obtain a fused audio projection feature sequence; and replacing the embedding of the audio placeholder with the corresponding feature vector in the fused audio projection feature sequence to obtain the extended sequence.

[0125] To improve the robustness and generalization ability of audio understanding, the external audio encoder can include multiple audio encoders of different types, fusing audio features from different pre-trained models. Specifically, audio encoders are extracted from multiple full-modal large language models or dedicated audio models. For example, a first audio encoder is extracted from the Whisper model, a second audio encoder is extracted from a full-modal large language model (such as Qwen2.5-Omni), and a third audio encoder is extracted from a contrastive language-audio pre-trained model (CLAP). Each encoder independently processes the same original audio signal, outputting a first audio feature sequence, a second audio feature sequence, and a third audio feature sequence, respectively. A corresponding sub-projector is configured for each audio encoder, and each sub-projector has the same two-layer MLP structure as the aforementioned audio projector. Each sub-projector maps the output feature sequence of the corresponding encoder to the latent space of the language backbone network, obtaining the first audio projection feature sequence, the second audio projection feature sequence, and the third audio projection feature sequence.

[0126] During the inference phase, multiple audio projection feature sequences are weighted and fused to obtain the final audio projection feature sequence. The fusion method can employ: fixed weights (e.g., equal weights for each encoder), learnable weights (weight parameters are updated via backpropagation during training), or dynamic weights (weights are adaptively calculated based on the type or confidence level of the input audio). The fused audio projection feature sequence replaces the audio placeholders in the input sequence, and start and end markers are inserted to obtain an expanded sequence. This expanded sequence is then input into the language backbone network for encoding, ultimately yielding the query embedding vector.

[0127] This implementation integrates multiple complementary audio encoders, which can combine the advantages of different models in speech recognition, audio event detection, music understanding, etc., and improve the ability to represent complex audio (such as noisy environments and multi-source sounds), thereby further improving the accuracy and robustness of cross-modal retrieval.

[0128] In another embodiment, the audio projector includes a first projection layer and a second projection layer connected in sequence. The first projection layer includes a first linear transformation layer, a layer normalization layer, and an activation function layer. The first linear transformation layer is used to linearly transform each feature vector of the audio feature sequence from the output dimension of the external audio encoder to the latent space dimension of the language backbone network. The layer normalization layer is used to normalize the output of the first linear transformation layer. The activation function layer performs a nonlinear transformation on the output of the layer normalization layer. The second projection layer includes a second linear transformation layer and a residual connection module. The second linear transformation layer is used to maintain the dimension of the output of the first projection layer as the latent space dimension of the language backbone network. The residual connection module is used to add the input of the first projection layer to the output of the second projection layer after dimension matching.

[0129] To enhance the training stability of the audio projector and alleviate the gradient vanishing problem in deep networks, the structure of the first and second projection layers is improved by introducing residual connections and layer normalization layers. The specific structure is as follows: The first projection layer includes a first linear transformation layer, a layer normalization layer, and an activation function layer. The first linear transformation layer, Linear(2048, 4096), linearly transforms the input feature vector from the output dimension (e.g., 2048) of the external audio encoder to the latent space dimension (e.g., 4096) of the language backbone network. Then, it is processed sequentially by the layer normalization (LayerNorm) and Gaussian error linear unit (GELU) activation function layers to output an intermediate feature vector.

[0130] The second projection layer consists of a second linear transformation layer and a residual connection module. The second linear transformation layer, Linear(4096,4096), linearly transforms the intermediate feature vectors to the same latent space dimension. Then, the residual connection module adds the output of the second projection layer to the input of the first projection layer (after dimension-matched projection).

[0131] In this embodiment, the layer normalization layer can stabilize the feature distribution and accelerate convergence; the residual connection module provides a direct propagation path for gradients across layers, enabling the audio projector to remain stable during a longer training process.

[0132] In another embodiment, the method further includes: acquiring target non-native modal data, the target non-native modal data including at least one of the following: 3D point cloud data, depth image data, infrared image data, or tabular structured data; converting the target non-native modal data into a target modal feature sequence using a target modal encoder corresponding to the target non-native modal data; mapping the target modal feature sequence to the latent space of the language backbone network through a target modal projector to obtain a target modal projection feature sequence; determining an input sequence containing target modal placeholders, replacing the embedding of the target modal placeholders with the corresponding feature vector in the target modal projection feature sequence and inserting a start marker before the start position and an end marker after the end position to obtain an extended sequence, and inputting the extended sequence into the language backbone network for encoding processing to obtain a target modal evaluation embedding vector or a target modal query embedding vector.

[0133] The modal extension framework provided in this application can be extended to other non-native modalities besides audio, such as 3D point clouds, depth sensor data, infrared thermal images, and structured tabular data.

[0134] For example, in a scenario where a 3D point cloud encoder is used for 3D retrieval, a dedicated 3D point cloud projector is first designed. The structure of this projector can be designed according to the characteristics of the 3D point cloud data. For instance, a combination of multi-layer convolutional and fully connected layers can be used to map the features output by the 3D point cloud encoder to the same dimension as the latent space of the language backbone network. Simultaneously, 3D point cloud placeholders are designed. When inputting 3D point cloud features into the language backbone network, an input sequence containing these placeholders is determined. The embeddings of the 3D point cloud placeholders are replaced with the corresponding feature vectors from the 3D point cloud feature sequence mapped by the projector. A start marker is inserted before the start position and an end marker is inserted after the end position to obtain an extended sequence. Finally, the extended sequence is input into the language backbone network for encoding processing to achieve the 3D point cloud data retrieval task.

[0135] For example, in scenarios where depth / infrared retrieval is achieved by integrating a depth / infrared sensor encoder, a corresponding depth / infrared projector is designed for the encoder. Based on the characteristics of depth / infrared data, the projector can employ a neural network structure suitable for processing this type of data, such as a structure containing special filters or pooling layers, mapping the features output by the depth / infrared sensor encoder to the same dimension as the latent space of the language backbone network. Depth / infrared placeholders are designed. When inputting depth / infrared features into the language backbone network, an input sequence containing these placeholders is determined. The embeddings of the placeholders are replaced with the corresponding feature vectors from the depth / infrared feature sequence mapped by the projector. A start marker is inserted before the start position and an end marker is inserted after the end position to obtain an extended sequence. Finally, the extended sequence is input into the language backbone network for encoding processing to achieve the depth / infrared data retrieval task.

[0136] For example, in scenarios involving structured data retrieval using a tabular data encoder, a tabular data projector is designed. Considering the structured nature of tabular data, the projector can employ a structure capable of processing tabular structure information, such as processing different columns of the table separately and then fusing them, mapping the output of the tabular data encoder to the same dimension as the latent space of the language backbone network. Table data placeholders are designed. When inputting tabular data features into the language backbone network, the input sequence containing these placeholders is determined. The embeddings of the placeholders are replaced with the corresponding feature vectors from the mapped tabular data feature sequence. A start marker is inserted before the start position and an end marker is inserted after the end position to obtain an extended sequence. Finally, the extended sequence is input into the language backbone network for encoding to achieve the structured data retrieval task. By designing corresponding projectors and special markers for each new modality, this method becomes universal and can easily access different non-native modal data.

[0137] In another implementation, in some embodiments, extracting the feature vector from the last non-padding marker position from the encoded output sequence includes at least one of the following pooling strategies: Last position pooling strategy: Extract the feature vector of the last non-padded marker position from the encoded output sequence as the query embedding vector; the last non-padded marker position is the output position corresponding to the <|audio_end|> marker in the extended sequence; Average pooling strategy: Take the arithmetic mean of the feature vectors of the output positions corresponding to all audio projection features in the encoded output sequence to obtain the query embedding vector; Attention pooling strategy: Introduce a learnable attention weight vector, and perform a weighted summation of the feature vectors of the output positions corresponding to all audio projection features in the encoded output sequence to obtain the query embedding vector; Multi-scale pooling strategy: Extract feature vectors corresponding to the audio projection features from the encoding outputs of multiple different layers of the language backbone network, perform average pooling on the features of each layer to obtain multiple pooled feature vectors of different scales, and concatenate the multiple pooled feature vectors of different scales to obtain the query embedding vector.

[0138] Specifically, different pooling strategies are suitable for different audio retrieval scenarios. The principle of positional pooling is that after L layers of self-attention encoding, the output vector at the last non-filled position has already aggregated information from all time steps to its left through a causal attention mechanism, essentially a compressed representation of the global audio semantics. This strategy has the lowest computational cost and is suitable for latency-sensitive online retrieval scenarios. The principle of average pooling is to average the feature vectors of all time steps, making each time step contribute equally to the final embedding. This is suitable for scenarios where the audio content is evenly distributed in the time dimension (such as continuous ambient sound). The principle of attention pooling is to automatically identify the time step that contributes most to the semantics in the audio through learnable attention weights (such as the chirping segment being more important than the silent segment in bird calls), enabling the model to adaptively focus on keyframes. This is suitable for audio containing obvious event boundaries. The principle of multi-scale pooling is that different layers of the language backbone network capture features of different granularities. For example, shallow layers capture local acoustic features (such as phoneme-level), while deep layers capture global semantic features (such as scene-level). Concatenating and fusing multi-layer features can simultaneously preserve local details and global semantics, making it suitable for complex audio scenarios. In actual deployment, the optimal pooling strategy can be selected based on the performance of the nDCG@10 metric on the evaluation dataset, or different strategies can be adopted in different tasks.

[0139] In another embodiment, the external audio encoder includes multiple coding layers stacked sequentially. The step of training the external audio encoder, the audio projector, and the language backbone network through contrastive learning includes: in the initial stage of training, freezing the parameters of the first M coding layers of the external audio encoder, and updating only the parameters of the last N coding layers of the external audio encoder, the parameters of the audio projector, and the parameters of the language backbone network; after training reaches a preset number of steps, unfreezing the parameters of all coding layers of the external audio encoder, and continuing to jointly update the parameters of all coding layers of the external audio encoder, the parameters of the audio projector, and the parameters of the language backbone network; wherein, M+N equals the total number of coding layers of the external audio encoder.

[0140] Specifically, to balance the general audio understanding capabilities of the audio encoder with the adaptation effect to the retrieval task during training, while reducing the number of training parameters, a layer-wise freezing strategy is adopted. The core principle of the layer-wise freezing strategy is that the lower coding layers of the external audio encoder (such as the first 16 layers) have already learned general audio feature extraction capabilities (such as spectral texture, fundamental frequency and harmonic structure, etc.) during large-scale audio pre-training. These features have strong universality and are still applicable in the retrieval task, so they do not need to be relearned. On the other hand, the higher coding layers (such as the last 16 layers) capture more specific semantic features related to the pre-training task and need to be adapted to the semantic alignment target of the retrieval task. By fine-tuning only the higher coding layers in the early stage of training (e.g., freezing the first 16 layers and training only the last 16 layers), the number of trainable parameters can be significantly reduced, thus lowering the memory requirements and training time, while avoiding the destruction of the underlying general features due to large updates in the early stage of training. After training reaches a preset number of steps (e.g., 30% of the total training steps), the low-level features have been adapted to a certain extent through backpropagation of high-level gradients. At this point, all encoding layers are unfrozen for joint fine-tuning training, further aligning the low-level features with the semantic space of the retrieval task. During the inference phase, all encoding layers of the external audio encoder participate in forward computation without affecting inference efficiency.

[0141] The embodiments of this application have at least the following beneficial effects: I. Achieving Unified Cross-Modal Retrieval: This application enables the mapping of data from multiple modalities, such as text, images, videos, visual documents, and audio, to the same vector space, achieving unified cross-modal retrieval and eliminating the semantic gap between different modalities in traditional methods. Users only need to deploy one model to complete retrieval tasks for all modalities, eliminating the need to maintain multiple independent single-modal retrieval systems, significantly reducing system complexity and operational costs.

[0142] II. Low-cost capability expansion: This application achieves 8B-level (approximately 8 billion parameters) visual capabilities by adding only 18.9M parameters to the audio projector, less than 0.2% of the total model parameters. The language embedding model adds audio retrieval capabilities. Compared to a fully modal model trained from scratch, this application reduces training costs by several orders of magnitude and achieves efficient modality expansion.

[0143] III. Plug-and-play audio encoder: This application extracts an external audio encoder from a full-modal large language model. This encoder has been tested on hundreds of thousands of audio tracks covering various audio types such as speech, music, and ambient sound. The text pairs are fully pre-trained, providing out-of-the-box audio understanding capabilities. No pre-training from scratch is required, enabling plug-and-play audio encoder integration.

[0144] IV. High training efficiency: The training data for the audio-related modules in this application only requires approximately 180,000 audio files. Text-paired data allows for short training times and low resource consumption. Furthermore, the use of a low-rank adapter (LoRA) for efficient parameter fine-tuning further reduces the number of trainable parameters, enabling a single training session to be completed within hours.

[0145] V. High Scalability: The modal access method provided in this application is universal and can be extended to non-native modalities such as 3D point clouds, depth / infrared sensors, and tabular data. When adding a new modality, only the corresponding projector and special markers need to be designed, without modifying the core architecture of the model, thus standardizing and templating the modal extension process.

[0146] VI. End-to-end differentiability: The computation graph between the audio encoder, audio projector and language backbone network in this application is fully differentiable, supporting end-to-end gradient backpropagation, which enables contrastive learning training to optimize all components simultaneously, thereby achieving globally optimal cross-modal retrieval performance.

[0147] All of the above technical solutions can be combined in any way to form optional embodiments of this application, and will not be described in detail here.

[0148] This application embodiment receives query data; when the query data includes audio data, it uses an external audio encoder extracted from a full-modal large language model to convert the audio data into an audio feature sequence; it then maps the audio feature sequence to the latent space of the language backbone network using an audio projector to obtain an audio projection feature sequence; it determines an input sequence based on the query data, the input sequence containing at least one audio placeholder; it replaces the embedding of the audio placeholder in the input sequence with the corresponding feature vector in the audio projection feature sequence, inserts a start marker before the start position and an end marker after the end position of the audio projection feature sequence to obtain an extended sequence; it inputs the extended sequence into the language backbone network for encoding processing to obtain a query embedding vector corresponding to the query data; and it outputs the retrieval result based on the query embedding vector. This application embodiment, by utilizing an external audio encoder in a full-modal large language model, in conjunction with an audio projector, effectively converts audio data into an audio projection feature sequence that can be processed by the language backbone network, and by combining the replacement and insertion of audio placeholders, achieves efficient encoding processing of multimodal query data, obtains accurate query embedding vectors to output retrieval results, and significantly improves the efficiency and accuracy of multimodal data retrieval containing audio.

[0149] To facilitate better implementation of the data processing method of this application embodiment, this application embodiment also provides a data processing apparatus. Please refer to... Figure 6 , Figure 6 This is a schematic diagram of the structure of a data processing apparatus provided in an embodiment of this application. The data processing apparatus 200 may include: The receiving unit 210 is used to receive query data; The conversion unit 220 is used to convert the audio into an audio feature sequence using an external audio encoder extracted from a full-modal large language model when the query data contains audio data. The mapping unit 230 is used to map the audio feature sequence to the latent space of the language backbone network through the audio projector to obtain the audio projection feature sequence. Determining unit 240 is configured to determine an input sequence based on the query data, the input sequence containing at least one audio placeholder; Processing unit 250 is used to replace the embedding of the audio placeholder in the input sequence with the corresponding feature vector in the audio projection feature sequence, and insert a start marker before the start position and an end marker after the end position of the audio projection feature sequence to obtain an extended sequence; Encoding unit 260 is used to input the extended sequence into the language backbone network for encoding processing to obtain the query embedding vector corresponding to the query data; The retrieval unit 270 is used to output retrieval results based on the query embedding vector.

[0150] In some embodiments, the audio projector is used to transform the dimension of the audio feature sequence from the output dimension of the external audio encoder to the latent space dimension of the language backbone network; the audio projector includes a first projection layer and a second projection layer, the first projection layer includes a first linear transformation layer and a Gaussian error linear unit, and the second projection layer includes a second linear transformation layer; the first projection layer is used to linearly transform each feature vector of the audio feature sequence from the output dimension of the external audio encoder to the latent space dimension of the language backbone network through the first linear transformation layer, and then process it through an activation function layer to output an intermediate feature sequence; the second projection layer is used to linearly transform each feature vector in the intermediate feature sequence through the second linear transformation layer to output the audio projected feature sequence, wherein the input dimension and output dimension of the second linear transformation layer are both the latent space dimension of the language backbone network. In some embodiments, the mapping unit 230 is configured to: sequentially input each feature vector in the audio feature sequence into the first linear transformation layer for linear transformation, and process it through the activation function layer to output the corresponding intermediate feature vector in the intermediate feature sequence; input the intermediate feature vector into the second linear transformation layer for linear transformation to output the corresponding audio projection feature vector in the audio projection feature sequence.

[0151] In some embodiments, the processing unit 250 is further configured to: insert one or more padding markers after the end marker and at the end of the extended sequence when the length of the extended sequence is less than a preset length, so that the length of the extended sequence reaches the preset length.

[0152] In some embodiments, the encoding unit 260 is configured to: superimpose positional encoding on the vectors of each position in the extended sequence to obtain a positionally encoded extended sequence; input the positionally encoded extended sequence into the language backbone network and encode it through the self-attention mechanism of the language backbone network to obtain an encoded output sequence; extract the feature vector of the last non-filled marker position from the encoded output sequence and perform normalization processing to obtain the query embedding vector.

[0153] In some embodiments, the determining unit 240 is further configured to: when the query data also contains text data, add a text placeholder corresponding to the text data to the input sequence, wherein the embedding of the text placeholder is directly obtained from the vocabulary of the language backbone network; and / or when the query data also contains image data and / or video data, add a visual placeholder corresponding to the image data and / or video data to the input sequence, wherein the visual placeholder is used to replace a visual feature sequence extracted from the image data and / or video data.

[0154] In some implementations, when the query data also includes image data and / or video data, the processing unit 250 is further configured to: convert the image data and / or video data into a visual feature sequence using a visual encoder; and when generating the extended sequence, replace the embedding of the visual placeholder in the input sequence with the corresponding feature vector in the visual feature sequence; wherein the visual encoder is a pre-trained visual encoder in the visual-language embedding model corresponding to the language backbone network.

[0155] In some embodiments, the encoding unit 260 is configured to: input the extended sequence into the language backbone network, the extended sequence containing features of different modalities, the features of different modalities including: the audio projection feature sequence, and at least one of the embedding of the text placeholder and the visual feature sequence; jointly encode the features of different modalities in the extended sequence through the self-attention mechanism of the language backbone network, so that the features of different modalities are fused together and aligned to the same semantic space during the encoding process, to obtain an encoded output sequence; extract the feature vector of the last non-filler mark position from the encoded output sequence and perform normalization processing to obtain the query embedding vector.

[0156] In some embodiments, the retrieval unit 270 is configured to: calculate the similarity between the query embedding vector and each candidate vector in the candidate library, the candidate library containing pre-stored embedding vectors corresponding to at least one modality among images, videos, documents, audio, and text; select at least one candidate vector according to the similarity from high to low; and output the candidate data corresponding to the candidate vector as the retrieval result.

[0157] In some embodiments, the data processing apparatus 200 further includes a training unit, configured to: acquire audio query samples and text description samples paired with the audio query samples; convert the audio query samples into audio training feature sequences using the external audio encoder; map the audio training feature sequences to the latent space of the language backbone network using the audio projector to obtain audio training projection feature sequences; determine a training input sequence based on the audio query samples, the training input sequence containing at least one audio placeholder; replace the embedding of the audio placeholder in the training input sequence with the corresponding feature vector in the audio training projection feature sequence, and apply the feature vector to the audio training projection feature sequence. A start marker is inserted before the start position and an end marker is inserted after the end position of the shadow feature sequence to obtain a training expansion sequence. The training expansion sequence is input into the language backbone network for encoding to obtain an audio training embedding vector. The text description sample is converted into a text tag sequence, and the text tag sequence is input into the language backbone network for encoding to obtain a text training embedding vector. Based on the audio training embedding vector and the text training embedding vector, the external audio encoder, the audio projector, and the language backbone network are trained through contrastive learning to align the audio training embedding vector and the text training embedding vector in the semantic space of the language backbone network.

[0158] In some embodiments, the training unit is configured to train the external audio encoder, the audio projector, and the language backbone network through contrastive learning based on the audio training embedding vector and the text training embedding vector, comprising: for each audio training embedding vector in the current batch, calculating the similarity between the audio training embedding vector and all text training embedding vectors in the current batch, wherein all text training embedding vectors include positive sample embedding vectors of text description samples paired with the audio training embedding vector and negative sample embedding vectors corresponding to other audio training embedding vectors; dividing the similarity between the audio training embedding vector and its positive sample embedding vector by the sum of the similarities between the audio training embedding vector and all text training embedding vectors in the current batch to obtain a ratio; taking the logarithm and negative of the ratio to obtain the sub-loss corresponding to the audio training embedding vector; averaging the sub-losses corresponding to all audio training embedding vectors in the current batch to obtain the contrastive loss; updating the parameters of the audio projector, and updating the parameters of the language backbone network and the external audio encoder according to the contrastive loss.

[0159] In some embodiments, the training unit is further configured to: update the parameters of the visual encoder according to the contrastive loss during training via contrastive learning; wherein the visual encoder is jointly optimized with the language backbone network, the audio projector, and the external audio encoder; and the visual encoder is configured to convert image data and / or video data into visual feature sequences.

[0160] In some embodiments, the data processing device 200 further includes an evaluation unit, configured to: acquire an audio retrieval evaluation dataset, the audio retrieval evaluation dataset comprising text-to-audio retrieval tasks and audio-to-text retrieval tasks, each query data in each task pre-configured with at least one correct retrieval result, and the evaluation dataset also pre-stores candidate embedding vectors corresponding to each candidate result in a candidate library; for each text query data in the text-to-audio retrieval task, converting the text query data into a text tag sequence, inputting the text tag sequence into the language backbone network for encoding processing to obtain a text evaluation embedding vector; for each audio query data in the audio-to-text retrieval task, using the external audio encoder to convert the audio query data into an audio feature sequence, and projecting the audio through the audio projector. The feature sequence is mapped to the latent space of the language backbone network to obtain the audio projection feature sequence. An input sequence containing audio placeholders is determined. The embedding of the audio placeholders is replaced with the corresponding feature vector in the audio projection feature sequence, and a start marker is inserted before the start position and an end marker is inserted after the end position to obtain an extended sequence. The extended sequence is input into the language backbone network for encoding processing to obtain the audio evaluation embedding vector. The similarity between the text evaluation embedding vector or the audio evaluation embedding vector and each candidate embedding vector in the candidate library is calculated respectively. The candidate results are sorted from high to low according to the similarity to obtain a first retrieval result list. The first retrieval result list is compared with the correct retrieval results corresponding to the corresponding query data, and the audio retrieval evaluation result is output according to the standardized loss cumulative gain index.

[0161] In some embodiments, the evaluation unit is further configured to: acquire a visual retrieval evaluation dataset, the visual retrieval evaluation dataset including image retrieval tasks and visual document retrieval tasks, each visual query data in each task being pre-configured with at least one correct retrieval result, and the evaluation dataset also pre-storing candidate embedding vectors corresponding to each candidate result in the candidate library; for each visual query data in the image retrieval task or visual document retrieval task, using the visual encoder to convert the visual query data into a visual feature sequence, determining an input sequence containing visual placeholders, replacing the embedding of the visual placeholders with the corresponding feature vector in the visual feature sequence and inserting a start marker before the start position and an end marker after the end position to obtain an extended sequence, inputting the extended sequence into the language backbone network for encoding processing to obtain a visual evaluation embedding vector; calculating the similarity between the visual evaluation embedding vector and each candidate embedding vector in the candidate library, sorting the candidate results from high to low similarity to obtain a second retrieval result list; comparing the second retrieval result list with the correct retrieval results corresponding to the corresponding query data, and outputting the visual retrieval evaluation result according to the standardized loss cumulative gain index.

[0162] It should be noted that the functions of each module in the data processing device 200 in this application embodiment can be referred to the specific implementation of any embodiment in the above method embodiments, and will not be repeated here.

[0163] Each unit in the above-described device can be implemented entirely or partially through software, hardware, or a combination thereof. Each unit can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each unit.

[0164] For example, the data processing device 200 may be integrated into a terminal or server that has storage and a processor and thus computing power, or the data processing device 200 may be the terminal or server.

[0165] In some embodiments, this application also provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0166] Figure 7 A schematic diagram of the structure of the computer device provided in the embodiments of this application, such as... Figure 7As shown, the computer device 300 may include: a communication interface 301, a memory 302, a processor 303, and a communication bus 304. The communication interface 301, memory 302, and processor 303 communicate with each other via the communication bus 304. The communication interface 301 is used for data communication between the device 300 and external devices. The memory 302 can be used to store software programs and modules, and the processor 303 runs the software programs and modules stored in the memory 302, such as the software programs for the corresponding operations in the aforementioned method embodiments.

[0167] In some embodiments, the processor 303 may invoke software programs and modules stored in the memory 302 to perform the following operations: receiving query data; when the query data contains audio data, converting the audio data into an audio feature sequence using an external audio encoder extracted from a full-modal large language model; mapping the audio feature sequence to the latent space of a language backbone network using an audio projector to obtain an audio projection feature sequence; determining an input sequence based on the query data, the input sequence containing at least one audio placeholder; replacing the embedding of the audio placeholder in the input sequence with the corresponding feature vector in the audio projection feature sequence, and inserting a start marker before the start position and an end marker after the end position of the audio projection feature sequence to obtain an extended sequence; inputting the extended sequence into the language backbone network for encoding processing to obtain a query embedding vector corresponding to the query data; and outputting retrieval results based on the query embedding vector. In some embodiments, the computer device 300 may be integrated into a terminal or server that has storage and a processor and thus computing power, or the computer device 300 may be the terminal or server.

[0168] This application also provides a computer-readable storage medium for storing a computer program. This computer-readable storage medium can be applied to a computer device, and the computer program causes the computer device to execute the corresponding processes in the methods described above in the embodiments of this application; for brevity, further details are omitted here.

[0169] This application also provides a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the corresponding processes in the methods described above in the embodiments of this application. For brevity, these details will not be elaborated further here.

[0170] This application also provides a computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the corresponding processes in the methods described above in the embodiments of this application. For brevity, these details will not be elaborated further here.

[0171] It should be understood that the processor in the embodiments of this application may be an integrated circuit chip with signal processing capabilities. In implementation, the steps of the above method embodiments can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor described above can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0172] It is understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0173] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0174] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0175] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0176] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0177] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0178] In addition, the functional units in the embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0179] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer or a server) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0180] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A data processing method, characterized in that, The method includes: Receive query data; When the query data contains audio data, the audio data is converted into an audio feature sequence using an external audio encoder extracted from a full-modal large language model; The audio feature sequence is mapped to the latent space of the language backbone network through an audio projector to obtain the audio projection feature sequence. An input sequence is determined based on the query data, the input sequence containing at least one audio placeholder; The embedded audio placeholders in the input sequence are replaced with the corresponding feature vectors in the audio projection feature sequence, and a start marker is inserted before the start position and an end marker is inserted after the end position to obtain the extended sequence. The extended sequence is input into the language backbone network for encoding processing to obtain the query embedding vector corresponding to the query data; The search results are output based on the query embedding vector.

2. The data processing method as described in claim 1, characterized in that, The audio projector is used to transform the dimension of the audio feature sequence from the output dimension of the external audio encoder to the latent space dimension of the language backbone network; The audio projector includes a first projection layer and a second projection layer. The first projection layer includes a first linear transformation layer and a Gaussian error linear unit, and the second projection layer includes a second linear transformation layer. The first projection layer is used to linearly transform each feature vector of the audio feature sequence from the output dimension of the external audio encoder to the latent space dimension of the language backbone network through the first linear transformation layer, and then process it through the activation function layer to output the intermediate feature sequence. The second projection layer is used to perform a linear transformation on each feature vector in the intermediate feature sequence through the second linear transformation layer to output the audio projection feature sequence, wherein the input dimension and output dimension of the second linear transformation layer are both the latent space dimensions of the language backbone network.

3. The data processing method as described in claim 2, characterized in that, The step of mapping the audio feature sequence to the latent space of the language backbone network through an audio projector to obtain an audio projection feature sequence includes: Each feature vector in the audio feature sequence is sequentially input into the first linear transformation layer for linear transformation, and then processed by the activation function layer to output the corresponding intermediate feature vector in the intermediate feature sequence. The intermediate feature vector is input into the second linear transformation layer for linear transformation, and the corresponding audio projection feature vector in the audio projection feature sequence is output.

4. The data processing method as described in claim 1, characterized in that, The process of obtaining the extended sequence further includes: When the length of the extended sequence is less than the preset length, one or more padding markers are inserted after the end marker and at the end of the extended sequence to make the length of the extended sequence reach the preset length.

5. The data processing method as described in claim 4, characterized in that, The step of inputting the extended sequence into the language backbone network for encoding processing to obtain the query embedding vector corresponding to the query data includes: The vectors at each position in the extended sequence are superimposed with position codes to obtain the position-coded extended sequence; The position-encoded extended sequence is input into the language backbone network and encoded through the self-attention mechanism of the language backbone network to obtain the encoded output sequence. The feature vector of the last non-padded marker position is extracted from the encoded output sequence and normalized to obtain the query embedding vector.

6. The data processing method as described in claim 1, characterized in that, The step of determining the input sequence based on the query data further includes: When the query data also contains text data, text placeholders corresponding to the text data are added to the input sequence, and the embedding of the text placeholders is directly obtained from the vocabulary of the language backbone network; and / or When the query data also includes image data and / or video data, visual placeholders corresponding to the image data and / or video data are added to the input sequence, and the visual placeholders are used to replace the visual feature sequences extracted from the image data and / or video data.

7. The data processing method as described in claim 6, characterized in that, When the query data also includes image data and / or video data, the method further includes: The image data and / or video data are converted into a visual feature sequence using a visual encoder; When generating the extended sequence, the embedding of the visual placeholder in the input sequence is replaced with the corresponding feature vector in the visual feature sequence; The visual encoder is a pre-trained visual encoder in the visual-language embedding model corresponding to the language backbone network.

8. The data processing method as described in claim 6, characterized in that, The step of inputting the extended sequence into the language backbone network for encoding processing to obtain the query embedding vector corresponding to the query data includes: The extended sequence is input into the language backbone network. The extended sequence contains features of different modalities, including at least one of the following: the audio projection feature sequence, the embedding of the text placeholder, and the visual feature sequence. The self-attention mechanism of the language backbone network is used to jointly encode the features of different modalities in the extended sequence, so that the features of different modalities are fused together and aligned to the same semantic space during the encoding process, resulting in an encoded output sequence. The feature vector of the last non-padded marker position is extracted from the encoded output sequence and normalized to obtain the query embedding vector.

9. The data processing method as described in claim 1, characterized in that, The step of outputting retrieval results based on the query embedding vector includes: Calculate the similarity between the query embedding vector and each candidate vector in the candidate library, which contains pre-stored embedding vectors corresponding to at least one modality among images, videos, documents, audio, and text; At least one candidate vector is selected from high to low similarity, and the candidate data corresponding to the candidate vector is output as the retrieval result.

10. The data processing method according to any one of claims 1 to 9, characterized in that, The method further includes: Obtain audio query samples and text description samples that are paired with the audio query samples; The external audio encoder is used to convert the audio query samples into audio training feature sequences; The audio training feature sequence is mapped to the latent space of the language backbone network through the audio projector to obtain the audio training projection feature sequence. A training input sequence is determined based on the audio query sample, and the training input sequence contains at least one audio placeholder. The embedding of the audio placeholder in the training input sequence is replaced with the corresponding feature vector in the audio training projection feature sequence, and a start marker is inserted before the start position and an end marker is inserted after the end position to obtain the training expansion sequence; The training extended sequence is input into the language backbone network for encoding processing to obtain the audio training embedding vector; The text description sample is converted into a text tag sequence, and the text tag sequence is input into the language backbone network for encoding processing to obtain the text training embedding vector; Based on the audio training embedding vector and the text training embedding vector, the external audio encoder, the audio projector, and the language backbone network are trained through contrastive learning, so that the audio training embedding vector and the text training embedding vector are aligned in the semantic space of the language backbone network.

11. The data processing method as described in claim 10, characterized in that, The step of training the external audio encoder, the audio projector, and the language backbone network through contrastive learning based on the audio training embedding vector and the text training embedding vector includes: For each audio training embedding vector in the current batch, calculate the similarity between the audio training embedding vector and all text training embedding vectors in the current batch, where all text training embedding vectors include positive sample embedding vectors of text description samples paired with the audio training embedding vector and negative sample embedding vectors corresponding to other audio training embedding vectors. The similarity between the audio training embedding vector and its positive sample embedding vector is divided by the sum of the similarities between the audio training embedding vector and all text training embedding vectors in the current batch to obtain the ratio. Taking the negative of the logarithm of the ratio yields the sub-loss corresponding to the audio training embedding vector; The contrastive loss is obtained by averaging the sub-losses corresponding to all audio training embedding vectors in the current batch. The parameters of the audio projector, the language backbone network, and the external audio encoder are updated based on the contrast loss.

12. The data processing method as described in claim 10, characterized in that, The method further includes: During training via contrastive learning, the parameters of the visual encoder are updated based on the contrastive loss; wherein the visual encoder is jointly optimized with the language backbone network, the audio projector, and the external audio encoder; the visual encoder is used to convert image data and / or video data into visual feature sequences.

13. The data processing method according to any one of claims 1 to 9, characterized in that, The method further includes: Obtain an audio retrieval evaluation dataset, which includes text-to-audio retrieval tasks and audio-to-text retrieval tasks. Each query in each task is pre-configured with at least one correct retrieval result. The evaluation dataset also pre-stores candidate embedding vectors corresponding to each candidate result in the candidate library. For each text query data in the text-to-audio retrieval task, the text query data is converted into a text tag sequence, and the text tag sequence is input into the language backbone network for encoding processing to obtain a text evaluation embedding vector; For each audio query data in the audio-to-text retrieval task, the external audio encoder converts the audio query data into an audio feature sequence. The audio projector maps the audio feature sequence to the latent space of the language backbone network to obtain an audio projection feature sequence. An input sequence containing audio placeholders is determined. The embedding of the audio placeholders is replaced with the corresponding feature vector in the audio projection feature sequence. A start marker is inserted before the start position and an end marker is inserted after the end position to obtain an extended sequence. The extended sequence is input into the language backbone network for encoding processing to obtain an audio evaluation embedding vector. Calculate the similarity between the text evaluation embedding vector or the audio evaluation embedding vector and each candidate embedding vector in the candidate library, and sort the candidate results from high to low similarity to obtain the first search result list; The first list of search results is compared with the correct search results corresponding to the query data, and the audio search evaluation results are output according to the standardized loss cumulative gain index.

14. The data processing method as described in claim 13, characterized in that, The method further includes: Obtain a visual retrieval evaluation dataset, which includes image retrieval tasks and visual document retrieval tasks. Each visual query data in each task is pre-configured with at least one correct retrieval result. The evaluation dataset also pre-stores candidate embedding vectors corresponding to each candidate result in the candidate library. For each visual query data in the image retrieval task or visual document retrieval task, the visual encoder is used to convert the visual query data into a visual feature sequence, an input sequence containing visual placeholders is determined, the embedding of the visual placeholders is replaced with the corresponding feature vector in the visual feature sequence, a start marker is inserted before the start position and an end marker is inserted after the end position to obtain an extended sequence, and the extended sequence is input into the language backbone network for encoding processing to obtain a visual evaluation embedding vector; Calculate the similarity between the visual evaluation embedding vector and each candidate embedding vector in the candidate library, sort the candidate results from high to low similarity, and obtain a second search result list; The second search result list is compared with the correct search results corresponding to the query data, and the visual search evaluation results are output according to the standardized loss cumulative gain index.

15. A data processing apparatus, characterized in that, The device includes: The receiving unit is used to receive query data; A conversion unit is used to convert audio into an audio feature sequence using an external audio encoder extracted from a full-modal large language model when the query data contains audio data. The mapping unit is used to map the audio feature sequence to the latent space of the language backbone network through the audio projector to obtain the audio projection feature sequence; A determining unit is configured to determine an input sequence based on the query data, the input sequence containing at least one audio placeholder; The processing unit is configured to replace the embedding of the audio placeholder in the input sequence with the corresponding feature vector in the audio projection feature sequence, and insert a start marker before the start position and an end marker after the end position of the audio projection feature sequence to obtain an extended sequence; An encoding unit is used to input the extended sequence into the language backbone network for encoding processing to obtain the query embedding vector corresponding to the query data. The retrieval unit is used to output retrieval results based on the query embedding vector.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted for loading by a processor to perform the data processing method as described in any one of claims 1 to 14.

17. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing a computer program, and the processor executing the data processing method as described in any one of claims 1 to 14 by calling the computer program stored in the memory.

18. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the data processing method according to any one of claims 1 to 14.