Cross-language speech text retrieval method based on pre-trained automatic speech recognition model

By adopting a dual-tower network structure and contrast learning strategy based on pre-trained automatic speech recognition model in the speech text retrieval system, the shortcomings of the existing system in model architecture and multimodal information modeling capabilities are solved, and efficient cross-language speech text retrieval is achieved, especially in multilingual and low-resource language scenarios.

CN120086354APending Publication Date: 2025-06-03SHANDONG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510257096.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing speech text retrieval system has shortcomings in model architecture, training data scale and multimodal information modeling capabilities, resulting in poor retrieval performance, especially in multilingual environments, which leads to low retrieval accuracy.

Method used

A cross-language speech text retrieval method based on pre-trained automatic speech recognition model is adopted. By removing complex cross-attention layers, adding pooling layers and projection layers, a dual-tower network structure is formed, and a comparison learning strategy is introduced to reduce the complexity of the model and training costs, while a low-rank adaptive fine-tuning method is introduced to reduce the amount of training parameters.

Benefits of technology

It significantly improves the performance and efficiency of cross-language speech text retrieval, reduces computing and storage overhead, enhances the cross-modal and multilingual retrieval capabilities of the model, and performs well in low-resource language scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086354A_ABST
    Figure CN120086354A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-language speech text retrieval method based on a pre-trained automatic speech recognition model. According to the method, a pre-trained automatic speech recognition model is expanded to a speech text retrieval system, and the model is finely adjusted in combination with comparative learning and a low-rank adaptive method, so that an efficient cross-language speech text retrieval function is realized. The method is based on an encoder-decoder structure initialized by a pre-training model, and comprises the following steps: firstly, converting voice data and text data into high-dimensional feature vectors through an encoding module respectively, and mapping the high-dimensional feature vectors to a unified embedding space; the model then minimizes the matching speech and text embedding distance in the embedding space. And finally, through a similarity matching algorithm, the model can efficiently match the query voice with the text in the text library, thereby returning the most relevant text data. Experimental results show that the retrieval precision and efficiency of the method on a test data set are close to or superior to those of an existing public model, and it is proved that the method has wide application prospects and remarkable practical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of deep learning natural language processing and signal processing, and particularly relates to a cross - language speech - text retrieval method based on a pre - trained automatic speech recognition model. Technical Background

[0002] With the rapid development of speech technology and natural language processing technology, speech - text retrieval has been widely applied in practical scenarios such as intelligent voice assistants, multilingual translation, and cross - modal retrieval. By mapping speech and text to the same embedding space, speech - text retrieval technology can achieve efficient cross - modal information retrieval. However, the performance and application scope of existing speech - text retrieval systems are limited to a certain extent by the model architecture, the scale of training data, and the multi - modal information modeling ability.

[0003] Existing retrieval models usually use single - modal pre - trained models (such as models only for text or speech) as initialization. These models lack the optimization ability for cross - modal speech - text matching, and it is difficult to efficiently capture the semantic relationship between speech and text, resulting in poor retrieval performance. Some methods attempt to improve performance through task - specific models, but fully training such models from scratch requires a large amount of computing resources and storage space, which is difficult to meet the actual application requirements.

[0004] In addition, most current cross - modal retrieval methods fail to effectively integrate global and local features. Especially in a multilingual environment, the ability to model semantic alignment is limited, resulting in a low retrieval precision. For example, some existing methods attempt to achieve cross - modal retrieval by simply aligning speech and text features, but lack sufficient semantic modeling and efficient contrast learning strategies, and cannot perform well in multilingual and multi - modal scenarios. In addition, the high complexity of the models makes them require high - level hardware support during deployment, which further limits their application scope. Summary of the Invention

[0005] In view of the above problems, the present invention proposes a cross - language speech - text retrieval method based on a pre - trained automatic speech recognition model. By removing complex cross - attention layers, adding pooling layers and projection layers to form a two - tower network structure, and introducing a contrast learning strategy, the present invention significantly reduces the model complexity and training cost, while improving the retrieval performance in multilingual and multi - modal scenarios. The present invention also further reduces the number of training parameters through a low - rank adaptation fine - tuning method, solving the deficiencies of the prior art in retrieval accuracy, generality, and deployment efficiency.

[0006] The technical solution adopted by the present invention is as follows:

[0007] A cross - language speech - text retrieval method based on a pre - trained automatic speech recognition model. This method uses an automatic speech recognition model trained on audio - text pairs to initialize the network structure. By removing the complex cross - attention mechanism in the pre - trained model, the connection between the speech encoder and the text decoder is decoupled, a two - tower structure is constructed, and a pooling layer and a projection layer are added. An efficient speech - text retrieval system is trained by fine - tuning. The specific steps are as follows:

[0008] Step 1: Construct speech and text embeddings: Prepare a speech - text dataset containing N pairs of paired cross - language speech and text data, denoted as where x i represents speech data, and y i represents the corresponding text data. Use the audio encoding module of the pre - trained model to convert the speech data x into a 128 - channel Mel - spectrum vector, denoted as where T represents the number of time frames. Use the text encoding module to encode the text data into text token embeddings, denoted as where d m is the dimension of the model, and M is the number of text tokens. In this way, N pairs of speech and text embedding pairs {X i , Y i} i=1:N ;

[0009] Step 2: Initialize the network structure: Use the pre - trained automatic speech recognition model as the base model, remove the cross - attention layer in the text decoder, and further decouple the connection between the audio encoder and the text decoder, enabling the model to independently process speech and text data. At the same time, in order to map speech and text data to a unified embedding space, a pooling layer and an audio projection layer are added after the output of the audio encoder, and a text projection layer is added after the output of the text decoder. In addition, freeze the pre - trained weights of the model to make full use of the rich prior knowledge of the pre - trained model, and introduce a low - rank adaptation layer in the self - attention layer of the text decoder to fine - tune the model, enabling the model to better adapt to the speech - text retrieval task;

[0010] Step 3: Extract features and map them to a unified embedding space: Input the Mel - spectrum vector X obtained in Step 1 into the audio encoder to extract features. Each encoder block contains a self - attention layer with 20 self - attention heads and a two - layer multi - layer perceptron with a Gaussian error linear unit as the activation function. The output of the encoder is processed by the pooling layer and the audio projection layer and then mapped to a unified embedding space. Input the text token embedding Y obtained in Step 1 into the text decoder to extract features. The text decoder has the same structure as the audio encoder except for the low - rank adaptation layer. The output of the decoder is processed by the text projection layer and then mapped to a unified embedding space;

[0011] Step 4: Train the cross - language speech - text retrieval model: In the embedding space, use the contrastive loss function to minimize the distance between similar speech and text embeddings and maximize the distance between dissimilar speech and text embeddings. The formula for the loss function is as follows:

[0012]

[0013] where represents the i - th speech - text embedding pair, τ is a learnable temperature parameter for scaling the loss, N is the number of speech - text embedding pairs. The training process is optimized using the stochastic gradient descent algorithm, and the model parameters are updated through backpropagation to achieve the training goal of minimizing the loss function, making the embeddings of speech and text closer in the unified space;

[0014] Step 5: Inference of the cross - language speech - text retrieval system: Taking speech retrieval of text as an example, when the user inputs any speech, the model uses the audio encoder to obtain the speech embedding Then calculate the similarity between the speech embedding and all text embeddings in the text database z db The similarity calculation uses cosine similarity:

[0015]

[0016] where ||·|| 2 is the L2 norm, used to standardize the embedding vector. Finally, according to the calculated similarity, return the text result most relevant to the speech for the user to view.

[0017] Preferably, in the above - mentioned step 1, the speech - text dataset should contain paired speech - text data in 21 languages in the CoVoST - 2 dataset.

[0018] Preferably, in the above - mentioned step 1, for the text data y, the model first appends a special token <|eot|> after each text token sequence to obtain y′ = [<|TextTokens|>,<|eot|>], and then encodes the text token y′ into a text token embedding, denoted as where d m is the dimension of the model, and M is the number of text tokens.

[0019] Preferably, in the above - mentioned step 2, the added pooling layer adopts the attention pooling method. Specifically, after the Mel - spectrogram vector C is extracted by the audio encoder to obtain a two - dimensional hidden feature where d m is the dimension of the model, T represents the number of time frames. First, perform average pooling on H speech along the T dimension to obtain In the attention pooling calculation of the pooling layer, h′ speech is used as the query, H speech itself is used as the key and value, and the scaled dot-product attention calculation is performed to obtain a one-dimensional speech embedding

[0020] h speech = CrossAttn(query = h′ speech ; key, value = H speech ),

[0021] where CrossAttn(·) is the attention calculation, and the three parameters query, key, and value are the query, key, and value in the attention calculation respectively.

[0022] Preferably, in step 2, the added projection layer adopts a linear transformation operation.

[0023] Preferably, in step 3, the text embedding Y will extract features through a text decoder, and finally obtain the text embedding

[0024] h text = TextDecoder(Y)[:, -1],

[0025] where TextDecoder(·) is the operation of extracting features by the text decoder.

[0026] Preferably, in step 4, during the training process, the parameters of the encoder and decoder are frozen, and only the parameters of the low-rank adaptation layer, audio pooling layer, and projection layer are trained.

[0027] The beneficial effects of the present invention are as follows:

[0028] 1. Make full use of the rich speech-text prior knowledge of the pre-trained automatic speech recognition model. Existing speech-text retrieval systems all use pure text pre-trained models to initialize the network structure and lack paired speech-text data. The Whisper model is trained based on 6.8 million hours of paired speech and text data. It has very excellent performance in tasks such as speech recognition, machine translation, and zero-shot generalization, and is currently an automatic speech recognition model with excellent performance. The present invention selects the Whisper-large-v3 version with the largest number of parameters and the best performance in the Whisper model to initialize the network, so that the present model has rich prior knowledge, can effectively capture the matching relationship between speech and text, significantly improve the matching accuracy between the two, especially in multi-language and low-resource language scenarios, the alignment between audio and text is more accurate and the performance is more prominent.

[0029] 2. Reduce computational and storage overheads. Existing cross-modal retrieval methods, especially task-specific models, usually contain a large number of parameters and high computational requirements, resulting in excessive computational and storage overheads during the training and inference processes. By removing redundant cross-attention layers, the present invention not only decouples the connection between the encoder and the decoder but also significantly reduces the computational amount of the model. In addition, the present invention also introduces the low-rank adaptation fine-tuning technology, enabling the model to efficiently operate with limited computational resources and storage space while ensuring rich prior knowledge.

[0030] 3. Have strong cross-modal and multilingual speech-text retrieval capabilities. Through a unified audio and text embedding space, the model of the present invention can better adapt to multi-modal and multilingual tasks. This model is only trained on the CoVoST-2 dataset containing 21 languages and still shows good performance metrics on the FLEURS dataset containing 102 languages. Compared with the prior art, the present invention performs excellently in cross-modal multilingual retrieval tasks, can handle datasets with up to 102 languages, and maintains stable accuracy in multi-modal tasks. Especially in the performance of low-resource languages, it exceeds the accuracy level of existing methods, verifying the superiority of the present invention in dealing with multi-language and low-resource language scenarios.

[0031] 4. Have strong capabilities for speech retrieval text translation. Most traditional speech retrieval text translation models rely on automatic speech recognition methods, first transcribing speech into text in the corresponding language, and then retrieving the translation text in the target language through the transcribed text. Such a process is prone to information loss and error transmission, especially when the speech recognition accuracy is not high. Moreover, current research on speech retrieval text translation is not sufficient. The present invention randomly extracts training data from the CoVoST-2 dataset in the combination of "75% transcribed text + 25% translation text" to train a speech direct retrieval text translation model. Compared with the model not trained on translation text, the BLEU index has been significantly improved. The model can directly compare and match audio and translation text within the same embedding space, reducing information loss in traditional multi-step processes. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a schematic diagram of the overall process of the present invention.

[0033] Figure 2 It is the network framework of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0035] The present invention provides a cross-language speech-text retrieval method based on a pre-trained automatic speech recognition model, combined withFigure 1 The overall process shown, the implementation method includes the following steps:

[0036] Step 1: Construct audio and text embeddings

[0037] First, prepare a training dataset containing N pairs of paired multi-speaker audio and text data. In the present invention, each pair of speech and text datasets comes from speech and text pairs in different languages, such as 21 different languages including English, Chinese, French, etc., and there is a direct semantic matching relationship between each speech and text pair.

[0038] Given the dataset where x i represents the speech data, and y i represents the corresponding text. The model encodes the speech and text respectively. Specifically, the model converts the speech data x into a 128-channel Mel spectrogram vector, denoted as where T represents the number of time frames. For the text data y, the model first appends a special token <|eot|> after each text token sequence to obtain y′ = [<|TextTokens|>, <|eot|>], and then encodes the text token y′ into a text token embedding, denoted as where d m is the dimension of the model, and M is the number of text tokens. In particular, the present invention does not add other special tokens such as language, task, or timestamp to y. In this way, N pairs of speech and text embedding pairs {X i , Y i} i=1:N are constructed.

[0039] Step 2: Initialize the network structure

[0040] As Figure 2 shown in the network framework, in this step, the present invention uses the pre-trained Whisper-large-v3 as the base model and makes appropriate modifications to it to adapt to the speech-text retrieval task.

[0041] In the Whisper-large-v3 model, the audio encoder and the text decoder interact through a complex cross-attention mechanism. In the present invention, to decouple the connection between the audio encoder and the text decoder, the cross-attention layer in the text decoder is removed. This operation helps improve the efficiency of the model and reduce unnecessary computational complexity, enabling the model to independently process speech and text data. Specifically, by adjusting the module structure of the model, the cross-modal interaction between audio and text is removed. Mathematically, this operation can be achieved by removing the matrices Q, K, V (representing query, key, and value respectively) of the cross-attention mechanism. This means that the self-attention layer of the text decoder only focuses on its own text information and no longer directly obtains information from the audio encoder.

[0042] In addition, to enable the model to better adapt to the tasks of speech retrieval of text and speech retrieval of text translation, the present invention introduces a low-rank adaptation layer in the self-attention layer of the text decoder for fine-tuning. When fine-tuning the pre-trained model using the low-rank adaptation method, the initial weights of the model are kept unchanged, and only the low-rank incremental weight matrix is updated, greatly reducing the computational cost. The self-attention calculation formula in the text decoder can be expressed as:

[0043]

[0044] After fine-tuning by the low-rank adaptation method, the weight matrix is updated as follows:

[0045]

[0046] where Softmax(·) is an activation function, are the weights of the pre-trained model; is an adjustable low-rank matrix, d m is the dimension of the model, k is the rank of the low-rank adaptation matrix, and in the present invention, k = 8 is set.

[0047] Step 3: Extract features and map them to a unified embedding space

[0048] In this step, the high-dimensional feature vectors of speech and text are mapped to a unified embedding space for cross-modal contrast learning.

[0049] Based on the N pairs of speech and text embedding pairs {X i , Y i} i=1:N obtained in Step 1, the present invention inputs the Mel spectrogram vector X into a convolutional layer containing two one-dimensional convolutions, where the dimension is increased to the model dimension d m, the context length is reduced by half and then input into the audio encoder of the pre-trained model, which is a stack composed of 32 Transformer encoder layers, so as to obtain a two-dimensional hidden feature

[0050] H speech = AudioEncoder(Conv(X)),

[0051] where AudioEncoder(·) is the operation of extracting features by the model audio encoder, and Conv(·) is the convolution operation. Each encoder block contains a self-attention layer with 20 self-attention heads and a two-layer multi-layer perceptron with Gaussian error linear unit as the activation function.

[0052] After obtaining the hidden speech feature H speech , the present invention uses attention pooling to pool it into a one-dimensional speech embedding Specifically, the cross-attention layer performs average pooling on H speech along the T dimension, that is:

[0053]

[0054] where MeanPool(·) is average pooling, taking h′ speech as the query in the attention calculation, and H speech itself as the key and value, and performing scaled dot-product attention calculation:

[0055] h speech = CrossAttn(query = h′ speech ; key, value = H speech ),

[0056] where CrosAttn(·) is attention pooling, and the three parameters query, key, and value are respectively the query, key, and value in the attention calculation. Finally, the present invention uses a projection layer (i.e., a linear transformation layer) to map the pooled speech feature h speech to a d e -dimensional speech-text joint embedding space to obtain

[0057] e speech = AudioProjection(h speech ),

[0058] where AudioProjection(·) is the audio projection mapping operation.

[0059] For text data, the present invention removes the cross-attention layer in the text decoder of the pre-trained model. Therefore, the structure of the text decoder is the same as that of the audio encoder. The only difference is that there is a causal attention mask in the text decoder. The text embedding Y will extract features through the decoder block, and finally obtain the text embedding

[0060] h text = TextDecoder(Y)[:,-1],

[0061] where TextDecoder(·) extracts features of the text decoder. The present invention selects the last hidden token, corresponding to the position of <|eot|>, to capture the global sentence features. Then, through the text projection layer, it is mapped to the joint speech-text embedding space to obtain the final text embedding As shown in the following formula:

[0062] e text = TextProjection(h text ),

[0063] where TextProjection(·) is the text projection mapping operation.

[0064] Through these steps, the features of audio and text are mapped to a unified embedding space, facilitating subsequent contrastive learning.

[0065] Step 4: Train the cross-lingual speech-text retrieval model

[0066] In the joint embedding space, the present invention defines a contrastive loss function to minimize the distance between similar speech and text embeddings and maximize the distance between dissimilar speech and text embeddings. The loss function we use is as follows:

[0067]

[0068] where represents the i-th speech-text embedding pair, τ is a learnable temperature parameter for scaling the loss, and N is the number of speech-text embedding pairs.

[0069] The pseudocode for calculating the contrastive loss function is as follows:

[0070] #audio_features: audio feature matrix, text_features: text feature matrix

[0071] #logit_scale: temperature parameter

[0072] #Step 1:Normalize features

[0073] audio_features = normalize(audio_features, p = 2, dim = -1)

[0074] text_features = normalize(text_features, p = 2, dim = -1)

[0075] # Step 2: Compute logits

[0076] logits_per_audio = logit_scale * (audio_features @ all_text_features.T)

[0077] logits_per_text = logit_scale * (text_features @ all_audio_features.T)

[0078] # Step3: Generate ground truth labels

[0079] num_samples = logits_per_audio.shape[0]

[0080] labels = range(0, num_samples)

[0081] if world_size > 1 and local_loss:

[0082] labels = labels + num_samples * global_rank

[0083] # Step 4: Compute loss

[0084] loss_audio_to_text = cross_entropy(logits_per_audio, labels)

[0085] loss_text_to_audio = cross_entropy(logits_per_text, labels)

[0086] loss = (loss_audio_to_text + loss_text_to_audio) / 2

[0087] The training process is optimized using the Stochastic Gradient Descent optimization algorithm to minimize the loss function. The model parameters are updated through backpropagation to make the embeddings of audio and text closer in a unified space.

[0088] Step 5: Inference of the cross - language speech - text retrieval system

[0089] During the inference process, the user needs to first prepare a database of texts (voices) to be queried. After inputting the query voice (text), following the operations described in Step 1 and Step 3, the system extracts and encodes features through the network model, and performs inference calculations according to the model trained in Step 4 and returns the most relevant results.

[0090] When the user inputs any query voice or query text, the query is first encoded through the encoding module of the pre - trained model. For the query voice, the audio encoder is used to obtain the query voice embedding. For the query text, the text decoder is used to obtain the query text embedding. Calculate the similarity between the query embedding and all text (voice) embeddings in the database z db The similarity calculation uses cosine similarity. Taking the query audio as an example:

[0091]

[0092] According to the calculated similarity, the audio or text results most relevant to the query are returned for the user to view.

[0093] To verify the effectiveness and robustness of the method described in the present invention, the present invention conducts a performance comparison with existing open - source models on the multi - language speech - text dataset FEURS. To evaluate the method performance, the evaluation metrics adopted by the present invention include Recall@K, Word Error Rate (WER), and BLEU score.

[0094] Recall@K measures the proportion of correct matching items in the top K retrieval results. Specifically: The present invention mainly uses the Recall@1 (R1) metric, that is, whether the correct matching text is included in the top 1 result returned by the model. Suppose there are N queries (audio or text embeddings). For each query i, the model returns K results and marks whether there is a correct matching text. If the correct matching text appears in the top K results, it is regarded as a correct match. The formula for Recall@1 is as follows:

[0095]

[0096] where 1 is the indicator function, which is 1 when the correct matching item appears in the top K retrieval results, and 0 otherwise.

[0097] WER measures the difference between the text predicted by the model and the reference text. WER is calculated based on the edit distance (Levenshtein distance), which represents the minimum number of operations (insertions, deletions, substitutions) required to transform one text into another. The formula for WER is:

[0098]

[0099] where S is the number of substituted words, D is the number of deleted words, I is the number of inserted words, and N is the total number of words in the reference text.

[0100] BLEU is a commonly used metric for automatically evaluating the quality of machine translation, which measures the similarity between the generated text and the reference text. The BLEU score is calculated based on the n-gram overlap, where an n-gram refers to a consecutive sequence of n words in the text. The formula for the BLEU score is:

[0101]

[0102] where p n is the n-gram precision, representing the overlapping ratio of the generated n-gram and the n-gram in the reference text. w n is the weighting coefficient of the n-gram precision, usually N is the maximum value of n-gram, usually taken as 4 (i.e., considering 1-gram to 4-gram). BP is the penalty term (Brevity Penalty), which is used to penalize the case where the generated text is shorter than the reference text. The formula is:

[0103]

[0104] where c is the length of the generated text and r is the length of the reference text.

[0105] Through a unified audio and text embedding space, the present invention enables the model to better adapt to multilingual and multimodal tasks. The present invention is only trained on the CoVoST-2 data containing 21 languages and still shows good metric performance on the FLEURS dataset containing 102 languages, as shown in Table 1 and Table 2. Compared with the prior art, the present invention performs excellently in multilingual cross-modal retrieval tasks, can simultaneously process datasets with up to 102 languages, and maintains stable precision in multimodal tasks. Especially in the performance of low-resource languages, it exceeds the precision level of the existing methods, verifying the superiority of the present invention in dealing with multilingual and low-resource language scenarios.

[0106] The present invention divides 102 languages in the FLEURS dataset into 15 language groups according to factors such as common features, historical origins, and grammatical structures of the languages, and tests the performance of the method proposed by the present invention for each language group. The results are shown in Table 3. The test results of this model are better than those of the SONAR model in 13 out of 15 language groups, verifying the superiority of this model in multilingual retrieval performance.

[0107] Table 1. Text test results of voice retrieval for this model and SONAR Note: The test set is 21 languages trained in FLEURS

[0108]

[0109] Table 2. Text test results of voice retrieval for this model and SONAR Note: The test set is all 102 languages in FLEURS

[0110]

[0111] Table 3. Test results of this model on the FLEURS dataset Note: Shown as the average of the test results of all languages within the language group

[0112]

[0113] In addition, the present invention extracts "75% transcribed text + 25% translated text" from the CoVoST-2 dataset to train a model for direct voice-to-text translation of retrieval. Compared with the model not trained on translated text, the BLEU metric has been significantly improved, as shown in Table 4. The model can directly compare and match audio and translated text within the same embedding space, reducing information loss in the traditional multi-step process.

[0114] Table 4. WhisperRet voice retrieval text translation test results

[0115]

Claims

1. A cross-language speech-to-text retrieval method based on a pre-trained automatic speech recognition model. The method uses an automatic speech recognition model trained on audio-text pairs to initialize the network structure. By removing the complex cross-attention mechanism in the pre-trained model, decoupling the connection between the audio encoder and the text decoder, building a double-tower structure, and adding a pooling layer and a projection layer, an efficient speech-to-text retrieval system is trained by fine-tuning. The specific steps include: Step 1: Build speech and text embeddings: Prepare a speech-text dataset containing N pairs of paired cross-language speech and text data, denoted as where x i Represents voice data, y i Represents the corresponding text data, and uses the audio encoding module of the pre-trained model to convert the speech data x into a 128-channel Mel spectrum vector, recorded as Where T represents the number of time frames, and the text encoding module is used to encode the text data into text token embeddings, denoted as where d m is the dimension of the model, M is the number of text tokens, and in this way, N pairs of speech and text embedding pairs {X i ,Y i } i=1:N ; Step 2: Initialize the network structure: Use the pre-trained automatic speech recognition model as the base model, remove the cross-attention layer in the text decoder, and then decouple the connection between the audio encoder and the text decoder, so that the model can process speech and text data independently; at the same time, in order to map speech and text data to a unified embedding space, add a pooling layer and an audio projection layer after the output of the audio encoder, and add a text projection layer after the output of the text decoder; in addition, freeze the pre-trained weights of the model to make full use of the rich prior knowledge of the pre-trained model, and introduce a low-rank adaptive layer in the self-attention layer of the text decoder to fine-tune the model, so that the model can better adapt to speech and text retrieval tasks; Step 3: Extract features and map them to a unified embedding space: Input the Mel-spectrogram vector X obtained in step 1 into the audio encoder to extract features. Each encoder block contains a self-attention layer with 20 self-attention heads and a two-layer multilayer perceptron with Gaussian error linear unit as the activation function. The output of the encoder is mapped to a unified embedding space after being processed by the pooling layer and the audio projection layer. The text token embedding Y obtained in step 1 is input into the text decoder to extract features. The text decoder has the same structure as the audio encoder except for the low-rank adaptive layer. The output of the decoder is mapped to a unified embedding space after being processed by the text projection layer. Step 4: Train the cross-language speech-text retrieval model: In the embedding space, use the contrastive loss function to minimize the distance between similar speech and text embeddings and maximize the distance between dissimilar speech and text embeddings. The loss function calculation formula is as follows: in represents the i-th speech-text embedding pair, τ is a learnable temperature parameter for scaling the loss, N is the number of speech-text embedding pairs, and the training process is optimized using the stochastic gradient descent algorithm. The model parameters are updated through back propagation to achieve the training goal of minimizing the loss function, so that the embedding of speech and text is closer in the unified space; Step 5: Cross-language speech-to-text retrieval system reasoning: Taking speech-to-text retrieval as an example, the user inputs any speech, and the model uses the audio encoder to obtain speech embedding Then calculate the speech embedding and text database z db The similarity of all text embeddings in , the similarity calculation uses cosine similarity: Among them, ||·||2 is the L2 norm, which is used to standardize the embedding vector. Finally, based on the calculated similarity, the text result most relevant to the speech is returned for the user to view.

2. The cross-language speech text retrieval method based on a pre-trained automatic speech recognition model according to claim 1, characterized in that: In step 1, the speech-text dataset should include paired speech-text data in 21 languages ​​in the CovoST-2 dataset.

3. The cross-language speech text retrieval method based on a pre-trained automatic speech recognition model according to claim 1, characterized in that: In step 1, for the text data y, the model first appends a special token <|eot|> to each text token sequence to obtain y′=[<|TextTokens|>,<|eot|>], and then encodes the text token y′ into a text token embedding, expressed as where d m is the dimension of the model and M is the number of text tokens.

4. The cross-language speech text retrieval method based on a pre-trained automatic speech recognition model according to claim 1, characterized in that: In step 2, the added pooling layer adopts the attention pooling method. Specifically, after the Mel spectrum vector X is extracted by the audio encoder, a two-dimensional hidden feature is obtained. where d m is the dimension of the model, T represents the number of time frames, first H speech Perform average pooling along the T dimension to obtain In the attention pooling calculation of the pooling layer, h′ speech As a query, H speech itself as the key and value, and perform the scaled dot product attention calculation to obtain a one-dimensional speech embedding h speech =CrossAttn(query=h′ speech ;key,value=H speech ), where CrossAttn(·) is the attention calculation, and the three parameters query, key, and value are the query, key, and value in the attention calculation respectively.

5. The cross-language speech text retrieval method based on a pre-trained automatic speech recognition model according to claim 1, characterized in that: In step 2, the added projection layer adopts a linear transformation operation.

6. The cross-language speech text retrieval method based on a pre-trained automatic speech recognition model according to claim 1, characterized in that: In step 3, the text embedding Y will be extracted through the text decoder to finally obtain the text embedding h text =TextDecoder(Y)[:,-1], Among them, TextDecoder(·) is the feature extraction operation of the text decoder.

7. The cross-language speech text retrieval method based on a pre-trained automatic speech recognition model according to claim 1, characterized in that: In step 4, the training process freezes the parameters of the encoder and decoder, and only trains the parameters of the low-rank adaptive layer, the audio pooling layer, and the projection layer.

Citation Information

Cited By

  • Intelligent tour guide method and system based on intelligent token and semantic fusion

    CN121301533A

  • An intelligent guide method and system based on intelligent token and semantic fusion

    CN121301533B