An Adaptive Hierarchical Representation Alignment Training Method and Device for Speech Large Models
By introducing internal speech adapter and adaptive hierarchical optimization selection strategy into the speech big model, cross-modal semantic retrieval and representation alignment training is used to use the mapping relationship between source speech and transcription text for cross-modal semantic retrieval and representation alignment training, the problem of modal differences between speech and text and insufficient understanding ability is solved, and the performance and robustness of the speech big model is significantly improved.
Patent Information
- Application Number
- CN202510206425.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-25
AI Technical Summary
Existing speech models have shortcomings in the modal gap and comprehension ability between speech and text, resulting in poor performance in complex application scenarios.
A method of adaptive hierarchical representation alignment training for speech large models is proposed. By introducing internal speech adapter and adaptive hierarchical optimization selection strategy, cross-modal semantic retrieval and representation alignment training is carried out by using the mapping relationship between source speech and transcription text.
It significantly improves the semantic understanding and representation alignment capabilities of the speech model in cross-modal speech-text tasks, narrows the modal gap between speech and text, and improves the performance and robustness of the model.
Smart Images

Figure CN119721258B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and particularly to an adaptive hierarchical representation alignment training method and device for a speech large model. Background Art
[0002] The training method of a speech large model currently mainly relies on paired speech and text data. The speech part of this data pair is usually in the audio form of natural language, and the text part is the written description associated with the semantics in the speech. In the training process, by introducing a series of speech language understanding tasks such as speech recognition tasks and speech annotation tasks, the model is prompted to be able to predict the corresponding text semantic information when receiving a speech input, realizing the implicit alignment of the speech and text modalities.
[0003] The training paradigms of speech large models are mainly divided into two categories. One category utilizes the structural characteristics of the large model itself and mature training paradigms. Based on the existing pre-trained text large model, by expanding the model vocabulary, the model can identify and process discrete speech units (such as phonemes, words, sub-words, etc.) in the speech signal. In this process, the model not only models the text content but also needs to convert the speech signal into discrete representation units during training and predict the associated text. This training method requires continuous pre-training using a large amount of speech-text data pairs, and at the same time trains the embedding representation of the expanded part of the large model vocabulary and the parameters inside the large model, ultimately enabling the text large model to obtain speech understanding capabilities.
[0004] Another training paradigm enables the speech input to be first converted into a continuous representation rich in semantic information through a dedicated speech understanding module (such as a speech pre-trained model). The task of the speech pre-trained model is to extract high-level and semantically rich features from the original speech signal, and these features can better represent the semantic content of the speech signal. The continuous representation undergoes a representation space transformation through a trainable adapter, enabling the transformed speech representation to be understood by the large model. This training method only uses fewer speech-text data pairs than the former, enabling the text large model to understand the speech signal input without training the original parameters of the text large model.
[0005] Although existing large speech models can initially complete speech-related understanding and generation tasks, there are still significant performance gaps compared to similar tasks in the text scenario (such as speech translation and machine translation). There is a monotonic mapping relationship between speech and text, but existing large speech model training techniques do not explicitly utilize the close connection between speech and text. Instead, they only implicitly model it as a supervision signal for text generation, failing to further narrow the modality gap between speech and text, which hinders the in-depth utilization of speech data in complex application scenarios. Existing research lacks exploration of the internal understanding mechanism of speech models, which not only limits the in-depth understanding of the model's working principle but also hinders the ability to further optimize and improve these models.
[0006] In the existing technology, there is a lack of an efficient and accurate adaptive hierarchical representation alignment training method that fully utilizes the mapping relationship between the source speech and the transcribed text. Summary of the Invention
[0007] To solve the technical problems of the modality differences between the source speech and the transcribed text, the differences in the performance of similar tasks under different modalities, and the understanding limitations of different modality information existing in the prior art, embodiments of the present invention provide an adaptive hierarchical representation alignment training method and device for a large speech model. The technical solutions are as follows:
[0008] On the one hand, an adaptive hierarchical representation alignment training method for a large speech model is provided. This method is implemented by an adaptive hierarchical representation alignment training device, and the method includes:
[0009] Obtain the source speech and the source speech target text; input the source speech into a candidate large speech model for text transcription to obtain the source speech transcription text;
[0010] Based on an internal speech adapter, according to the candidate large speech model, use the source speech and text prompts to train the model to obtain a first large speech model;
[0011] Based on a cross-modal semantic retrieval task, according to the source speech and the source speech transcription text, screen the semantic retrieval capabilities of multiple neural network levels of the first large speech model to obtain the optimal neural network level with semantic representation alignment;
[0012] Based on the optimal neural network level, input the source speech and the text prompts into the first large speech model for text prediction to obtain a speech representation, a speech attention weight matrix, and a predicted text;
[0013] Based on the optimal neural network level, input the source speech transcription text and the text prompts into the first large speech model for text representation extraction to obtain the optimal text representation and the text attention weight matrix;
[0014] Based on the target text, the speech representation, the speech attention weight matrix, the predicted text, the optimal text representation, and the text attention weight matrix, calculate the loss function to obtain the model prediction loss;
[0015] According to the model prediction loss, optimize the parameters of the first speech large model to obtain the second speech large model.
[0016] On the other hand, an adaptive hierarchical representation alignment training device for a speech large model is provided. This device is applied to the adaptive hierarchical representation alignment training method of the speech large model. The device includes:
[0017] A data acquisition module, configured to acquire the source speech and the source speech target text; input the source speech into the candidate speech large model for text transcription to obtain the source speech transcription text;
[0018] A model training module, configured to, based on an internal speech adapter, use the source speech and the text prompt words to train the model according to the candidate speech large model to obtain the first speech large model;
[0019] A cross-modal semantic retrieval module, configured to, based on the cross-modal semantic retrieval task, screen the semantic retrieval capabilities of multiple neural network levels of the first speech large model according to the source speech and the source speech transcription text to obtain the optimal neural network level with aligned semantic representations;
[0020] A text prediction module, configured to, based on the optimal neural network level, input the source speech and the text prompt words into the first speech large model for text prediction to obtain the speech representation, the speech attention weight matrix, and the predicted text;
[0021] A text representation extraction module, configured to, based on the optimal neural network level, input the source speech transcription text and the text prompt words into the first speech large model for text representation extraction to obtain the optimal text representation and the text attention weight matrix;
[0022] A loss calculation module, which calculates the loss function according to the target text, the speech representation, the speech attention weight matrix, the predicted text, the optimal text representation, and the text attention weight matrix to obtain the model prediction loss;
[0023] A model optimization module, configured to optimize the parameters of the first speech large model according to the model prediction loss to obtain the second speech large model.
[0024] On the other hand, an adaptive hierarchical representation alignment training device is provided. The adaptive hierarchical representation alignment training device includes: a processor; a memory storing computer-readable instructions, which when executed by the processor, implement any one of the methods in the above-mentioned adaptive hierarchical representation alignment training method for the speech large model.
[0025] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement any one of the methods in the above-mentioned adaptive hierarchical representation alignment training method for the speech large model.
[0026] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:
[0027] The present invention proposes an adaptive hierarchical representation alignment training method for a speech large model. By introducing an internal speech adapter and an adaptive hierarchical optimization selection strategy, the semantic understanding and representation alignment capabilities of the speech large model in cross-modal speech-text tasks are significantly improved, and the performance of the model is enhanced. The adaptive hierarchical optimization selection strategy allows the model to automatically select the optimal layer for training according to the actual semantic understanding ability, avoiding resource waste, accelerating the convergence speed, and improving the training efficiency. By constructing a first representation loss, the difference between the speech representation and the text representation is directly minimized, ensuring a high degree of alignment between the two at the semantic level, and improving the model's understanding and robustness of complex semantic structures. Utilizing the natural mapping relationship between the source speech and the transcribed text, by splicing the speech and text prompts and inputting them into the model, the deep fusion of the two modal representations is promoted, the semantic parsing ability is enhanced, and the modal gap between speech and text is effectively reduced. The present invention is an efficient and accurate adaptive hierarchical representation alignment training method that makes full use of the mapping relationship between the source speech and the transcribed text. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0029] Figure 1 is a flowchart of an adaptive hierarchical representation alignment training method for a speech large model provided by an embodiment of the present invention;
[0030] Figure 2 is a schematic structural diagram of an internal speech adapter provided by an embodiment of the present invention;
[0031] Figure 3It is a block diagram of an adaptive hierarchical representation alignment training device for a speech large model provided by an embodiment of the present invention;
[0032] Figure 4 It is a schematic structural diagram of an adaptive hierarchical representation alignment training device provided by an embodiment of the present invention. Specific embodiments
[0033] Next, in combination with the accompanying drawings, the technical solutions in the present invention will be described.
[0034] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of the word "example" is intended to present concepts in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two can be selected.
[0035] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when their differences are not emphasized, the meanings they express are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when their differences are not emphasized, the meanings they express are the same.
[0036] In the embodiments of the present invention, sometimes subscripts such as W 1 may be written in a non-subscript form such as W1. When their differences are not emphasized, the meanings they express are the same.
[0037] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail in combination with the accompanying drawings and specific embodiments.
[0038] The embodiments of the present invention provide an adaptive hierarchical representation alignment training method for a speech large model. This method can be implemented by an adaptive hierarchical representation alignment training device, and this adaptive hierarchical representation alignment training device can be a terminal or a server. As Figure 1 shown in the flowchart of the adaptive hierarchical representation alignment training method for the speech large model, the processing flow of this method can include the following steps:
[0039] S1. Obtain the source speech and the source speech target text; input the source speech into the candidate speech large model for text transcription to obtain the source speech transcription text.
[0040] In a feasible implementation manner, in the present invention, the source voice refers to the original voice data at the input end. The voice data refers to the sound signals generated by the vibration of the vocal cords of humans recognized and processed by a computer system, and is a carrier of natural language. Each type of voice itself represents the corresponding type of natural language, such as Chinese voice, English voice, etc.
[0041] The source voice transcription text refers to the text data of the first natural language type corresponding to the source voice. Natural language generally refers to the spoken language that has evolved naturally in culture, such as: Chinese, English, Japanese, German, etc. The first natural language can be implemented as any natural language. Taking the first natural language being implemented as Chinese as an example, the text data of the first natural language type can be "Hello!".
[0042] The target text refers to the text data expected to be generated or converted after the source voice or the source voice transcription text is input into the speech large model. The specific form and content of the target text depend on the tasks and application scenarios of the speech large model. The second natural language type of the target text can be the same as the first natural language type of the source voice transcription text, or different from the first natural language type of the source voice transcription text.
[0043] Obtain the voice and transcription text data of the first natural language type in the field of speech and the target text data of the second natural language type from the database.
[0044] Or obtain arbitrary voice data of the first natural language type and the target text data of the second natural language from the database, and use a professional speech recognition model to transcribe the voice data to obtain the corresponding source voice transcription text data of the first natural language type.
[0045] In addition, due to the different pronunciation habits and pronunciation methods in different regions, in the same text natural language scenario, there correspond different voice natural language scenarios, such as Cantonese, Minnan dialect, Mandarin, etc. corresponding to the Chinese text natural language. If the natural language type of the source voice is a non-written language form, the source voice transcription text should be similar to or within the same language system as the natural language type of the source voice. For example, if the source voice is Minnan dialect, and Minnan dialect and Chinese both belong to the Chinese branch of the Sino-Tibetan language family, the source voice transcription text can be Chinese text. For example, if the first natural language of the source voice is Chinese voice and the first natural language of the source voice transcription text is Chinese, the trained speech large model can implement tasks such as recognition, processing, and analysis of Chinese voice.
[0046] S2. Based on the internal voice adapter, according to the candidate speech large model, use the source voice and the text prompt word for model training to obtain the first speech large model.
[0047] Optionally, based on an internal voice adapter, according to a candidate voice large model, using a source voice and text prompts to train the model to obtain a first voice large model, including:
[0048] Insert the internal voice adapter into the candidate voice large model to obtain an improved candidate language large model;
[0049] Use the source voice and text prompts to train the improved candidate voice large model to obtain a first voice large model;
[0050] The language type of the text prompts is the same as that of the source voice transcription text.
[0051] In a feasible implementation manner, in the present invention, the candidate voice large model includes a candidate voice encoder, a candidate feature adapter, a candidate text generation model, etc. Among them, the candidate voice encoder is responsible for signal processing and feature extraction of the source voice; the candidate feature adapter is used to convert the features extracted by the voice encoder; the candidate text generation model is responsible for understanding the voice representation after the feature adapter and generating text. When the source voice is input into the candidate voice large model, the start position and end position of the voice input will be recorded for subsequent training of the internal voice adapter.
[0052] The training of the candidate voice large model is generally a speech recognition task, that is, inputting the source voice and text prompts, and requiring the model to output the transcription text corresponding to the source voice. After training, a first voice large model is obtained. The text prompts refer to the text instruction input of the voice large model, and it is the same natural language type as the source voice transcription text, that is, the first natural language type of the text prompts is the same as the first natural language type of the source voice transcription text. This instruction guides the voice large model to complete a series of tasks such as recognition, processing, analysis, and text generation of the source voice.
[0053] Among them, the structure of the internal voice adapter includes two linear mapping layers, a non-linear activation function, a residual connection, and a mean squared error regularization module.
[0054] In a feasible implementation manner, the structure of the internal voice adapter is as Figure 2 shown, the internal voice is inserted behind each neural network layer, and its scope of action is only for the voice representation part in the concatenated representation.
[0055] The input representation first undergoes dimensionality reduction through the first linear mapping layer, passes through a non-linear activation function to obtain the non-linear mapping ability in the high-dimensional space; then undergoes dimensionality increase through the second linear mapping layer; the input representation is added to the representation after dimensionality increase through a residual connection, and is output after passing through a root mean square normalization module. According to the starting position and ending position in the source speech encoding process, the corresponding speech representations are obtained, and this part of the speech representations is converted, and the converted representations replace the previous positions, and the replaced representations are continuously input into the next layer of the candidate speech large model.
[0056] S3. Based on the cross-modal semantic retrieval task, according to the source speech and the source speech transcription text, screen the semantic retrieval capabilities of multiple neural network levels of the first speech large model to obtain the optimal neural network level with aligned semantic representations.
[0057] Optionally, based on the cross-modal semantic retrieval task, according to the source speech and the source speech transcription text, screen the semantic retrieval capabilities of multiple neural network levels of the first speech large model to obtain the optimal neural network level with aligned semantic representations, including:
[0058] According to the source speech, extract the data representations of multiple neural network levels through the first speech large model to obtain a speech representation matrix;
[0059] According to the source speech transcription text, extract the data representations of multiple neural network levels through the first speech large model to obtain a text representation matrix;
[0060] Construct a cross-modal semantic retrieval task based on the Wasserstein metric of the optimal transport theory;
[0061] Based on the cross-modal semantic retrieval task, calculate the similarity according to the speech representation matrix and the text representation matrix to obtain a similarity matrix;
[0062] Based on the retrieval metrics, according to the similarity matrix, screen the semantic retrieval capabilities of multiple neural network levels of the first speech large model to obtain the optimal neural network level with aligned semantic representations; the retrieval metrics include the top-K recall rate and the mean reciprocal rank.
[0063] In a feasible implementation manner, in the present invention, using the pairing relationship between the source speech and the source speech transcription text, randomly select several pieces of data, and consider the paired data as positive examples, and the rest as negative examples.
[0064] Input the source speech into the first speech large model, extract the speech representations of each layer passing through the internal speech adapter, input the source speech transcription text into the first speech large model, extract the text representations of each layer, compress the lengths of the representations, and then calculate the similarity between all the speech representations and the text representations of each layer to obtain a similarity matrix.
[0065] Considering that common metrics such as cosine similarity only reflect the correlation of the global semantic information between speech and text and cannot reflect the correlation of fine-grained semantic information, therefore, the present invention adopts a loss based on the optimal transport theory and uses the Wasserstein Distance to measure the degree of fine-grained semantic similarity between speech and text.
[0066] The Wasserstein Distance, also known as the Earth Mover's Distance, is a measure between two discrete probability distributions. It evaluates the minimum cost required to redistribute one probability distribution into another, and is commonly used in style transfer in the field of images, training in generative adversarial networks, etc. The goal of optimal transport is to find a mapping from X to Y that minimizes the "transportation cost" from μ to ν, and the mathematical expressions are shown as the following formulas (1), (2), and (3):
[0067] (1);
[0068] (2);
[0069] (3);
[0070] Among them, given two probability distributions μ and ν on spaces X and Y respectively, they represent two different material distributions or resource distributions. The length of distribution μ is n and the length of ν is m. The space X corresponding to the distribution is n d-dimensional vectors, and the space Y is m d-dimensional vectors. The transport plan , represents to the transport quality, and the cost matrix , represents the corresponding space to the corresponding space of the transport cost (usually the Euclidean norm between two vectors). The final problem can be described as finding an optimal transport plan Z that minimizes the overall transport cost W.
[0071] Define X and Y as the speech representation space and the text representation space, and assume that the distributions u and v are uniform distributions. Then the Wasserstein distance between the speech representation and the text representation can be calculated. This distance measures the influence of all element pairs in the speech / text representation on the distance and can reflect the fine-grained semantic alignment information. The smaller the distance, the more aligned the semantic information of the two representations, and the farther the distance, the more distant the semantic information of the two representations.
[0072] The present invention uses a series of retrieval metrics such as the top-K recall rate and the mean reciprocal rank to measure the semantic retrieval capabilities of each layer, and obtains the optimal layer of semantic representation alignment through multi-metric trade-off.
[0073] S4. Based on the optimal neural network layer, input the source speech and the text prompt into the first speech large model for text prediction to obtain a speech representation, a speech attention weight matrix, and a predicted text.
[0074] Optionally, based on the optimal neural network layer, input the source speech and the text prompt into the first speech large model for text prediction to obtain a speech representation, a speech attention weight matrix, and a predicted text, including:
[0075] Concatenate the source speech and the text prompt to obtain a first input data;
[0076] Input the first input data into a speech encoder for feature extraction to obtain a concatenated speech representation;
[0077] Input the concatenated speech representation into an internal speech adapter for dimension conversion to obtain a speech representation;
[0078] Input the speech representation into a feature adapter for adaptation to obtain an adapted speech representation;
[0079] Based on the optimal neural network layer, input the adapted speech representation into a text generation model for text prediction to obtain a predicted text and a speech attention weight matrix.
[0080] In a feasible implementation manner, in the present invention, the speech representation includes the representation after the optimal layer and its internal speech adapter and the attention representation in the middle of the optimal layer.
[0081] The attention representation in the middle of the optimal layer is obtained through operations from the optimal layer of the first speech large model. Among them, the number of layers of the optimal layer is and the input representation of the optimal layer is The overall attention representation in the middle of the optimal layer is calculated by the following formulas (4), (5), and (6):
[0082] (4);
[0083] (5);
[0084] (6)
[0085] Among them, RMSNorm represents the root mean square normalization module, are two different linear mapping matrices respectively, represents the linear mapping matrix The hidden dimension of Represents the first Position to The degree of attention paid to each location; Represents the vector of the i-th position after the input representation is transformed by linear mapping; similarly, Represents the vector at the jth position after the input representation is transformed by another linear mapping.
[0086] For the speech attention weight matrix in the middle of the optimal layer, by recording the starting position of the speech representation and end position And the starting position of the target text after the first speech model segmentation and end position , we can confirm the corresponding speech attention weight matrix, that is, the overall attention representation matrix in the middle of the optimal layer The first dimension of to Position, pair matrix The second dimension of to Position, thus obtaining the speech attention weight matrix in the middle of the optimal layer .
[0087] The construction of text prompt words is related to the candidate speech task and the target text. The downstream speech task can be a speech translation task or a speech question-answering task. Taking the speech translation task as an example, the text prompt word can be "Please help me translate this speech into English."
[0088] The natural language type of the predicted text may be the same as or different from the first natural language type of the text prompt word.
[0089] If the source voice input is "Hello" in Mandarin, the first natural language type is the Chinese text prompt "Please help me translate this voice into English", and the second natural language type is the English target text "Hello", after training, the first voice model can generate the first predicted text "Hello" in English. The first language model can complete the cross-language and cross-modal generation task. If the source voice input is "Hello" in Mandarin, the first natural language type is the Chinese text prompt "Please give the Chinese synonym expression of this voice", and the second natural language type is the Chinese target text "Hello", after training, the first voice model can generate the Chinese "Hello". The first language model can complete the cross-modal generation task.
[0090] S5. Based on the optimal neural network layer, input the source speech transcription text and the text prompt into the first large speech model for text representation extraction to obtain the optimal text representation and the text attention weight matrix.
[0091] Optionally, based on the optimal neural network layer, input the source speech transcription text and the text prompt into the first large speech model for text representation extraction to obtain the optimal text representation and the text attention weight matrix, including:
[0092] Concatenate the source speech transcription text and the text prompt to obtain the second input data;
[0093] Based on the optimal neural network layer, according to the second input data, perform text representation extraction through a text generation model to obtain the optimal text representation and the text attention weight matrix.
[0094] In a feasible implementation, concatenate and input the source speech transcription text and the text prompt into the first large speech model to obtain the text representation of the optimal layer, including the representation passing through the optimal layer and the attention representation in the middle of the optimal layer.
[0095] The calculation of the overall text attention representation matrix in the middle of the optimal layer is similar to the calculation of the overall speech attention representation matrix in the middle of the optimal layer. For the text attention weight matrix in the middle of the optimal layer, by recording the starting position and the ending position of the representation input by the source speech transcription text and the starting position and the ending position after the target text is tokenized by the first speech model, the corresponding text attention weight matrix can be confirmed, that is, take the to position of the first dimension of the overall attention representation matrix in the middle of the optimal layer, and take the to to position of the second dimension of the matrix, thus obtaining the text attention weight matrix in the middle of the optimal layer .
[0096] S6. Calculate the loss function according to the target text, speech representation, speech attention weight matrix, predicted text, optimal text representation and text attention weight matrix to obtain the model prediction loss.
[0097] Optionally, calculate the loss function according to the target text, speech representation, speech attention weight matrix, predicted text, optimal text representation and text attention weight matrix to obtain the model prediction loss, including:
[0098] The Wasserstein metric based on the optimal transport theory constructs an optimal transport loss function according to the speech attention weight matrix and the text attention weight matrix;
[0099] Based on the optimal transport loss function, calculate the loss function according to the speech representation and the optimal text representation to obtain the first loss;
[0100] Based on the cross-entropy loss function, calculate the loss function according to the predicted text and the target text to obtain the second loss.
[0101] In a feasible implementation manner, the present invention takes into account the length difference between the speech representation and the text representation, and cannot directly use the L2 norm as the model prediction loss. Therefore, the Wasserstein Distance based on the optimal transport theory is adopted as the optimization objective.
[0102] Different from the existing work's use of this loss, the existing loss based on the Wasserstein Distance has an assumption, that is, it is assumed that the distributions of the speech representation and the text representation in terms of length are uniform, or an empirically assumed distribution. Such a distribution has certain limitations, that is, it cannot well reflect what kind of high-level distribution mapping relationship exists between speech and text, and any assumed distribution cannot adapt to all situations.
[0103] Therefore, the present invention proposes a loss calculation method guided by the model characteristics. Transpose the speech attention weight matrix of the optimal layer and the text attention weight matrix and then perform a softmax transformation to obtain the distribution u and the distribution v of this optimal layer. The speech representation and the text representation extracted from this layer obtain the space X and the space Y. Based on the above method, the characteristics that the model pays different attentions to different samples can be utilized to dynamically guide the calculation of the optimal transport loss, making the calculation of the loss more accurate, so as to improve the optimization accuracy and further improve the model performance; the first loss is the optimal transport loss guided by the model characteristics, and the second loss includes at least one of the cross-entropy loss, the mean square error loss, etc.
[0104] S7. Optimize the parameters of the first speech large model according to the model prediction loss to obtain the second speech large model.
[0105] In a feasible implementation, the training method proposed in the present invention can be applied to the training of large speech models under various architectures, enhance the speech comprehension ability of large speech models, and fully explore the rich knowledge and generalization ability of large text models in text scenarios. Explain the internal working mechanism of the model at the level of multimodal information understanding. In addition, the trained large speech model can be applied to a variety of downstream speech tasks. For example: when communicating across languages, users can input the speech to be translated into the large speech model, give text instructions to require the large speech model to translate in the direction of a specific target language, and the large speech model can generate a translated text in a given language.
[0106] The present invention proposes an adaptive hierarchical representation alignment training method for a large speech model. By introducing an internal speech adapter and an adaptive hierarchical optimization selection strategy, the semantic understanding and representation alignment capabilities of the large speech model in cross-modal speech-text tasks are significantly improved, and the performance of the model is enhanced. The adaptive hierarchical optimization selection strategy allows the model to automatically select the optimal layer for training according to the actual semantic understanding ability, avoiding resource waste, accelerating the convergence speed, and improving the training efficiency. By constructing the first representation loss, the difference between the speech representation and the text representation is directly minimized, ensuring a high degree of alignment between the two at the semantic level, and improving the model's understanding and robustness for complex semantic structures. By utilizing the natural mapping relationship between the source speech and the transcribed text, by splicing the speech and text prompt words into the model, the deep fusion of the two modal representations is promoted, the semantic parsing ability is enhanced, and the modal gap between speech and text is effectively narrowed. The present invention is an efficient and accurate adaptive hierarchical representation alignment training method that fully utilizes the mapping relationship between the source speech and the transcribed text.
[0107] Figure 3 The block diagram of a device for adaptive hierarchical representation alignment training of a large speech model according to an exemplary embodiment is shown. The device is used for an adaptive hierarchical representation alignment training method of a large speech model. Figure 3 The device includes a data acquisition module 310, a model training module 320, a cross-modal semantic retrieval module 330, a text prediction module 340, a text representation extraction module 350, a loss calculation module 360 and a model optimization module 370. Among them:
[0108] The data acquisition module 310 is used to acquire the source speech and the target text of the source speech; the source speech is input into the candidate speech large model for text transcription to obtain the source speech transcription text;
[0109] A model training module 320 is used to perform model training based on the internal speech adapter and the candidate speech large model using the source speech and the text prompt words to obtain a first speech large model;
[0110] The cross-modal semantic retrieval module 330 is used for semantic retrieval ability screening of multiple neural network layers of the first speech large model based on the cross-modal semantic retrieval task according to the source speech and the source speech transcription text, and obtaining the optimal neural network layer with semantic representation alignment;
[0111] The text prediction module 340 is used for inputting the source speech and the text prompt word into the first speech large model based on the optimal neural network layer for text prediction, and obtaining the speech representation, the speech attention weight matrix, and the predicted text;
[0112] The text representation extraction module 350 is used for inputting the source speech transcription text and the text prompt word into the first speech large model based on the optimal neural network layer for text representation extraction, and obtaining the optimal text representation and the text attention weight matrix;
[0113] The loss calculation module 360 calculates the loss function according to the target text, the speech representation, the speech attention weight matrix, the predicted text, the optimal text representation, and the text attention weight matrix, and obtains the model prediction loss;
[0114] The model optimization module 370 is used for parameter optimization of the first speech large model according to the model prediction loss, and obtaining the second speech large model.
[0115] Optionally, the model training module 320 is further used for:
[0116] Inserting the internal speech adapter into the candidate speech large model to obtain the improved candidate language large model;
[0117] Using the source speech and the text prompt word to train the improved candidate speech large model to obtain the first speech large model;
[0118] The language type of the text prompt word is the same as that of the source speech transcription text.
[0119] Among them, the structure of the internal speech adapter includes two linear mapping layers, a non-linear activation function, a residual connection, and a mean square error regularization module.
[0120] Optionally, the cross-modal semantic retrieval module 330 is further used for:
[0121] According to the source speech, performing data representation extraction of multiple neural network layers through the first speech large model to obtain the speech representation matrix;
[0122] According to the source speech transcription text, performing data representation extraction of multiple neural network layers through the first speech large model to obtain the text representation matrix;
[0123] Constructing a cross-modal semantic retrieval task based on the Wasserstein metric of the optimal transport theory;
[0124] Based on the cross-modal semantic retrieval task, similarity calculation is performed according to the speech representation matrix and the text representation matrix to obtain a similarity matrix;
[0125] Based on the retrieval metrics, according to the similarity matrix, the semantic retrieval capabilities of multiple neural network layers of the first speech large model are screened to obtain the optimal neural network layer with aligned semantic representations; the retrieval metrics include the top-K recall rate and the mean reciprocal rank.
[0126] Optionally, the text prediction module 340 is further configured to:
[0127] Concatenate the source speech and the text prompt to obtain the first input data;
[0128] Input the first input data into the speech encoder for feature extraction to obtain the concatenated speech representation;
[0129] Input the concatenated speech representation into the internal speech adapter for dimensionality conversion to obtain the speech representation;
[0130] Input the speech representation into the feature adapter for adaptation to obtain the adapted speech representation;
[0131] Based on the optimal neural network layer, input the adapted speech representation into the text generation model for text prediction to obtain the predicted text and the speech attention weight matrix.
[0132] Optionally, the text representation extraction module 350 is further configured to:
[0133] Concatenate the source speech transcription text and the text prompt to obtain the second input data;
[0134] Based on the optimal neural network layer, according to the second input data, perform text representation extraction through the text generation model to obtain the optimal text representation and the text attention weight matrix.
[0135] Optionally, the loss calculation module 360 is further configured to:
[0136] Based on the Wasserstein metric of the optimal transport theory, construct an optimal transport loss function according to the speech attention weight matrix and the text attention weight matrix;
[0137] Based on the optimal transport loss function, perform loss function calculation according to the speech representation and the optimal text representation to obtain the first loss;
[0138] Based on the cross-entropy loss function, perform loss function calculation according to the predicted text and the target text to obtain the second loss.
[0139] The present invention proposes an adaptive hierarchical representation alignment training method for a large speech model. By introducing an internal speech adapter and an adaptive hierarchical optimization selection strategy, the semantic understanding and representation alignment capabilities of the large speech model in cross-modal speech-text tasks are significantly improved, and the performance of the model is enhanced. The adaptive hierarchical optimization selection strategy allows the model to automatically select the optimal layer for training according to the actual semantic understanding ability, avoiding resource waste, accelerating the convergence speed, and improving the training efficiency. By constructing the first representation loss, the difference between the speech representation and the text representation is directly minimized, ensuring a high degree of alignment between the two at the semantic level, and improving the model's understanding and robustness for complex semantic structures. By utilizing the natural mapping relationship between the source speech and the transcribed text, by splicing the speech and text prompt words into the model, the deep fusion of the two modal representations is promoted, the semantic parsing ability is enhanced, and the modal gap between speech and text is effectively narrowed. The present invention is an efficient and accurate adaptive hierarchical representation alignment training method that fully utilizes the mapping relationship between the source speech and the transcribed text.
[0140] Figure 4 is a schematic diagram of the structure of an adaptive hierarchical representation alignment training device provided by an embodiment of the present invention, such as Figure 4 As shown, the adaptive hierarchical representation alignment training device may include the above Figure 3 The adaptive hierarchical representation alignment training device for the large speech model shown. Optionally, the adaptive hierarchical representation alignment training device 410 may include a first processor 2001.
[0141] Optionally, the adaptive hierarchical representation alignment training device 410 may further include a memory 2002 and a transceiver 2003 .
[0142] The first processor 2001, the memory 2002 and the transceiver 2003 may be connected via a communication bus.
[0143] Combine the following Figure 4 The components of the adaptive hierarchical representation alignment training device 410 are specifically introduced as follows:
[0144] Among them, the first processor 2001 is the control center of the adaptive hierarchical representation alignment training device 410, which can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0145] Optionally, the first processor 2001 can execute various functions of the adaptive hierarchical representation alignment training device 410 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0146] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 4 the CPU0 and CPU1 shown in
[0147] In a specific implementation, as an embodiment, the adaptive hierarchical representation alignment training device 410 may also include multiple processors, such as Figure 4 the first processor 2001 and the second processor 2004 shown in
[0148] Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0149] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or any other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently and be coupled to the first processor 2001 through an interface circuit ( Figure 4 not shown) of the adaptive hierarchical representation alignment training device 410. The embodiments of the present invention do not make specific limitations thereto.
[0150] The transceiver 2003 is used to communicate with a network device or with a terminal device.
[0151] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 4 not shown separately). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.
[0152] Optionally, the transceiver 2003 may be integrated with the first processor 2001 or may exist independently and be coupled to the first processor 2001 through an interface circuit ( Figure 4 not shown) of the adaptive hierarchical representation alignment training device 410. The embodiments of the present invention do not make specific limitations thereto.
[0153] It should be noted that Figure 4 the structure of the adaptive hierarchical representation alignment training device 410 shown in does not constitute a limitation to the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0154] In addition, the technical effects of the adaptive hierarchical representation alignment training device 410 may refer to the technical effects of the adaptive hierarchical representation alignment training method of the speech large model described in the above method embodiments, and will not be elaborated here.
[0155] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0156] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0157] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0158] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.
[0159] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0160] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0161] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0162] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0163] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0164] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0165] In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0166] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0167] As described above, the above are only specific implementation manners of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An adaptive hierarchical representation alignment training method for a large speech model, characterized in that: The method comprises: Acquire source speech and source speech target text; input the source speech into the candidate speech large model for text transcription to obtain the source speech transcription text; Based on the internal speech adapter, according to the candidate speech large model, the source speech and the text prompt word are used to perform model training to obtain a first speech large model; Based on the cross-modal semantic retrieval task, the semantic retrieval capability of multiple neural network layers of the first large speech model is screened according to the source speech and the source speech transcription text to obtain the optimal neural network layer for semantic representation alignment; The cross-modal semantic retrieval task is based on the source speech and the source speech transcription text, and the semantic retrieval capability of multiple neural network layers of the first speech model is screened to obtain the optimal neural network layer for semantic representation alignment, including: According to the source speech, extract data representations at multiple neural network levels through the first speech large model to obtain a speech representation matrix; According to the source speech transcription text, extract data representations at multiple neural network levels through the first speech large model to obtain a text representation matrix; Constructing a cross-modal semantic retrieval task based on the Wasserstein metric of optimal transfer theory; Based on the cross-modal semantic retrieval task, similarity calculation is performed according to the speech representation matrix and the text representation matrix to obtain a similarity matrix; Based on the retrieval index and the similarity matrix, multiple neural network layers of the first speech model are screened for semantic retrieval capabilities to obtain the optimal neural network layer for semantic representation alignment; the retrieval index includes the top K recall rate and the average reciprocal ranking; Based on the optimal neural network level, the source speech and the text prompt word are input into the first speech large model for text prediction to obtain speech representation, speech attention weight matrix and predicted text; Based on the optimal neural network level, the source speech transcription text and the text prompt words are input into the first speech macro model to extract text representation, and obtain the optimal text representation and text attention weight matrix; A loss function is calculated based on the target text, the speech representation, the speech attention weight matrix, the predicted text, the optimal text representation and the text attention weight matrix to obtain a model prediction loss; According to the model prediction loss, the first large speech model is optimized in parameters to obtain a second large speech model.
2. The adaptive hierarchical representation alignment training method for a large speech model according to claim 1, characterized in that: The method of performing model training based on the candidate speech large model using the source speech and the text prompt words based on the internal speech adapter to obtain the first speech large model includes: Inserting the internal speech adapter into the candidate speech large model to obtain an improved candidate speech large model; Using the source speech and the text prompt word, the improved candidate speech large model is trained to obtain a first speech large model; The language type of the text prompt word is the same as that of the source speech transcription text.
3. The adaptive hierarchical representation alignment training method for a large speech model according to claim 2, characterized in that: The structure of the internal speech adapter includes two linear mapping layers, a nonlinear activation function, a residual connection and a mean square error regularization module.
4. The adaptive hierarchical representation alignment training method for a large speech model according to claim 1, characterized in that: The method of inputting the source speech and the text prompt word into the first speech large model for text prediction based on the optimal neural network level to obtain speech representation, speech attention weight matrix and predicted text includes: splicing the source voice and the text prompt word to obtain first input data; Inputting the first input data into a speech encoder for representation extraction to obtain a concatenated speech representation; Inputting the concatenated speech representation into an internal speech adapter for dimensional conversion to obtain a speech representation; Inputting the speech representation into a feature adapter for adaptation to obtain an adapted speech representation; Based on the optimal neural network level, the adapted speech representation is input into a text generation model for text prediction to obtain predicted text and a speech attention weight matrix.
5. The adaptive hierarchical representation alignment training method for a large speech model according to claim 1, characterized in that: The method of inputting the source speech transcription text and the text prompt words into the first speech model to extract text representation based on the optimal neural network level to obtain the optimal text representation and the text attention weight matrix includes: splicing the source speech transcription text and the text prompt word to obtain second input data; Based on the optimal neural network level and according to the second input data, text representation extraction is performed through a text generation model to obtain an optimal text representation and a text attention weight matrix.
6. The adaptive hierarchical representation alignment training method for a large speech model according to claim 1, characterized in that: The step of calculating the loss function according to the target text, the speech representation, the speech attention weight matrix, the predicted text, the optimal text representation and the text attention weight matrix to obtain the model prediction loss includes: Based on the Wasserstein metric of optimal transmission theory, an optimal transmission loss function is constructed according to the speech attention weight matrix and the text attention weight matrix; Based on the optimal transmission loss function, a loss function is calculated according to the speech representation and the optimal text representation to obtain a first loss; Based on the cross entropy loss function, a loss function is calculated according to the predicted text and the target text to obtain a second loss.
7. An adaptive hierarchical representation alignment training device for a large speech model, the adaptive hierarchical representation alignment training device for a large speech model is used to implement the adaptive hierarchical representation alignment training method for a large speech model as claimed in any one of claims 1 to 6, characterized in that: The device comprises: A data acquisition module is used to acquire source speech and source speech target text; the source speech is input into the candidate speech large model for text transcription to obtain the source speech transcription text; A model training module, configured to perform model training based on an internal speech adapter and the candidate speech large model using the source speech and the text prompt words to obtain a first speech large model; A cross-modal semantic retrieval module, for performing semantic retrieval capability screening on multiple neural network layers of the first large speech model based on the cross-modal semantic retrieval task and according to the source speech and the source speech transcription text, to obtain the optimal neural network layer for semantic representation alignment; A text prediction module, configured to input the source speech and the text prompt word into the first speech large model for text prediction based on the optimal neural network layer, and obtain speech representation, speech attention weight matrix and predicted text; A text representation extraction module, for inputting the source speech transcription text and the text prompt words into the first speech macro model to extract text representation based on the optimal neural network layer, and obtaining an optimal text representation and a text attention weight matrix; A loss calculation module calculates a loss function according to the target text, the speech representation, the speech attention weight matrix, the predicted text, the optimal text representation and the text attention weight matrix to obtain a model prediction loss; The model optimization module is used to optimize the parameters of the first large speech model according to the model prediction loss to obtain the second large speech model.
8. An adaptive hierarchical representation alignment training device, characterized in that: The adaptive hierarchical representation alignment training device comprises: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Cross-modal video clip retrieval method based on pre-training language model adaptation network
CN116662609A
Cross-modal representation alignment-based English-beyond end-to-end speech translation method
CN116663577A