Audio understanding model training method and device, audio understanding method and device, storage medium and program product
By building an audio text classification model and using the gradient inversion mechanism to train the first fusion module, combined with the powerful language understanding ability of large language models, the problems of low training efficiency and poor results in streaming speech recognition tasks are solved, and efficient speech recognition in the case of scarce data is achieved.
Patent Information
- Application Number
- CN202510828122.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-19
AI Technical Summary
Existing speech recognition systems are inefficient in training in streaming speech recognition tasks and are not effective, especially in the case of scarcity of data, which makes it difficult for traditional models to handle complex speech tasks.
Using pre-trained speech recognition model and large language model, the first fusion module is trained by building an audio text classification model and using the gradient inversion mechanism to learn modal-independent feature representations, and combining the powerful language understanding and generation capabilities of large language models, an audio understanding model is built.
It improves the accuracy and efficiency of streaming speech recognition, and can significantly improve the speech recognition effect even with less data, avoiding oscillations or biases caused by modal differences during training.
Smart Images

Figure CN120356465A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of speech recognition technology, and in particular to a training method for an audio understanding model, an audio understanding method, a training device for an audio understanding model, an audio understanding device, a non-volatile computer-readable storage medium, and a computer program product. Background Art
[0002] The application scope of speech recognition technology has expanded to many industries, but its effectiveness depends largely on the quality and diversity of the training data used. Since modern speech recognition systems must deal with complex sound variables including various accents, speaking speeds, and background noise, obtaining accurate, diverse, and sufficient training data is crucial to building an efficient speech recognition system. In addition, in order to enhance the performance of the model in a specific vertical industry, detailed annotation of the speech data in the industry scenario is also required.
[0003] In this context, it is particularly important to use large language models (LLMs) to train audio understanding models. Large language models have proven their strong capabilities in language understanding and generation by pre-training on massive amounts of text data. In contrast, traditional models often rely on large amounts of training data to achieve acceptable performance. This reliance on big data increases the difficulty of resource acquisition, especially when dealing with languages with scarce data. Large language models, due to their powerful pre-training capabilities, can achieve good results in specific audio understanding tasks even with only a small amount of data. In addition, traditional models usually have a small number of parameters, which limits their learning ability and ability to handle complex problems, resulting in poor robustness. In contrast, large language models have a larger number of parameters, which enables them to show stronger robustness and adaptability when facing speech data in complex, unknown or noisy environments. Traditionally, different models need to be trained and deployed separately for different source and target languages, which not only increases the complexity of model management, but also makes it unrealistic to develop high-quality models for each language combination when resources are limited. Large language models, especially those pre-trained in multiple languages, can more easily adapt to different languages, reducing the need for large amounts of language-specific data. Finally, traditional models are often limited in effectiveness, especially when dealing with complex language phenomena. Large language models, with their powerful language understanding and generation capabilities, can achieve better results in a wider range of speech tasks.
[0004] Speech recognition systems are generally divided into two scenarios: non-streaming speech recognition and streaming speech recognition, and different scenarios are suitable for different business requirements. Non-streaming speech recognition means that after the user uploads a complete audio, the system performs speech recognition on the audio in an offline state and finally outputs the recognition result. Streaming speech recognition, on the other hand, is that while the user is speaking, that is, during the process of inputting speech, the speech recognition system synchronously recognizes the audio and displays the recognition result to the user in real time. Since users can obtain instant feedback, streaming speech recognition has more advantages in terms of interaction experience and is therefore commonly used in scenarios such as real-time captions, speech search, and speech interaction (such as smart homes, in-vehicle devices, digital humans, etc.). How to improve the training efficiency and speech recognition effect of large models for streaming audio understanding is an important research topic. Summary of the Invention
[0005] The present disclosure provides a technical solution for audio understanding.
[0006] According to one aspect of the present disclosure, there is provided a method for training an audio understanding model, including:
[0007] Obtaining a pre-trained speech recognition model and a pre-trained large language model; wherein, the speech recognition model includes an encoding module, a prediction module, and a first fusion module, the first fusion module is respectively connected to the encoding module and the prediction module, and the large language model includes a large language model main body and a large language model head that are connected to each other;
[0008] Constructing an audio text classification model according to the encoding module, the large language model main body, the first fusion module, and a binary classification layer, wherein the first fusion module is respectively connected to the encoding module and the large language model main body, and the binary classification layer is connected to the first fusion module;
[0009] Processing the acoustic features output by the encoding module or the semantic features output by the large language model main body through the first fusion module and the binary classification layer to obtain a modality prediction result;
[0010] Determining the reverse gradient value of the first fusion module according to the modality prediction result and the true modality corresponding to the modality prediction result, and updating the parameters of the first fusion module according to the reverse gradient value of the first fusion module to train the audio text classification model;
[0011] In response to the audio text classification model being trained to meet a preset condition, constructing an audio understanding model according to the encoding module, the large language model main body, and a second fusion module, and training the audio understanding model, wherein the second fusion module includes the first fusion module.
[0012] In a possible implementation manner, determining the reverse gradient value of the first fusion module according to the modality prediction result and the true modality corresponding to the modality prediction result includes:
[0013] Determining the value of the classification loss function corresponding to the audio text classification model according to the modality prediction result and the true modality corresponding to the modality prediction result;
[0014] Calculating the gradient value of the first fusion module according to the value of the classification loss function;
[0015] Reversing the gradient value of the first fusion module to obtain the reverse gradient value of the first fusion module.
[0016] In a possible implementation manner, processing the acoustic features output by the encoding module or the semantic features output by the large language model main body through the first fusion module and the binary classification layer to obtain a modality prediction result includes:
[0017] For any audio-text pair in the training set, extracting the acoustic features corresponding to the audio sample in the audio-text pair through the encoding module in the audio text classification model, and extracting the semantic features corresponding to the target text label sequence in the audio-text pair through the large language model main body in the audio text classification model;
[0018] Inputting the acoustic features or the semantic features into the first fusion module, and processing the acoustic features or the semantic features through the first fusion module to obtain fused features;
[0019] Processing the fused features through the binary classification layer to obtain a modality prediction result.
[0020] In a possible implementation manner, during the training process of the audio text classification model, the parameters of the encoding module and the large language model main body are kept fixed.
[0021] In a possible implementation manner, training the audio understanding model includes:
[0022] For any audio sample in the training set, extracting the acoustic feature sequence corresponding to the audio sample through the encoding module in the audio understanding model, and extracting the semantic feature sequence corresponding to the audio sample through the large language model main body in the audio understanding model;
[0023] Based on the acoustic feature sequence and the semantic feature sequence, determining all valid alignment paths that can generate the target text label sequence corresponding to the audio sample through the second fusion module, and calculating the total probability of all the valid alignment paths;
[0024] Take the negative logarithm of the total probability to obtain the value of the streaming sequence alignment loss function corresponding to the audio understanding model;
[0025] Update the parameters of the audio understanding model according to the value of the streaming sequence alignment loss function.
[0026] In a possible implementation, the second fusion module further includes a first head;
[0027] The method further includes:
[0028] Obtain the target token table corresponding to the audio understanding model, where the number of tokens in the target token table is less than the number of tokens in the original token table corresponding to the large language model;
[0029] According to the position indices of the respective tokens in the target token table in the original token table, extract the initial weights corresponding to the respective tokens in the target token table from the weight matrix of the large language model head;
[0030] Initialize the weight matrix of the first head according to the initial weights corresponding to the respective tokens in the target token table.
[0031] In a possible implementation, the first head further includes a dimension corresponding to the blank symbol.
[0032] In a possible implementation, obtaining the target token table corresponding to the audio understanding model includes:
[0033] Perform a word frequency statistics on the tokens in the training set;
[0034] Construct the target token table corresponding to the audio understanding model according to the preset number of tokens with the highest word frequencies in the training set.
[0035] In a possible implementation, the audio understanding model further includes a second head, the second head is connected to the large language model body, and the second head is initialized according to the large language model head;
[0036] The method further includes:
[0037] Determine the value of the knowledge distillation loss function according to the first text prediction result output by the first head and the second text prediction result output by the second head;
[0038] Update the parameters of the audio understanding model according to the value of the knowledge distillation loss function.
[0039] In a possible implementation, the method further includes:
[0040] Determine the value of the token distribution loss function corresponding to the audio understanding model according to the second text prediction result and the target text label sequence corresponding to the second text prediction result;
[0041] Update the parameters of the audio understanding model according to the value of the token distribution loss function.
[0042] In a possible implementation, during the training process of the audio understanding model, the parameters of the second head are kept fixed.
[0043] In a possible implementation, the audio understanding model is trained in a first stage and a second stage based on a training set;
[0044] Wherein, during the first stage of training, the parameters of the encoding module are kept fixed; during the second stage of training, the parameters of the encoding module are adjusted.
[0045] In a possible implementation, the speech recognition model uses a recurrent neural network transducer.
[0046] According to one aspect of the present disclosure, there is provided an audio understanding method, including:
[0047] Obtain an audio understanding model trained by the training method of the audio understanding model;
[0048] Input the audio to be processed into the audio understanding model, and output the text prediction result corresponding to the audio to be processed through the audio understanding model.
[0049] According to one aspect of the present disclosure, there is provided a training device for an audio understanding model, including:
[0050] A first obtaining module, configured to obtain a pre-trained speech recognition model and a pre-trained large language model; wherein, the speech recognition model includes an encoding module, a prediction module and a first fusion module, the first fusion module is respectively connected to the encoding module and the prediction module, and the large language model includes a large language model main body and a large language model head connected to each other;
[0051] A construction module, configured to construct an audio text classification model according to the encoding module, the large language model main body, the first fusion module and a binary classification layer, wherein the first fusion module is respectively connected to the encoding module and the large language model main body, and the binary classification layer is connected to the first fusion module;
[0052] A modality prediction module, configured to process the acoustic features output by the encoding module or the semantic features output by the large language model main body through the first fusion module and the binary classification layer to obtain a modality prediction result;
[0053] The first training module is used to determine the reverse gradient value of the first fusion module according to the modality prediction result and the true modality corresponding to the modality prediction result, and update the parameters of the first fusion module according to the reverse gradient value of the first fusion module, so as to train the audio-text classification model;
[0054] The second training module is used to, in response to the audio-text classification model being trained to meet a preset condition, construct an audio understanding model according to the encoding module, the large language model main body, and the second fusion module, and train the audio understanding model, where the second fusion module includes the first fusion module.
[0055] In a possible implementation manner, the first training module is used to:
[0056] Determine the value of the classification loss function corresponding to the audio-text classification model according to the modality prediction result and the true modality corresponding to the modality prediction result;
[0057] Calculate the gradient value of the first fusion module according to the value of the classification loss function;
[0058] Reverse the gradient value of the first fusion module to obtain the reverse gradient value of the first fusion module.
[0059] In a possible implementation manner, the modality prediction module is used to:
[0060] For any audio-text pair in the training set, extract the acoustic features corresponding to the audio sample in the audio-text pair through the encoding module in the audio-text classification model, and extract the semantic features corresponding to the target text label sequence in the audio-text pair through the large language model main body in the audio-text classification model;
[0061] Input the acoustic features or the semantic features into the first fusion module, and process the acoustic features or the semantic features through the first fusion module to obtain fusion features;
[0062] Process the fusion features through the binary classification layer to obtain a modality prediction result.
[0063] In a possible implementation manner, during the training process of the audio-text classification model, the parameters of the encoding module and the large language model main body remain fixed.
[0064] In a possible implementation manner, the second training module is used to:
[0065] For any audio sample in the training set, extract the acoustic feature sequence corresponding to the audio sample through the encoding module in the audio understanding model, and extract the semantic feature sequence corresponding to the audio sample through the large language model body in the audio understanding model;
[0066] Through the second fusion module, based on the acoustic feature sequence and the semantic feature sequence, determine all valid alignment paths that can generate the target text label sequence corresponding to the audio sample, and calculate the total probability of all the valid alignment paths;
[0067] Take the negative logarithm of the total probability to obtain the value of the streaming sequence alignment loss function corresponding to the audio understanding model;
[0068] Update the parameters of the audio understanding model according to the value of the streaming sequence alignment loss function.
[0069] In a possible implementation, the second fusion module further includes a first head;
[0070] The device further includes:
[0071] A second obtaining module, configured to obtain the target token table corresponding to the audio understanding model, where the number of tokens in the target token table is less than the number of tokens in the original token table corresponding to the large language model;
[0072] An extraction module, configured to extract the initial weights corresponding to each token in the target token table from the weight matrix of the large language model head according to the position indexes of each token in the target token table in the original token table;
[0073] An initialization module, configured to initialize the weight matrix of the first head according to the initial weights corresponding to each token in the target token table.
[0074] In a possible implementation, the first head further includes a dimension corresponding to the blank symbol.
[0075] In a possible implementation, the second obtaining module is configured to:
[0076] Perform word frequency statistics on the tokens in the training set;
[0077] Construct the target token table corresponding to the audio understanding model according to the preset number of tokens with the highest word frequencies in the training set.
[0078] In a possible implementation, the audio understanding model further includes a second head, the second head is connected to the large language model body, and the second head is initialized according to the large language model head;
[0079] The apparatus further includes:
[0080] A first determination module, configured to determine a value of a knowledge distillation loss function according to a first text prediction result output by the first head and a second text prediction result output by the second head;
[0081] A first update module, configured to update parameters of the audio understanding model according to the value of the knowledge distillation loss function.
[0082] In a possible implementation manner, the apparatus further includes:
[0083] A second determination module, configured to determine a value of a token distribution loss function corresponding to the audio understanding model according to the second text prediction result and a target text label sequence corresponding to the second text prediction result;
[0084] A second update module, configured to update parameters of the audio understanding model according to the value of the token distribution loss function.
[0085] In a possible implementation manner, during the training process of the audio understanding model, parameters of the second head are kept fixed.
[0086] In a possible implementation manner, the audio understanding model is trained in a first stage and a second stage based on a training set;
[0087] Wherein, during the first stage of training, parameters of the encoding module are kept fixed; during the second stage of training, parameters of the encoding module are adjusted.
[0088] In a possible implementation manner, the speech recognition model adopts a recurrent neural network transducer.
[0089] According to one aspect of the present disclosure, there is provided an audio understanding apparatus, including:
[0090] An acquisition module, configured to acquire an audio understanding model trained by a training apparatus of the audio understanding model;
[0091] An audio understanding module, configured to input an audio to be processed into the audio understanding model, and output a text prediction result corresponding to the audio to be processed through the audio understanding model.
[0092] According to one aspect of the present disclosure, there is provided a training apparatus of an audio understanding model, including a memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the steps of the above-mentioned training method of the audio understanding model.
[0093] According to one aspect of the present disclosure, an audio understanding device includes a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the above audio understanding method.
[0094] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above method is implemented.
[0095] According to one aspect of the present disclosure, a computer program product is provided, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, and when the computer-readable code runs in an electronic device, the processor in the electronic device executes the above method.
[0096] In an embodiment of the present disclosure, a pre-trained speech recognition model and a pre-trained large language model are obtained; wherein, the speech recognition model includes an encoding module, a prediction module, and a first fusion module, the first fusion module is respectively connected to the encoding module and the prediction module, the large language model includes a large language model main body and a large language model head that are connected to each other, an audio text classification model is constructed according to the encoding module, the large language model main body, the first fusion module, and a binary classification layer, wherein the first fusion module is respectively connected to the encoding module and the large language model main body, the binary classification layer is connected to the first fusion module, the acoustic features output by the encoding module or the semantic features output by the large language model main body are processed through the first fusion module and the binary classification layer to obtain a modality prediction result, the reverse gradient value of the first fusion module is determined according to the modality prediction result and the true modality corresponding to the modality prediction result, and the parameters of the first fusion module are updated according to the reverse gradient value of the first fusion module to train the audio text classification model. In response to the audio text classification model being trained to meet a preset condition, an audio understanding model is constructed according to the encoding module, the large language model main body, and a second fusion module, and the audio understanding model is trained, wherein the second fusion module includes the first fusion module. Thus, before training the audio understanding model, the first fusion module for fusing the audio modality and the text modality is pre-trained, and the first fusion module is enabled to learn modality-agnostic representations through a gradient reversal mechanism, so that the first fusion module can better learn how to balance and fuse the acoustic features from the encoding module and the semantic features from the large language model main body, and avoid problems such as oscillations or biases towards a certain modality caused by modality differences during the training process. In addition, the embodiment of the present disclosure utilizes the powerful language understanding and generation capabilities of the pre-trained large language model, as well as its ability to quickly learn through a small amount of audio data, to achieve efficient streaming speech recognition. Even when the audio training data is scarce, the accuracy of speech recognition can be significantly improved.
[0097] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure.
[0098] According to the following detailed description of exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. BRIEF DESCRIPTION OF THE DRAWINGS
[0099] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.
[0100] Figure 1The flowchart showing the training method of the audio understanding model provided by the embodiments of the present disclosure.
[0101] Figure 2 The schematic diagram showing the recurrent neural network transducer.
[0102] Figure 3 The schematic diagram showing the audio text classification model in the training method of the audio understanding model provided by the embodiments of the present disclosure.
[0103] Figure 4 The schematic diagram showing the audio understanding model provided by the embodiments of the present disclosure.
[0104] Figure 5 The schematic diagram showing the dynamic alignment process of the audio understanding model in the decoding stage provided by the embodiments of the present disclosure.
[0105] Figure 6 The block diagram showing the training device of the audio understanding model provided by the embodiments of the present disclosure.
[0106] Figure 7 The block diagram showing a training device or an audio understanding device of an audio understanding model according to an exemplary embodiment. Detailed implementation manners
[0107] The various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0108] The special word "exemplary" herein means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" herein is not necessarily to be construed as superior to or better than other embodiments.
[0109] The term "and / or" herein merely describes an association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent including any one or more elements selected from the set composed of A, B, and C.
[0110] In addition, for better illustrating the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can also be implemented without some specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0111] Embodiments of the present disclosure provide a training method for an audio understanding model, which includes obtaining a pre-trained speech recognition model and a pre-trained large language model. Among them, the speech recognition model includes an encoding module, a prediction module, and a first fusion module. The first fusion module is respectively connected to the encoding module and the prediction module. The large language model includes a large language model main body and a large language model head that are connected to each other. An audio text classification model is constructed according to the encoding module, the large language model main body, the first fusion module, and a binary classification layer. Among them, the first fusion module is respectively connected to the encoding module and the large language model main body, and the binary classification layer is connected to the first fusion module. The first fusion module and the binary classification layer are used to process the acoustic features output by the encoding module or the semantic features output by the large language model main body to obtain a modality prediction result. According to the modality prediction result and the true modality corresponding to the modality prediction result, the reverse gradient value of the first fusion module is determined, and according to the reverse gradient value of the first fusion module, the parameters of the first fusion module are updated to train the audio text classification model. In response to the audio text classification model being trained to meet a preset condition, an audio understanding model is constructed according to the encoding module, the large language model main body, and a second fusion module, and the audio understanding model is trained. Among them, the second fusion module includes the first fusion module. Thus, before training the audio understanding model, the first fusion module for fusing the audio modality and the text modality is pre-trained. Through the gradient reversal mechanism, the first fusion module learns modality-agnostic representations, enabling the first fusion module to better learn how to balance and fuse the acoustic features from the encoding module and the semantic features from the large language model main body, avoiding problems such as oscillations or biases towards a certain modality caused by modality differences during the training process. In addition, the embodiments of the present disclosure utilize the powerful language understanding and generation capabilities of the pre-trained large language model, as well as its ability to quickly learn through a small amount of audio data, to achieve efficient streaming speech recognition. Even when the audio training data is scarce, the accuracy of speech recognition can be significantly improved.
[0112] The following will describe in detail the training method for the audio understanding model provided by the embodiments of the present disclosure with reference to the accompanying drawings.
[0113] Figure 1The flowchart showing the training method of the audio understanding model provided by the embodiments of the present disclosure is shown. In a possible implementation manner, the execution main body of the training method of the audio understanding model can be a training device of the audio understanding model. For example, the training method of the audio understanding model can be executed by a terminal device, a server, or other electronic devices. Among them, the terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, or a wearable device, etc. In some possible implementation manners, the training method of the audio understanding model can be implemented by a processor calling computer-readable instructions stored in a memory. As Figure 1 shown, the training method of the audio understanding model includes steps S11 to S15.
[0114] In step S11, a pre-trained speech recognition model and a pre-trained large language model are obtained; wherein, the speech recognition model includes an encoding module, a prediction module, and a first fusion module, the first fusion module is respectively connected to the encoding module and the prediction module, and the large language model includes a large language model main body and a large language model head that are connected to each other.
[0115] In step S12, an audio text classification model is constructed according to the encoding module, the large language model main body, the first fusion module, and a binary classification layer, wherein the first fusion module is respectively connected to the encoding module and the large language model main body, and the binary classification layer is connected to the first fusion module.
[0116] In step S13, the acoustic features output by the encoding module or the semantic features output by the large language model main body are processed through the first fusion module and the binary classification layer to obtain a modality prediction result.
[0117] In step S14, according to the modality prediction result and the true modality corresponding to the modality prediction result, the reverse gradient value of the first fusion module is determined, and according to the reverse gradient value of the first fusion module, the parameters of the first fusion module are updated to train the audio text classification model.
[0118] In step S15, in response to the audio text classification model being trained to meet a preset condition, an audio understanding model is constructed according to the encoding module, the large language model main body, and a second fusion module, and the audio understanding model is trained, wherein the second fusion module includes the first fusion module.
[0119] In an embodiment of the present disclosure, the pre-trained speech recognition model may refer to a speech recognition model pre-trained on a speech dataset. By learning and extracting features from speech samples, the speech recognition model can capture the acoustic features, prosody patterns, and the correspondence between speech and text of the speech, and convert the input speech into the corresponding text content.
[0120] In an embodiment of the present disclosure, the pre-trained speech recognition model may include an encoding module (encoder), a prediction module (predictor), and a first fusion module (jointer). In some application scenarios, the encoding module may also be referred to as an encoder, etc., the prediction module may also be referred to as a predictor, a prediction network, etc., and the first fusion module may also be referred to as a joint network, a joint module, etc., which are not limited herein. In a possible implementation manner, the encoding module may be used to extract acoustic features, the prediction module may be used to extract semantic features, and the first fusion module may be used to fuse the acoustic features and semantic features.
[0121] In a possible implementation manner, the speech recognition model employs a Recurrent Neural Network Transducer (RNNT).
[0122] Figure 2 A schematic diagram showing the Recurrent Neural Network Transducer is as follows. Figure 2 As shown, the Recurrent Neural Network Transducer may include an encoding module, a prediction module, a first fusion module, and an activation layer.
[0123] Among them, the encoding module may be used to model the acoustic features. For example, the encoding module may convert the input audio frame sequence into a high-dimensional acoustic feature representation. For example, the encoding module may input the audio frame at the current time step t and output the acoustic feature corresponding to the current time step t . Among them, the time step may refer to the smallest discrete time unit for audio processing, corresponding to the single-frame data processed or output by the model at a certain moment. For the audio frame sequence , the encoding module may output an acoustic feature sequence . The encoding module may be implemented using a Conformer or a Long Short-Term Memory (LSTM) network, which is not limited herein.
[0124] The prediction module may be used to model the semantic features. In the training phase, the prediction module may input the historical text labels before the current text step u (such as or ) and output the semantic feature of the current text step u . In the prediction stage, the prediction module can take as input the already generated (i.e., already predicted) text labels and output the semantic features of the current text step. Here, a text step can refer to a discrete unit in the text generation process, corresponding to a single token processed or output by the model. The prediction module can adopt an LSTM or Transformer structure, which is not limited here.
[0125] The training samples in the training set can consist of audio-text pairs {x, y}, where x can represent an audio sample and y can represent the target text label sequence (i.e., the annotated text) corresponding to the audio sample. For example: the target text label sequence y is: "What's the weather like today". Then the text format input to the prediction module during training can be: <sos>What's the weather like today. Among them, <sos>is the start of sentence (SOS), which is used to initialize the hidden state of the prediction module and indicates the start of text generation. The format of the target text label sequence for calculating the loss can be: What's the weather like today <eos>Among them, <eos>Is the end-of-sentence marker (End Of Sentence). <eos>As part of the training objective, help the model learn when to terminate the output.
[0126] The first fusion module can fuse the outputs of the encoding module and the prediction module. For example, the output of the first fusion module can be: . The first fusion module can include a linear layer and an activation layer (such as softmax).
[0127] As Figure 2 shown, after the first fusion module, an activation layer (such as softmax) can be further connected, and through this activation layer, the probability distribution of all possible tokens (such as characters, words, symbols) is generated .
[0128] In this implementation, the pre-trained speech recognition model can effectively achieve real-time dynamic alignment of audio and text sequences by adopting a recurrent neural network transducer structure, and shows significant advantages in streaming speech recognition tasks. The recurrent neural network transducer structure extracts acoustic features through an encoding module, models text context dependencies through a prediction module, and fuses multimodal information through a first fusion module, and has a powerful feature representation ability, which can significantly improve the recognition accuracy and real-time performance of the model in complex scenarios.
[0129] In the embodiments of the present disclosure, a pre-trained large language model (LLM) can refer to a model pre-trained on a large-scale text data. In some application scenarios, a large language model can also be referred to as a large language model, a text large model, etc., which are not limited herein. In one possible implementation, the pre-trained large language model can be an open-source pre-trained large language model, such as Llama, Qwen, etc. These open-source models have learned and trained on a vast amount of text, and have accumulated rich language knowledge and semantic understanding capabilities. In another possible implementation, the pre-trained large language model can be a model self-trained by developers.
[0130] The large language model shows powerful language understanding and generation capabilities through pre-training on a vast amount of text data. Compared with traditional models that rely on a large amount of data, the large language model has more advantages in data-scarce languages and complex speech tasks, and its large number of parameters makes it more robust and adaptable in noisy environments. In addition, a multi-lingual pre-trained large language model can reduce the need for specific language data and reduce the complexity of model management.
[0131] In an embodiment of the present disclosure, the large language model may include a mutually connected large language model body (LLM-body) and a large language model head (LLM-head). The large language model body may be composed of multiple layers of neural networks (such as Transformer modules), and may be responsible for performing in-depth feature extraction and semantic understanding on the input text, and obtaining general language representation capabilities through pre-training on a large amount of text data; the large language model head, as the output layer of the large language model, may be composed of a linear layer and an activation layer, and may be responsible for mapping the high-dimensional features extracted by the large language model body to the target output space (such as the probability distribution of a token table). In some application scenarios, the large language model head may also be referred to as the output layer of the large language model.
[0132] In an embodiment of the present disclosure, an audio-text classification model may be constructed based on a pre-trained speech recognition model and a pre-trained large language model. Figure 3 A schematic diagram of the audio-text classification model in the training method of the audio understanding model provided by an embodiment of the present disclosure is shown. As Figure 3 shown, the audio-text classification model may include an encoding module in the pre-trained speech recognition model, a large language model body (LLM body) in the pre-trained large language model, a first fusion module in the pre-trained speech recognition model, and a binary classification layer. The encoding module may input audio, the large language model body may input text, and the binary classification layer may output a modality prediction result.
[0133] In a possible implementation manner, an audio-text classification model may be constructed based on the pre-trained speech recognition model. Specifically, the encoding module and the first fusion module in the pre-trained speech recognition model may be retained, the prediction module in the pre-trained speech recognition model may be replaced with the large language model body, and a binary classification layer may be added after the first fusion module.
[0134] In an embodiment of the present disclosure, during the training process of the audio-text classification model, the acoustic features output by the encoding module or the semantic features output by the large language model body may be processed by the first fusion module and the binary classification layer to obtain a modality prediction result, and according to the modality prediction result and the corresponding true modality, the reverse gradient value of the first fusion module may be determined, and the parameters of the first fusion module may be updated according to the reverse gradient value of the first fusion module.
[0135] In a possible implementation, processing the acoustic features output by the encoding module or the semantic features output by the large language model body through the first fusion module and the binary classification layer to obtain a modality prediction result includes: for any audio-text pair in the training set, extracting the acoustic features corresponding to the audio sample in the audio-text pair through the encoding module in the audio-text classification model, and extracting the semantic features corresponding to the target text label sequence in the audio-text pair through the large language model body in the audio-text classification model; inputting the acoustic features or the semantic features into the first fusion module, and processing the acoustic features or the semantic features through the first fusion module to obtain a fused feature; and processing the fused feature through the binary classification layer to obtain a modality prediction result.
[0136] In this implementation, the training set may include multiple audio-text pairs. For any audio-text pair {x, y} in the training set, where x may represent an audio sample (such as a speech signal), and y may represent the target text label sequence corresponding to the audio sample (such as the annotated text content).
[0137] The audio sample x can be processed by the encoding module in the audio-text classification model to output acoustic features. The acoustic features can capture acoustic information such as the spectrum and prosody of the audio. The target text label sequence y can be processed by the large language model body in the audio-text classification model to output semantic features. The semantic features can contain context and language structure information of the text.
[0138] The acoustic features or the semantic features can be input into the first fusion module. That is, one of the acoustic features and the semantic features can be input into the first fusion module, rather than inputting both the acoustic features and the semantic features into the first fusion module simultaneously.
[0139] The first fusion module (for example, including a linear layer + activation layer) can perform a non-linear transformation on the input features to generate a fused feature. The purpose is to map features of different modalities to the same space. The fused feature can be processed by the binary classification layer (such as a fully connected layer + softmax) to output a modality prediction result. Among them, the modality prediction result can represent the probability that the current feature is determined to come from the audio modality or the text modality.
[0140] In a possible implementation manner, determining the reversed gradient value of the first fusion module according to the modal prediction result and the true modality corresponding to the modal prediction result includes: determining the value of the classification loss function corresponding to the audio-text classification model according to the modal prediction result and the true modality corresponding to the modal prediction result; calculating the gradient value of the first fusion module according to the value of the classification loss function; and reversing the gradient value of the first fusion module to obtain the reversed gradient value of the first fusion module.
[0141] In one example, if the input of the first fusion module is acoustic features (i.e., the input of the first fusion module comes from the encoding module), the label of the true modality can be 0; if the input of the first fusion module is semantic features (i.e., the input of the first fusion module comes from the large language model main body), the label of the true modality can be 1.
[0142] In one example, the cross-entropy loss function can be used to calculate the difference between the modal prediction result and the label of the true modality to obtain the value of the classification loss function:
[0143]
[0144] where y represents the label of the true modality (0 or 1), and p represents the modal prediction result (i.e., the predicted probability).
[0145] In this implementation manner, the gradient can be backpropagated from the classification loss function to the parameters of the first fusion module through the chain rule. Assuming the parameters of the first fusion module are θ, its gradient can be:
[0146]
[0147] This is the conventional gradient descent update direction, and the goal is to optimize the accuracy of the classification task.
[0148] To enable the first fusion module to generate modality-agnostic features (i.e., it is impossible to distinguish between audio or text modalities), in this implementation manner, the gradient direction can be reversed through a gradient reversal layer (GRL):
[0149]
[0150] where λ represents the reversal weight (a hyperparameter, for example, it can be set to 1).
[0151] Normal gradient descent will minimize the classification loss (i.e., improve the modality discrimination ability), while gradient reversal is equivalent to maximizing the classification loss, so that the feature output of the first fusion module cannot be distinguished by the binary classification layer in terms of modality.
[0152] In this implementation, the parameters of the first fusion module can be updated using the reversed gradient:
[0153]
[0154] Among them, η represents the learning rate. Since the gradient is reversed, the actual effect is to encourage the first fusion module to confuse the modality features rather than distinguish them.
[0155] In this implementation, through gradient reversal, the first fusion module is forced to generate features that are "indistinguishable" for the audio and text modalities, so as to more fairly fuse the two modalities in the subsequent audio understanding model and avoid training bias towards a certain modality (such as over-relying on text semantics and ignoring acoustic features).
[0156] In one possible implementation, during the training process of the audio-text classification model, the parameters of the encoding module and the large language model body are kept fixed.
[0157] The encoding module of the pre-trained speech recognition model (such as RNNT) already has the ability to extract effective acoustic features (such as spectrograms, intonations) from audio. Fixing the parameters of the encoding module can retain its pre-trained acoustic representation ability and avoid the degradation of acoustic features caused by new task training. The pre-trained large language model body has learned general language representations through a large amount of text data. Fixing its parameters can maintain its semantic understanding ability and prevent the loss of original language knowledge due to fine-tuning for audio tasks.
[0158] In this implementation, during the training process of the audio-text classification model, only the parameters of the first fusion module and the binary classification layer participate in the training. These two parts are responsible for mapping the acoustic features and semantic features to the same space and learning modality-agnostic fusion representations through the gradient reversal mechanism.
[0159] By adopting this implementation, it can be ensured that the pre-training capabilities of the encoding module and the large language model body are not interfered by new tasks. Moreover, it can significantly reduce the number of training parameters and reduce the computational overhead (especially when the number of LLM parameters is large).
[0160] In the embodiments of the present disclosure, an audio understanding model can be constructed according to the pre-trained speech recognition model, the pre-trained large language model, and the audio-text classification model. Among them, the audio understanding model can include an encoding module in the pre-trained speech recognition model, a large language model body in the pre-trained large language model, and a second fusion module. The second fusion module can include the first fusion module in the pre-trained audio-text classification model. In the audio understanding model, the second fusion module is respectively connected to the encoding module and the large language model body, and the second fusion module can be used to perform fusion processing on the output of the encoding module and the output of the large language model body.
[0161] Among them, the encoding module in the audio understanding model can be used to process audio input and output acoustic features . The large language model body in the audio understanding model can take text as input and output semantic features . In the audio understanding model, by replacing the prediction module in the pre-trained speech recognition model with the large language model body, the text context modeling ability can be enhanced. The input of the second fusion module in the audio understanding model can include the acoustic features output by the encoding module and the semantic features output by the large language model body . The second fusion module can dynamically align the acoustic features and semantic features
[0162] In a possible implementation manner, training the audio understanding model includes: for any audio sample in the training set, extracting the acoustic feature sequence corresponding to the audio sample through the encoding module in the audio understanding model, and extracting the semantic feature sequence corresponding to the audio sample through the large language model body in the audio understanding model; based on the acoustic feature sequence and the semantic feature sequence through the second fusion module, determining all valid alignment paths that can generate the target text label sequence corresponding to the audio sample, and calculating the total probability of all the valid alignment paths; taking the negative logarithm of the total probability to obtain the value of the streaming sequence alignment loss function corresponding to the audio understanding model; updating the parameters of the audio understanding model according to the value of the streaming sequence alignment loss function
[0163] In this implementation manner, during the training process of the audio understanding model, for each audio sample in the training set, the encoding module can process the original audio signal frame by frame and output an acoustic feature sequence, and the large language model body can autoregressively generate a corresponding semantic feature sequence based on the historical text labels. The acoustic feature sequence and the semantic feature sequence provide multimodal input for the subsequent second fusion module
[0164] In this implementation manner, the second fusion module can be used to determine all valid alignment paths that can generate the target text label sequence corresponding to the audio sample based on the acoustic feature sequence and the semantic feature sequence, and calculate the total probability of all the valid alignment paths
[0165] Among them, the alignment path can refer to the mapping relationship between the acoustic feature sequence (time step t) and the semantic feature sequence (text step u). For example
[0166] Horizontal movement (→): time step t+1, text step u remains unchanged (corresponding to outputting a blank symbol (blank), indicating that text has not been generated yet)
[0167] Vertical movement (↑): The time step t remains unchanged, and the text step u + 1 (corresponding to generating an actual text label).
[0168] A valid alignment path can refer to all legal paths from the starting point (t = 0, u = 0) to the ending point (t = T, u = U), and the generated text sequence is exactly the same as the target text label sequence. Among them, T is the total time step of the audio sample, and U is the length of the target text label sequence.
[0169] The second fusion module can receive the acoustic feature sequence from the encoding module and the semantic feature sequence from the large language model body , and output the joint probability distribution through a linear layer and an activation layer (such as softmax) . Among them, can represent the predicted label of the current step.
[0170] The probability of each alignment path is the product of the probabilities of all steps on the path. For example, if a path is , then its probability is . Sum the probabilities of all valid paths: .
[0171] For example, if the target text is "A B", and the audio has 3 time steps, two possible valid paths are:
[0172] Path 1: (t = 1) → (t = 2) → (t = 3↑) → (t = 3↑)
[0173] (Output: blank → blank → A → B)
[0174] Path 2: (t = 1↑) → (t = 2) → (t = 3↑)
[0175] (Output: A → blank → B)
[0176] Total probability .
[0177] In this implementation, during the training process of the audio understanding model, after calculating the total probability of all valid alignment paths, the model parameters can be optimized to maximize this total probability (i.e., minimize the negative log probability). Specifically, the negative logarithm can be taken for the total probability to obtain the value of the streaming sequence alignment loss function . The probability value range is [0, 1]. After taking the negative logarithm, the higher the probability (close to 1), the lower the loss value (close to 0); conversely, the lower the probability (close to 0), the higher the loss value (approaching infinity). This conversion transforms the probability maximization problem into a loss minimization problem, which conforms to the gradient descent optimization framework.
[0178] For example, if an audio sample has two valid alignment paths with probabilities of 0.6 and 0.4 respectively, the total probability , and the loss value ; if the path probability distribution is not ideal (such as 0.1 and 0.1), the loss value . The audio understanding model will adjust the parameters through gradient descent to increase the generation possibility of the high-probability path.
[0179] In this implementation, the gradient can be calculated based on the value of the streaming sequence alignment loss function, and the parameters of the audio understanding model can be updated through backpropagation. For example, the parameters of at least some of the second fusion module, the large language model body, and the encoding module can be updated.
[0180] In this implementation, based on the streaming sequence alignment loss function, the alignment path between the acoustic features and the semantic features can be directly optimized to ensure the dynamic matching between the text sequence generated by the audio understanding model and the audio input. By minimizing the streaming sequence alignment loss function, the second fusion module can learn to effectively combine the acoustic features of the encoding module and the semantic features of the large language model body.
[0181] In a possible implementation, the second fusion module further includes a first head; the method further includes: obtaining a target token table corresponding to the audio understanding model, where the number of tokens in the target token table is less than the number of tokens in the original token table corresponding to the large language model; according to the position indexes of the respective tokens in the target token table in the original token table, extracting the initial weights corresponding to the respective tokens in the target token table from the weight matrix of the large language model head; and initializing the weight matrix of the first head according to the initial weights corresponding to the respective tokens in the target token table.
[0182] In this implementation, the second fusion module can be composed of a first fusion module and a first head. The first head can be initialized based on the large language model head. That is, the parameters of the first head can be inherited from the pre-trained large language model head to utilize the existing language modeling ability of the pre-trained large language model.
[0183] Large language models usually have a huge output layer. For example, their token table may contain hundreds of thousands or even more tokens (such as 150,000 tokens). Such a huge token table is not efficient in speech recognition tasks because speech recognition usually only needs to focus on a limited set of tokens (such as Chinese, English, punctuation marks, and high-frequency words, etc., which may be only about 10,000). Therefore, directly using the complete token table of the large language model will lead to problems such as low training efficiency, low inference efficiency, and resource waste.
[0184] In this implementation, a target token table (tiny vocab) is constructed to reduce the dimension of the output layer, so as to optimize the training and inference efficiency of the audio understanding model. Among them, the number of tokens in the target token table is less than that in the original token table corresponding to the large language model. For example, the number of tokens in the original token table corresponding to the large language model is 150,000, and the number of tokens in the target token table is 10,000.
[0185] In this implementation, the position index of the tokens in the target token table in the original token table corresponding to the large language model can be recorded to form a mapping relationship (map), for example: {v0: V5000, v1: V9000,..., v9999: V142900}. Among them, "v0: V5000" means that index 0 (v0) in the target token table corresponds to index 5000 (V5000) in the original token table, and the two represent the same token, such as "Hello". Based on this mapping relationship, the weight matrix corresponding to the target token table can be initialized. For example, the dimension of the weight matrix of the softmax layer in the large language model head is (2048, 150000), then the dimension of the weight matrix of the softmax layer in the first head will be (2048, 10000). Among them, each row in the weight matrix corresponding to the target token table corresponds to the row with the corresponding index in the weight matrix corresponding to the original token table. For example, the first row weight of the target token table is initialized with the 5000th row of the weight matrix of the original token table, and so on. In this way, the target token table inherits the pre-trained weights of the large language model, while significantly reducing the output dimension and optimizing the training and inference efficiency of the audio understanding model.
[0186] In a possible implementation, an audio understanding model can be constructed based on the audio text classification model. Specifically, the encoding module, the large language model body, and the first fusion module in the pre-trained audio text classification model can be retained, and a first head can be connected after the first fusion module to obtain a second fusion module.
[0187] In a possible implementation, the first head further includes a dimension corresponding to the blank symbol.
[0188] In this implementation, a dimension corresponding to the blank symbol (blank) is added to the large language model head to obtain the first head. This is to adapt to the streaming alignment requirements of the speech recognition model (such as the recurrent neural network transducer), because the speech recognition model needs to process the blank symbol (indicating no output or waiting for subsequent input) during decoding.
[0189] In this implementation, the language generation capability of the large language model is retained by reusing the parameters of the large language model head, while the alignment mechanism unique to speech recognition models (such as recurrent neural network transcribers) is supported by expanding the dimension. The introduction of blank symbols enables the audio understanding model to dynamically align audio and text (such as a path moving to the right to indicate a blank, and moving upward to indicate a generated character), thereby achieving streaming decoding.
[0190] In a possible implementation, obtaining the target word list corresponding to the audio understanding model includes: performing word frequency statistics on word units in a training set; and constructing a target word list corresponding to the audio understanding model based on a preset number of word units with the highest word frequency in the training set.
[0191] In this implementation, all text data in the training set (eg, all target text label sequences) may be segmented, and the frequency of occurrence of each word unit may be counted.
[0192] As an example of this implementation, the tokenizer of a large language model can be used to tokenize the text data in the training set. For example, the Chinese text "Hello, how is the weather today" can be decomposed by the tokenizer into "Hello|,|how is the weather today|", while the English text "good morning" can be decomposed into "good| |mor@|ning|". Here, each part separated by a vertical line "|" represents a separate token. In this way, the text data can be converted into discrete units that can be processed by the audio understanding model, providing a basis for the subsequent word frequency statistics and the construction of the target token table.
[0193] In this implementation, based on the word frequency statistics, a preset number of word units (eg, 10,000) with the highest frequency may be selected to construct a target word unit table, which are usually the most commonly used word units in speech recognition tasks.
[0194] In a possible implementation, the audio understanding model also includes a second head, which is connected to the large language model body and initialized according to the large language model head; the method also includes: determining the value of a knowledge distillation loss function based on a first text prediction result output by the first head and a second text prediction result output by the second head; and updating the parameters of the audio understanding model according to the value of the knowledge distillation loss function.
[0195] In this implementation, the second head can directly reuse the original head structure of the pre-trained large language model. For example, it can include a linear layer and a softmax layer to inherit the powerful language generation ability of the large language model. The parameters of the second head can be copied from the head of the large language model but are independent of the first head, forming a dual-task output branch during the training process of the audio understanding model. Moreover, the second head can be connected to the large language model body and can directly generate the second text prediction result (such as the probability distribution of tokens) corresponding to the audio sample based on the semantic features output by the large language model body.
[0196] In this implementation, the output layer of the second head has a large dimension (e.g., 150,000) and has powerful text generation ability; the output layer of the first head has a small dimension (e.g., 10,000) and is suitable for speech recognition tasks. Knowledge distillation can be used to make the first head learn the output distribution of the second head, thereby inheriting the knowledge of the large language model.
[0197] According to the mapping relationship between the target token table and the original token table, the predicted probability distribution of each token in the target token table can be extracted from the second text prediction result output by the second head. Based on the predicted probability distribution of each token in the target token table in the second text prediction result and the predicted probability distribution of each token in the target token table in the first text prediction result, the value of the knowledge distillation loss function can be determined. In one example, the value of the knowledge distillation loss function can be determined by calculating the Kullback-Leibler Divergence (KL divergence). The KL divergence measures the difference between two probability distributions. By minimizing this difference, the output of the first head can be made closer to the output of the second head. In this way, the audio understanding model can inherit the powerful language generation ability of the large language model while maintaining efficient inference.
[0198] In one possible implementation, the method further includes: determining the value of the token distribution loss function corresponding to the audio understanding model according to the second text prediction result and the target text label sequence corresponding to the second text prediction result; updating the parameters of the audio understanding model according to the value of the token distribution loss function.
[0199] In this implementation, by calculating the token distribution loss function (such as cross-entropy loss) between the second text prediction result and the true label (i.e., the target text label sequence), additional supervision signals can be provided to strengthen the language modeling ability.
[0200] In one possible implementation, during the training process of the audio understanding model, the parameters of the second head remain fixed.
[0201] In this implementation, during the training process of the audio understanding model, the parameters of the second head can be kept fixed (i.e., not participating in gradient updates). Thus, while introducing the pre-trained knowledge of the large language model, the interference of its text generation ability by the speech alignment task can be avoided. The second head directly reuses the parameters of the original head of the large language model. By freezing its weights, it ensures that the audio understanding model retains the powerful language modeling ability of the large language model. This design not only utilizes the prior knowledge of the large language model to improve the text prediction quality but also prevents parameter conflicts during multi-task training, enabling the audio understanding model to converge more stably.
[0202] Figure 4 A schematic diagram showing the audio understanding model provided by an embodiment of the present disclosure. As Figure 4 shown, the audio understanding model may include an encoding module, an LLM body, a first fusion module, a first head, and a second head. Among them, the input of the encoding module can be audio, and the input of the LLM body can be text. According to the output of the first head, the value of the streaming sequence alignment loss function can be calculated. According to the outputs of the first head and the second head, the value of the knowledge distillation loss function can be calculated. According to the output of the second head, the value of the token distribution loss function can be calculated.
[0203] In a possible implementation, the audio understanding model is trained in a first stage and a second stage based on a training set; wherein, during the first stage training, the parameters of the encoding module are kept fixed; during the second stage training, the parameters of the encoding module are adjusted.
[0204] In this implementation, a two-stage training strategy can be adopted to optimize the audio understanding model.
[0205] In the first - stage training, the parameters of the encoding module can be kept fixed to utilize the acoustic feature extraction ability that the pre - trained encoding module has already learned. Meanwhile, the parameters of other parts of the audio understanding model, such as the large - language model main body and the second fusion module, are adjusted. Among them, the encoding module in the pre - trained speech recognition model already has the ability to extract effective acoustic features from audio signals. The acoustic features can include the spectral information of the audio, prosodic patterns, etc., which can provide a basis for subsequent text generation. In the first stage, by keeping the parameters of the encoding module fixed, excessive adjustment of the encoding module in the preliminary training stage can be avoided, thus retaining its general acoustic feature extraction ability learned in the pre - training stage. This can ensure that the audio understanding model can still effectively extract acoustic features when processing new audio data. In the first stage, the focus of training is to combine the pre - trained large - language model main body with the encoding module, and align and fuse the acoustic features and semantic features through the second fusion module. At this time, the parameters of the large - language model main body and the second fusion module are mainly updated to enable the audio understanding model to better understand and generate text.
[0206] In one example, the audio understanding model obtained from the first - stage training can be called a non - streaming audio understanding model.
[0207] In the second - stage training, the parameters of the encoding module can be unfrozen to optimize the encoding module's ability to extract local acoustic features. That is, in the second stage, the audio understanding model can update the parameters of the encoding module, the large - language model main body, and the second fusion module simultaneously to achieve more accurate alignment and fusion of audio and text.
[0208] In one example, the audio understanding model obtained from the second - stage training can be called a streaming audio understanding model.
[0209] In this implementation, through staged training, the gradient conflict caused by optimizing all modules simultaneously can be avoided. Through progressive training, the encoding module is transitioned from global feature extraction to local real - time processing to be applicable to the streaming audio understanding scenario.
[0210] In one possible implementation, the method further includes: in the second - stage training, the audio sample is sliced into multiple audio segments according to a preset duration, and the multiple audio segments are sequentially input into the encoding module in the audio understanding model; the audio understanding model is restricted by an attention mask to only obtain the information of the current audio segment and the historical audio segments before the current audio segment when processing the current audio segment.
[0211] In this implementation, during the second-stage training, to achieve the ability of streaming audio processing, the audio understanding model can adopt a chunk training strategy combined with attention mask control. Specifically, the complete audio sample can be segmented into consecutive audio segments according to a preset duration (such as 320 ms / block), and then sequentially input into the encoding module. For example, a 3-second audio (3000 ms) can be segmented into approximately 9 audio segments of 320 ms (padding zeros for the insufficient part at the end) and processed in order. Moreover, in the self-attention layer of the audio understanding model, a unidirectional mask can be applied, so that the audio understanding model can only access the information of the current audio segment and historical audio segments, while the information of future audio segments is completely masked. Thus, the causality of the streaming scenario can be simulated, enabling the audio understanding model to output results in real time only relying on the received audio data.
[0212] The encoding module can learn to extract effective features from the local part of the audio sample (i.e., the audio segment), rather than relying on the complete context. Through segmented training, the encoding module can gradually optimize the ability to capture short-time acoustic patterns (such as phonemes, syllables).
[0213] In a possible implementation, during the training process of the audio understanding model, the first head is trained with all parameters, and the parameters of the large language model body are updated by fine-tuning.
[0214] In this implementation, in the audio understanding model, the first head is directly responsible for text generation (such as generating the token probability distribution), and needs to quickly adapt to the special requirements of the speech recognition task (such as increasing the dimension of the blank symbol). By training the first head with all parameters, that is, by completely updating the weights of the first head through gradient descent, the output layer of the audio understanding model can flexibly learn the language pattern after audio-text alignment.
[0215] The large language model body has already possessed the general semantic understanding ability through massive text pre-training. Directly training with all parameters is likely to lead to overfitting (especially when the audio data is scarce). In this implementation, parameter-efficient fine-tuning techniques (such as LoRA, Adapter) can be used to train only a small number of newly added low-rank matrices or adapter layers, and freeze the original parameters of the large language model body. Thus, while retaining the pre-trained knowledge, the feature representation can be slightly adjusted to adapt to the audio context.
[0216] In this implementation, by training the first head with all parameters, the output distribution can be accurately optimized, improving the recognition accuracy. By fine-tuning the parameters of the large language model body, the core ability of the pre-trained language model can be avoided from being damaged, reducing the training risk. Moreover, compared with full-model training, the computational cost can be significantly reduced (especially when the number of parameters of the large language model body is huge).
[0217] An embodiment of the present disclosure also provides an audio understanding method, including: obtaining an audio understanding model trained by the training method of the audio understanding model; inputting the audio to be processed into the audio understanding model, and outputting a text prediction result corresponding to the audio to be processed through the audio understanding model.
[0218] In an embodiment of the present disclosure, an optimized audio understanding model can be obtained through the training method of the audio understanding model described above. This audio understanding model integrates the acoustic feature extraction ability of the encoding module in the pre-trained speech recognition model, the semantic understanding ability of the large language model, and the ability of the first fusion module to fuse the audio modality and the text modality. In practical applications, the audio to be processed (such as a real-time speech stream or a complete recording) can be input into this audio understanding model. The audio understanding model extracts acoustic features through the encoding module, and jointly performs semantic reasoning with the main body of the large language model. Finally, the second fusion module dynamically aligns the audio and text sequences and outputs the corresponding text prediction result. The audio understanding model provided by the embodiment of the present disclosure supports streaming processing, can generate verbatim text (such as subtitles) in real time, and can also process non-streaming audio, and is applicable to scenarios such as intelligent assistants and conference transcription.
[0219] In a possible implementation manner, before the audio to be processed is input into the audio understanding model, the method further includes: deleting the second head in the audio understanding model. In this implementation manner, the second head is only used to assist in optimizing the text generation task during the training phase. Removing the second head can reduce the redundant calculation amount and reduce the inference latency, which is particularly crucial for real-time streaming processing (such as a voice assistant), and at the same time ensures that the audio understanding model only retains the necessary output paths.
[0220] In a possible implementation manner, the audio to be processed is streaming audio or non-streaming audio. In this implementation manner, according to the actual business scenario, the audio understanding model can be used for the inference of streaming audio or non-streaming audio.
[0221] In a possible implementation manner, the audio understanding model can use beam search for decoding to find the optimal output path. Among them, beam search is a heuristic search algorithm that can maintain a set of candidate solutions at each step. These candidate solutions are sorted according to a certain scoring mechanism, and usually several solutions with the highest scores are selected as the candidates for the next step.
[0222] Figure 5 A schematic diagram showing the dynamic alignment process of the audio understanding model provided by the embodiment of the present disclosure in the decoding stage. Figure 5 It can show the streaming alignment mechanism between the acoustic feature sequence (time step t) and the semantic feature sequence (text step u).
[0223] In Figure 5 In it, the horizontal axis (t) can represent the time steps of the audio input (such as a sequence of audio frames), and the acoustic feature A output by the encoding module t (black circles). The vertical axis (u) can represent the number of steps of text generation (such as the words or characters that have been output), and the semantic feature T output by the large language model body u (white circles). The node y(t,u) can represent the joint decoding result at time step t and text step u, which may be the actual text label (such as Chinese characters, words) or a blank symbol (∅). Among them, the blank symbol indicates that no text is generated at the current time step, and only the time step is advanced.
[0224] Figure 5 The red arrow in is an example of the decoding path. A horizontal move (→) represents time step t+1, and the text step u remains unchanged (output a blank symbol, such as ∅(1,2)). A vertical move (↑) represents that the time step t remains unchanged, and the text step u+1 (output the actual text, such as y(3,2)).
[0225] In this way, the audio understanding model can align the acoustic features with the semantic features in real time when processing the streaming audio input and generate the corresponding text results. This alignment ability not only improves the interaction experience but also significantly enhances the performance of the streaming speech recognition system.
[0226] The training method and audio understanding method of the audio understanding model provided by the embodiments of the present disclosure can be applied to technical fields such as artificial intelligence, speech recognition, audio understanding large models, and streaming speech recognition, which are not limited herein. Moreover, the audio understanding model trained by the training method of the audio understanding model provided by the embodiments of the present disclosure can be applied to scenarios such as real-time captions, voice search, and smart homes, which are not limited herein.
[0227] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to space limitations, the present disclosure will not elaborate further. Those skilled in the art can understand that in the above methods of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic.
[0228] In addition, the present disclosure also provides a training device for the audio understanding model, an audio understanding device, a non-volatile computer-readable storage medium, and a computer program product, all of which can be used to implement any training or audio understanding method provided by the present disclosure. The corresponding technical solutions and technical effects can be seen in the corresponding records in the method part and will not be elaborated further.
[0229] Figure 6 shows a block diagram of the training device for the audio understanding model provided by the embodiments of the present disclosure. As Figure 6 As shown in the figure, the training device of the audio understanding model includes:
[0230] A first acquisition module 61, configured to acquire a pre-trained speech recognition model and a pre-trained large language model; wherein, the speech recognition model includes an encoding module, a prediction module, and a first fusion module, the first fusion module is respectively connected to the encoding module and the prediction module, and the large language model includes a large language model main body and a large language model head that are connected to each other;
[0231] A construction module 62, configured to construct an audio text classification model according to the encoding module, the large language model main body, the first fusion module, and a binary classification layer, wherein the first fusion module is respectively connected to the encoding module and the large language model main body, and the binary classification layer is connected to the first fusion module;
[0232] A modality prediction module 63, configured to process the acoustic features output by the encoding module or the semantic features output by the large language model main body through the first fusion module and the binary classification layer to obtain a modality prediction result;
[0233] A first training module 64, configured to determine an inversion gradient value of the first fusion module according to the modality prediction result and the true modality corresponding to the modality prediction result, and update the parameters of the first fusion module according to the inversion gradient value of the first fusion module to train the audio text classification model;
[0234] A second training module 65, configured to, in response to the audio text classification model being trained to meet a preset condition, construct an audio understanding model according to the encoding module, the large language model main body, and a second fusion module, and train the audio understanding model, wherein the second fusion module includes the first fusion module.
[0235] In a possible implementation manner, the first training module 64 is configured to:
[0236] Determine the value of the classification loss function corresponding to the audio text classification model according to the modality prediction result and the true modality corresponding to the modality prediction result;
[0237] Calculate the gradient value of the first fusion module according to the value of the classification loss function;
[0238] Invert the gradient value of the first fusion module to obtain the inversion gradient value of the first fusion module.
[0239] In a possible implementation manner, the modality prediction module 63 is configured to:
[0240] For any audio-text pair in the training set, extract the acoustic features corresponding to the audio sample in the audio-text pair through the encoding module in the audio-text classification model, and extract the semantic features corresponding to the target text label sequence in the audio-text pair through the large language model body in the audio-text classification model;
[0241] Input the acoustic features or the semantic features into the first fusion module, and process the acoustic features or the semantic features through the first fusion module to obtain fusion features;
[0242] Process the fusion features through the binary classification layer to obtain a modality prediction result.
[0243] In a possible implementation manner, during the training process of the audio-text classification model, the parameters of the encoding module and the large language model body are kept fixed.
[0244] In a possible implementation manner, the second training module 65 is used for:
[0245] For any audio sample in the training set, extract the acoustic feature sequence corresponding to the audio sample through the encoding module in the audio understanding model, and extract the semantic feature sequence corresponding to the audio sample through the large language model body in the audio understanding model;
[0246] Through the second fusion module, based on the acoustic feature sequence and the semantic feature sequence, determine all valid alignment paths capable of generating the target text label sequence corresponding to the audio sample, and calculate the total probability of all the valid alignment paths;
[0247] Take the negative logarithm of the total probability to obtain the value of the streaming sequence alignment loss function corresponding to the audio understanding model;
[0248] Update the parameters of the audio understanding model according to the value of the streaming sequence alignment loss function.
[0249] In a possible implementation manner, the second fusion module further includes a first head;
[0250] The device further includes:
[0251] A second obtaining module, configured to obtain the target token table corresponding to the audio understanding model, where the number of tokens in the target token table is less than the number of tokens in the original token table corresponding to the large language model;
[0252] An extraction module, configured to extract the initial weights corresponding to each token in the target token table from the weight matrix of the head of the large language model according to the position indexes of the tokens in the target token table in the original token table;
[0253] An initialization module, configured to initialize the weight matrix of the first head according to the initial weights corresponding to each token in the target token table.
[0254] In a possible implementation, the first head further includes a dimension corresponding to a blank symbol.
[0255] In a possible implementation, the second obtaining module is configured to:
[0256] Perform word frequency statistics on the tokens in the training set;
[0257] Construct a target token table corresponding to the audio understanding model according to a preset number of tokens with the highest word frequencies in the training set.
[0258] In a possible implementation, the audio understanding model further includes a second head, the second head is connected to the large language model body, and the second head is initialized according to the head of the large language model;
[0259] The apparatus further includes:
[0260] A first determination module, configured to determine the value of the knowledge distillation loss function according to the first text prediction result output by the first head and the second text prediction result output by the second head;
[0261] A first update module, configured to update the parameters of the audio understanding model according to the value of the knowledge distillation loss function.
[0262] In a possible implementation, the apparatus further includes:
[0263] A second determination module, configured to determine the value of the token distribution loss function corresponding to the audio understanding model according to the second text prediction result and the target text label sequence corresponding to the second text prediction result;
[0264] A second update module, configured to update the parameters of the audio understanding model according to the value of the token distribution loss function.
[0265] In a possible implementation, during the training process of the audio understanding model, the parameters of the second head are kept fixed.
[0266] In a possible implementation, the audio understanding model is trained in a first stage and a second stage based on a training set;
[0267] Among them, in the training of the first stage, the parameters of the encoding module are kept fixed; in the training of the second stage, the parameters of the encoding module are adjusted.
[0268] In a possible implementation, the speech recognition model uses a recurrent neural network transducer.
[0269] According to one aspect of the present disclosure, there is provided an audio understanding device, including:
[0270] An acquisition module, configured to acquire an audio understanding model trained by a training device of the audio understanding model;
[0271] An audio understanding module, configured to input an audio to be processed into the audio understanding model, and output a text prediction result corresponding to the audio to be processed through the audio understanding model.
[0272] The embodiments of the present disclosure further provide a training device for an audio understanding model, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the above-mentioned training method of the audio understanding model.
[0273] The embodiments of the present disclosure further provide an audio understanding device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the above-mentioned audio understanding method.
[0274] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation and technical effects can refer to the descriptions of the above method embodiments. For the sake of brevity, they will not be elaborated here.
[0275] The embodiments of the present disclosure further provide a computer-readable storage medium, on which computer program instructions are stored, and the computer program instructions, when executed by a processor, implement the above method. Among them, the computer-readable storage medium can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium.
[0276] The embodiments of the present disclosure further propose a computer program, including computer-readable code, and when the computer-readable code runs in an electronic device, the processor in the electronic device executes the above method.
[0277] The embodiments of the present disclosure further provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, and when the computer-readable code runs in an electronic device, the processor in the electronic device executes the above method.
[0278] Figure 7 It is a block diagram of a training device for an audio understanding model or an audio understanding device 1 shown according to an exemplary embodiment. For example, device 1900 may be provided as a server or a terminal device. Referring to Figure 7 , device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0279] Device 1900 may also include a power supply component 1926 configured to perform power management of device 1900, a wired or wireless network interface 1950 configured to connect device 1900 to a network, and an input / output interface 1958 (I / O interface). Device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server TM , MacOS X TM , Unix TM , Linux TM , FreeBSD TM or the like.
[0280] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions, and the above computer program instructions can be executed by the processing component 1922 of the device 1900 to complete the above method.
[0281] A computer-readable storage medium can be a tangible device that can hold and store programs / instructions used by an instruction execution device. The computer-readable storage medium can be, for example -- but not limited to -- an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as being an instantaneous signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0282] The computer programs (or computer-readable program instructions) described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0283] A computer program (or computer program instructions) for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on a user's computer, partially on a user's computer, executed as a stand - alone software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, may be connected to an external computer (e.g., via an Internet service provider through the Internet). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0284] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0285] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0286] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0287] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.
[0288] The computer program product may be implemented specifically by hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is embodied as a computer storage medium. In another alternative embodiment, the computer program product is embodied as a software product, such as a Software Development Kit (SDK), and so on.
[0289] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or likenesses can be referred to each other. For the sake of brevity, they will not be elaborated herein.
[0290] If the technical solution of the embodiments of the present disclosure involves personal information, before the product applying the technical solution of the embodiments of the present disclosure processes personal information, it has clearly informed the personal information processing rules and obtained the personal's independent consent. If the technical solution of the embodiments of the present disclosure involves sensitive personal information, before the product applying the technical solution of the embodiments of the present disclosure processes sensitive personal information, it has obtained the personal's separate consent and at the same time meets the requirement of "express consent". For example, at a personal information collection device such as a camera, a clear and prominent logo is set to inform that the personal information collection scope has been entered and personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that the individual consents to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are informed by obvious logos / information, personal authorization is obtained through pop-up messages or by asking the individual to upload their personal information by themselves; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.
[0291] The embodiments of the present disclosure have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the art in the technical field to understand the disclosed embodiments.< / eos> < / eos> < / eos> < / sos> < / sos>
Claims
1. A training method for an audio understanding model, characterized in that, Including: Obtain a pre-trained speech recognition model and a pre-trained large language model; wherein, the speech recognition model includes an encoding module, a prediction module, and a first fusion module, the first fusion module is respectively connected to the encoding module and the prediction module, and the large language model includes a large language model main body and a large language model head that are connected to each other; Construct an audio text classification model according to the encoding module, the large language model main body, the first fusion module, and a binary classification layer, wherein the first fusion module is respectively connected to the encoding module and the large language model main body, and the binary classification layer is connected to the first fusion module; Process the acoustic features output by the encoding module or the semantic features output by the large language model main body through the first fusion module and the binary classification layer to obtain a modality prediction result; Determine the reverse gradient value of the first fusion module according to the modality prediction result and the true modality corresponding to the modality prediction result, and update the parameters of the first fusion module according to the reverse gradient value of the first fusion module to train the audio text classification model; In response to the audio text classification model being trained to meet a preset condition, construct an audio understanding model according to the encoding module, the large language model main body, and a second fusion module, and train the audio understanding model, wherein the second fusion module includes the first fusion module.
2. The method according to claim 1, wherein The determining the reverse gradient value of the first fusion module according to the modality prediction result and the true modality corresponding to the modality prediction result includes: Determine the value of the classification loss function corresponding to the audio text classification model according to the modality prediction result and the true modality corresponding to the modality prediction result; Calculate the gradient value of the first fusion module according to the value of the classification loss function; Reverse the gradient value of the first fusion module to obtain the reverse gradient value of the first fusion module.
3. The method according to claim 1, wherein The processing the acoustic features output by the encoding module or the semantic features output by the large language model main body through the first fusion module and the binary classification layer to obtain a modality prediction result includes: For any audio text pair in the training set, extract the acoustic features corresponding to the audio sample in the audio text pair through the encoding module in the audio text classification model, and extract the semantic features corresponding to the target text label sequence in the audio text pair through the large language model main body in the audio text classification model; Input the acoustic features or the semantic features into the first fusion module, and process the acoustic features or the semantic features through the first fusion module to obtain a fusion feature; Process the fusion feature through the binary classification layer to obtain a modality prediction result.
4. The method according to claim 1, wherein During the training process of the audio text classification model, the parameters of the encoding module and the large language model main body remain fixed.
5. The method according to claim 1, wherein The training the audio understanding model includes: For any audio sample in the training set, extract the acoustic feature sequence corresponding to the audio sample through the encoding module in the audio understanding model, and extract the semantic feature sequence corresponding to the audio sample through the large language model main body in the audio understanding model; Based on the acoustic feature sequence and the semantic feature sequence, the second fusion module determines all valid alignment paths capable of generating the target text label sequence corresponding to the audio sample, and calculates the total probability of all the valid alignment paths; Take the negative logarithm of the total probability to obtain the value of the streaming sequence alignment loss function corresponding to the audio understanding model; Update the parameters of the audio understanding model according to the value of the streaming sequence alignment loss function.
6. The method according to claim 1, wherein The second fusion module further includes a first head; The method further includes: Obtain the target token table corresponding to the audio understanding model, where the number of tokens in the target token table is less than the number of tokens in the original token table corresponding to the large language model; According to the position indexes of the respective tokens in the target token table in the original token table, extract the initial weights corresponding to the respective tokens in the target token table from the weight matrix of the large language model head; Initialize the weight matrix of the first head according to the initial weights corresponding to the respective tokens in the target token table.
7. The method according to claim 6, characterized in that, The first head further includes a dimension corresponding to the blank symbol.
8. The method according to claim 6, wherein The obtaining the target token table corresponding to the audio understanding model includes: Perform word frequency statistics on the tokens in the training set; Construct the target token table corresponding to the audio understanding model according to the preset number of tokens with the highest word frequencies in the training set.
9. The method according to claim 6, characterized in that, The audio understanding model further includes a second head, the second head is connected to the large language model main body, and the second head is initialized according to the large language model head; The method further includes: Determine the value of the knowledge distillation loss function according to the first text prediction result output by the first head and the second text prediction result output by the second head; Update the parameters of the audio understanding model according to the value of the knowledge distillation loss function.
10. The method according to claim 9, wherein The method further includes: Determine the value of the token distribution loss function corresponding to the audio understanding model according to the second text prediction result and the target text label sequence corresponding to the second text prediction result; Update the parameters of the audio understanding model according to the value of the token distribution loss function.
11. The method according to claim 9, wherein During the training process of the audio understanding model, the parameters of the second head are kept fixed.
12. The method according to claim 1, characterized in that, The audio understanding model is trained in a first stage and a second stage based on the training set; Wherein, in the first stage of training, the parameters of the encoding module are kept fixed; in the second stage of training, the parameters of the encoding module are adjusted.
13. The method according to claim 1, wherein The speech recognition model uses a recurrent neural network transducer.
14. An audio understanding method, characterized in that, Includes: Obtain an audio understanding model trained by the training method of the audio understanding model according to any one of claims 1 to 13; Input the audio to be processed into the audio understanding model, and output the text prediction result corresponding to the audio to be processed through the audio understanding model.
15. A training device for an audio understanding model, characterized in that, Includes: A first acquisition module, configured to acquire a pre-trained speech recognition model and a pre-trained large language model; wherein, the speech recognition model includes an encoding module, a prediction module, and a first fusion module, the first fusion module is respectively connected to the encoding module and the prediction module, and the large language model includes a large language model main body and a large language model head that are connected to each other; A construction module, configured to construct an audio text classification model according to the encoding module, the large language model main body, the first fusion module, and a binary classification layer, wherein the first fusion module is respectively connected to the encoding module and the large language model main body, and the binary classification layer is connected to the first fusion module; A modality prediction module, configured to process the acoustic features output by the encoding module or the semantic features output by the large language model main body through the first fusion module and the binary classification layer to obtain a modality prediction result; A first training module, configured to determine an inversion gradient value of the first fusion module according to the modality prediction result and the true modality corresponding to the modality prediction result, and update the parameters of the first fusion module according to the inversion gradient value of the first fusion module to train the audio text classification model; A second training module, configured to, in response to the audio text classification model being trained to meet a preset condition, construct an audio understanding model according to the encoding module, the large language model main body, and a second fusion module, and train the audio understanding model, wherein the second fusion module includes the first fusion module.
16. An audio understanding device, characterized in that, Comprising: An acquisition module, configured to acquire an audio understanding model trained by the training device of the audio understanding model according to claim 15; An audio understanding module, configured to input the audio to be processed into the audio understanding model, and output a text prediction result corresponding to the audio to be processed through the audio understanding model.
17. A training device for an audio understanding model, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 13.
18. An audio understanding device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to claim 14.
19. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, The computer program instructions, when executed by a processor, implement the method according to any one of claims 1 to 14.
20. A computer program product, comprising computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, characterized in that, When the computer-readable code runs in an electronic device, the processor in the electronic device executes the method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Speech recognition model training method, speech recognition method and speech recognition system
CN113936647A
Speech recognition method and device, medium and computing equipment
CN115064153A
Speech recognition model training method and device, speech recognition method and device, equipment and medium
CN117711386A
Speech recognition method based on large language model
CN118447827A
Speech recognition model training method and device, speech recognition method and device and storage medium
CN118711573A
Cited By
Multi-modal generative dialogue task processing method, device and equipment
CN120932648A
Speech recognition method and device, electronic equipment, storage medium and program product
CN121438812A
Training method and device of voice large model, equipment and medium
CN121438813A