Audio understanding model training method and device, audio understanding method and device, storage medium and program product

By building an audio understanding model, using pre-trained speech recognition model and large language model, the acoustic and semantic feature sequences of audio samples are extracted, the effective alignment path is determined and the total probability is calculated, which solves the problem of low training efficiency of streaming audio understanding model, and significantly improves speech recognition accuracy with less audio training data.

CN120544542AActive Publication Date: 2025-08-26MOORE THREADS TECH CO LTD

Patent Information

Application Number
CN202510779304.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-08-26
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

In the prior art, the training efficiency of streaming audio understanding models is low and the processing effect is poor, making it difficult to improve speech recognition accuracy with less audio training data.

Method used

By building an audio understanding model, using pre-trained speech recognition model and large language model, the acoustic and semantic feature sequences of audio samples are extracted, the effective alignment path is determined and the total probability is calculated, the model parameters are updated, and the powerful language understanding and generation capabilities of large language models are combined to achieve efficient streaming speech recognition.

Benefits of technology

Even with less audio training data, the accuracy of speech recognition and the processing effect of streaming audio understanding are significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544542A_ABST
    Figure CN120544542A_ABST
Patent Text Reader

Abstract

The invention relates to an audio understanding model training method, an audio understanding method, an audio understanding device, a storage medium and a program product. The method comprises the following steps: obtaining a pre-trained speech recognition model and a large language model; the speech recognition model comprises a coding module, a prediction module and a first fusion module; according to the speech recognition model and the large-scale language model, an audio understanding model is constructed, and the audio understanding model comprises a coding module, a large-scale language model main body and a second fusion module; extracting an acoustic feature sequence corresponding to the audio sample through a coding module, and extracting a semantic feature sequence corresponding to the audio sample through a large language model main body; determining all effective alignment paths capable of generating a target text tag sequence corresponding to the audio sample based on the acoustic feature sequence and the semantic feature sequence through a second fusion module, and calculating the total probability of all the effective alignment paths; and updating parameters of the audio understanding model according to the total probability. The speech recognition precision can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a training method for an audio understanding model, an audio understanding method, a training device for an audio understanding model, an audio understanding device, a non-volatile computer-readable storage medium, and a computer program product. Background Art

[0002] Audio understanding technology uses artificial intelligence models to perform semantic analysis and content recognition on audio signals, converting speech content into understandable text. Audio understanding can be categorized into non-streaming and streaming audio understanding. Non-streaming processing analyzes the entire audio stream and outputs results, suitable for scenarios such as meeting transcript analysis. Streaming processing analyzes audio in real time and provides dynamic feedback, widely used in applications requiring instant interaction, such as intelligent voice assistants, real-time subtitle generation, and smart homes. Improving the training efficiency and processing performance of streaming audio understanding models is a key research topic. Summary of the Invention

[0003] In view of this, the present disclosure provides an audio understanding technical solution.

[0004] According to one aspect of the present disclosure, a method for training an audio understanding model is provided, comprising:

[0005] Obtaining a pre-trained speech recognition model and a pre-trained large language model; wherein the speech recognition model includes an encoding module, a prediction module, and a first fusion module, the first fusion module being connected to the encoding module and the prediction module respectively; and the large language model includes a large language model body and a large language model head that are connected to each other;

[0006] Constructing an audio understanding model based on the speech recognition model and the large language model, wherein the audio understanding model includes the encoding module, the large language model body, and a second fusion module, and the second fusion module is connected to the encoding module and the large language model body respectively;

[0007] For any audio sample in the audio training set, extract the acoustic feature sequence corresponding to the audio sample through the encoding module in the audio understanding model, and extract the semantic feature sequence corresponding to the audio sample through the large language model body in the audio understanding model;

[0008] Determining, by the second fusion module, all valid alignment paths that can generate a target text label sequence corresponding to the audio sample based on the acoustic feature sequence and the semantic feature sequence, and calculating a total probability of all valid alignment paths;

[0009] According to the total probability, the parameters of the audio understanding model are updated.

[0010] In a possible implementation, the second fusion module includes the first fusion module and a first header, wherein the first header is initialized according to the large language model header, and the first header adds a dimension corresponding to the blank symbol blank on the basis of the large language model header.

[0011] In a possible implementation, the method further includes:

[0012] During the pre-training process, the output dimension of the first fusion module is controlled to be the same as the input dimension of the large language model head.

[0013] In a possible implementation, updating the parameters of the audio understanding model according to the total probability includes:

[0014] Taking the negative logarithm of the total probability to obtain a value of a first loss function corresponding to the audio understanding model;

[0015] Update the parameters of the audio understanding model according to the value of the first loss function.

[0016] In one possible implementation, the audio understanding model further includes a second head, the second head is connected to the large language model body, and the second head is initialized according to the large language model head;

[0017] The method further comprises:

[0018] Outputting a text prediction result corresponding to the audio sample through the second head;

[0019] Determining a value of a second loss function corresponding to the audio understanding model according to a text prediction result corresponding to the audio sample and the target text label sequence;

[0020] Update the parameters of the audio understanding model according to the value of the second loss function.

[0021] In a possible implementation, during the training of the audio understanding model, parameters of the second head remain fixed.

[0022] In a possible implementation, the audio understanding model performs first-stage training and second-stage training based on the audio training set;

[0023] In the first stage of training, the parameters of the encoding module remain fixed; in the second stage of training, the parameters of the encoding module are adjusted.

[0024] In a possible implementation, the method further includes:

[0025] In the second stage of training, the audio sample is divided into multiple audio segments according to a preset duration, and the multiple audio segments are sequentially input into the encoding module of the audio understanding model;

[0026] The attention mask is used to restrict the audio understanding model from only obtaining information about the current audio segment and historical audio segments before the current audio segment when processing the current audio segment.

[0027] In a possible implementation, during the training of the audio understanding model, the first head performs full parameter training, and the parameters of the large language model body are updated by fine-tuning.

[0028] In one possible implementation, the speech recognition model uses a recurrent neural network transcriber.

[0029] According to another aspect of the present disclosure, there is provided an audio understanding method, comprising:

[0030] Obtaining an audio understanding model trained using the audio understanding model training method;

[0031] The audio to be processed is input into the audio understanding model, and the audio understanding model outputs the text prediction result corresponding to the audio to be processed.

[0032] In a possible implementation, before inputting the audio to be processed into the audio understanding model, the method further includes:

[0033] The second head in the audio understanding model is deleted.

[0034] In a possible implementation, the audio to be processed is streaming audio or non-streaming audio.

[0035] According to another aspect of the present disclosure, a training device for an audio comprehension model is provided, comprising:

[0036] An acquisition module, configured to obtain a pre-trained speech recognition model and a pre-trained large language model; wherein the speech recognition model includes an encoding module, a prediction module, and a first fusion module, the first fusion module being connected to the encoding module and the prediction module, respectively; and the large language model includes a large language model body and a large language model head connected to each other;

[0037] A construction module, configured to construct an audio understanding model based on the speech recognition model and the large language model, wherein the audio understanding model includes the encoding module, the large language model body, and a second fusion module, wherein the second fusion module is connected to the encoding module and the large language model body respectively;

[0038] An extraction module is configured to extract, for any audio sample in the audio training set, an acoustic feature sequence corresponding to the audio sample using the encoding module in the audio understanding model, and to extract a semantic feature sequence corresponding to the audio sample using the large language model body in the audio understanding model;

[0039] a calculation module, configured to determine, through the second fusion module, based on the acoustic feature sequence and the semantic feature sequence, all valid alignment paths that can generate a target text label sequence corresponding to the audio sample, and calculate a total probability of all the valid alignment paths;

[0040] A first updating module is used to update the parameters of the audio understanding model according to the total probability.

[0041] In a possible implementation, the second fusion module includes the first fusion module and a first header, wherein the first header is initialized according to the large language model header, and the first header adds a dimension corresponding to the blank symbol blank on the basis of the large language model header.

[0042] In a possible implementation, the apparatus further includes:

[0043] A control module is used to control the output dimension of the first fusion module to be the same as the input dimension of the head of the large language model during pre-training.

[0044] In a possible implementation, the first update module is configured to:

[0045] Taking the negative logarithm of the total probability to obtain a value of a first loss function corresponding to the audio understanding model;

[0046] Update the parameters of the audio understanding model according to the value of the first loss function.

[0047] In one possible implementation, the audio understanding model further includes a second head, the second head is connected to the large language model body, and the second head is initialized according to the large language model head;

[0048] The device further comprises:

[0049] A second output module, configured to output a text prediction result corresponding to the audio sample through the second header;

[0050] A determination module, configured to determine a value of a second loss function corresponding to the audio understanding model based on a text prediction result corresponding to the audio sample and the target text label sequence;

[0051] A second updating module is used to update the parameters of the audio understanding model according to the value of the second loss function.

[0052] In a possible implementation, during the training of the audio understanding model, parameters of the second head remain fixed.

[0053] In a possible implementation, the audio understanding model performs first-stage training and second-stage training based on the audio training set;

[0054] In the first stage of training, the parameters of the encoding module remain fixed; in the second stage of training, the parameters of the encoding module are adjusted.

[0055] In a possible implementation, the apparatus further includes:

[0056] a segmentation module, configured to segment the audio sample into a plurality of audio segments according to a preset duration during the second stage of training, and sequentially input the plurality of audio segments into the encoding module in the audio understanding model;

[0057] A restriction module is used to restrict the audio understanding model through an attention mask to only obtain information of the current audio segment and historical audio segments before the current audio segment when processing the current audio segment.

[0058] In a possible implementation, during the training of the audio understanding model, the first head performs full parameter training, and the parameters of the large language model body are updated by fine-tuning.

[0059] In one possible implementation, the speech recognition model uses a recurrent neural network transcriber.

[0060] According to another aspect of the present disclosure, there is provided an audio understanding apparatus, comprising:

[0061] An acquisition module, configured to acquire an audio understanding model trained using the audio understanding model training method;

[0062] The first output module is used to input the audio to be processed into the audio understanding model, and output the text prediction result corresponding to the audio to be processed through the audio understanding model.

[0063] In a possible implementation, the apparatus further includes:

[0064] A deletion module is used to delete the second head in the audio understanding model.

[0065] In a possible implementation, the audio to be processed is streaming audio or non-streaming audio.

[0066] According to another aspect of the present disclosure, a training device for an audio understanding model is provided, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.

[0067] According to another aspect of the present disclosure, an audio understanding device is provided, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.

[0068] According to another aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.

[0069] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program implements the steps of the above method when executed by a processor.

[0070] In an embodiment of the present disclosure, a pre-trained speech recognition model and a pre-trained large language model are obtained, wherein the speech recognition model includes an encoding module, a prediction module and a first fusion module, the first fusion module is connected to the encoding module and the prediction module respectively, and the large language model includes a large language model body and a large language model head that are connected to each other. An audio understanding model is constructed according to the speech recognition model and the large language model, wherein the audio understanding model includes the encoding module, the large language model body and a second fusion module, the second fusion module is connected to the encoding module and the large language model body respectively. For any audio sample in the audio training set, the audio is extracted through the encoding module in the audio understanding model. The acoustic feature sequence corresponding to the frequency sample is obtained, and the semantic feature sequence corresponding to the audio sample is extracted through the large language model body in the audio understanding model. The second fusion module determines all valid alignment paths that can generate the target text label sequence corresponding to the audio sample based on the acoustic feature sequence and the semantic feature sequence, and calculates the total probability of all valid alignment paths. According to the total probability, the parameters of the audio understanding model are updated. By utilizing the powerful language understanding and generation capabilities of the pre-trained large language model and its ability to quickly learn through a small amount of audio data, combined with the audio and text alignment mechanism, efficient streaming speech recognition can be achieved, and the accuracy of speech recognition can be significantly improved even when there is less audio training data.

[0071] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.

[0073] Figure 1 A flowchart of a method for training an audio understanding model provided by an embodiment of the present disclosure is shown.

[0074] Figure 2 Schematic diagram showing a recurrent neural network transcriber.

[0075] Figure 3 A schematic diagram of an audio understanding model provided by an embodiment of the present disclosure is shown.

[0076] Figure 4 A schematic diagram illustrating the dynamic alignment process of the audio understanding model provided by an embodiment of the present disclosure during the decoding phase.

[0077] Figure 5A block diagram of a training device for an audio understanding model provided by an embodiment of the present disclosure is shown.

[0078] Figure 6 It is a block diagram of an audio understanding model training device or an audio understanding device 1900 according to an exemplary embodiment. DETAILED DESCRIPTION

[0079] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0080] As used herein, the terms "comprises," "comprising," "having," or variations thereof are open ended and include one or more stated features, integers, elements, steps, parts, or functions, but do not preclude the presence or addition of one or more other features, integers, elements, steps, parts, functions, or groups thereof.

[0081] When an element is referred to as being "connected," "coupled," "responsive" or variations thereof to another element, it can be directly connected, coupled or responsive to the other element or intervening elements may be present.

[0082] Although the terms first, second, third, etc. may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another element / operation. Therefore, without departing from the teachings of the present invention, the first element / operation in some embodiments may be referred to as the second element / operation in other embodiments.

[0083] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0084] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0085] The embodiment of the present disclosure provides a training method for an audio understanding model, by obtaining a pre-trained speech recognition model and a pre-trained large language model, wherein the speech recognition model includes an encoding module, a prediction module and a first fusion module, the first fusion module is connected to the encoding module and the prediction module respectively, the large language model includes a large language model body and a large language model head that are connected to each other, and an audio understanding model is constructed according to the speech recognition model and the large language model, wherein the audio understanding model includes the encoding module, the large language model body and a second fusion module, the second fusion module is connected to the encoding module and the large language model body respectively, for any audio sample in the audio training set, through the encoding in the audio understanding model The module extracts the acoustic feature sequence corresponding to the audio sample, and extracts the semantic feature sequence corresponding to the audio sample through the large language model body in the audio understanding model. The second fusion module determines all valid alignment paths that can generate the target text label sequence corresponding to the audio sample based on the acoustic feature sequence and the semantic feature sequence, and calculates the total probability of all the valid alignment paths. According to the total probability, the parameters of the audio understanding model are updated. In this way, the powerful language understanding and generation capabilities of the pre-trained large language model and its ability to quickly learn through a small amount of audio data are utilized, combined with the audio and text alignment mechanism, efficient streaming speech recognition can be achieved, and the accuracy of speech recognition can be significantly improved even when there is less audio training data.

[0086] The following describes in detail the training method of the audio understanding model provided by the embodiments of the present disclosure in conjunction with the accompanying drawings.

[0087] Figure 1 A flow chart of the audio understanding model training method provided by an embodiment of the present disclosure is shown. In one possible implementation, the large language model body that executes the audio understanding model training method may be a training device for the audio understanding model. For example, the audio understanding model training method may be executed by a terminal device or a server or other electronic device. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device or a wearable device, etc. In some possible implementations, the audio understanding model training method may be implemented by a processor calling computer-readable instructions stored in a memory. As Figure 1 As shown, the training method of the audio understanding model includes steps S11 to S15.

[0088] In step S11, a pre-trained speech recognition model and a pre-trained large language model are obtained; wherein the speech recognition model includes an encoding module, a prediction module and a first fusion module, the first fusion module is connected to the encoding module and the prediction module respectively, and the large language model includes a large language model body and a large language model head that are connected to each other.

[0089] In step S12, an audio understanding model is constructed based on the speech recognition model and the large language model, wherein the audio understanding model includes the encoding module, the large language model body and a second fusion module, and the second fusion module is connected to the encoding module and the large language model body respectively.

[0090] In step S13, for any audio sample in the audio training set, the acoustic feature sequence corresponding to the audio sample is extracted through the encoding module in the audio understanding model, and the semantic feature sequence corresponding to the audio sample is extracted through the large language model body in the audio understanding model.

[0091] In step S14, the second fusion module determines all valid alignment paths that can generate the target text label sequence corresponding to the audio sample based on the acoustic feature sequence and the semantic feature sequence, and calculates the total probability of all the valid alignment paths.

[0092] In step S15, the parameters of the audio understanding model are updated according to the total probability.

[0093] In the embodiments of the present disclosure, a pre-trained speech recognition model may refer to a speech recognition model that has been pre-trained on a speech dataset. By learning and extracting features from speech samples, the speech recognition model can capture the acoustic characteristics, prosodic patterns, and correspondence between speech and text, and convert the input speech into corresponding text content.

[0094] In an embodiment of the present disclosure, the pre-trained speech recognition model may include an encoding module (encoder), a prediction module (predictor) and a first fusion module (jointer). In some application scenarios, the encoding module may also be referred to as an encoder, the prediction module may also be referred to as a predictor, a prediction network, etc., and the first fusion module may also be referred to as a joint network, a joint module, etc., which are not limited here. In one possible implementation, the encoding module may be used to extract acoustic features, the prediction module may be used to extract semantic features, and the first fusion module may be used to fuse acoustic features and semantic features.

[0095] In one possible implementation, the speech recognition model uses a recurrent neural network transducer (RNNT).

[0096] Figure 2 Figure 2 shows a schematic diagram of a recurrent neural network transcriber. Figure 2 As shown, the recurrent neural network transcriber may include an encoding module, a prediction module, a first fusion module, and an activation layer.

[0097] The encoding module can be used to model acoustic features. For example, the encoding module can convert the input audio frame sequence into a high-dimensional acoustic feature representation. For example, the encoding module can input the audio frame x at the current time step t t , output the acoustic features corresponding to the current time step t The time step can refer to the smallest discrete time unit of audio processing, corresponding to a single frame of data processed or output by the model at a certain moment. For an audio frame sequence x=(x1,x2,...,x T ), the encoding module can output acoustic feature sequences The encoding module can be implemented using a Conformer or LSTM (Long Short-Term Memory) network, which is not limited here.

[0098] The prediction module can be used to model semantic features. During the training phase, the prediction module can input historical text labels before the current text step u (such as y u-1 or y <u =(y1,y2,...,y u-1 )), output the semantic features of the current text step u During the prediction phase, the prediction module takes as input the generated (i.e., predicted) text labels and outputs the semantic features of the current text step. A text step refers to a discrete unit in the text generation process, corresponding to a single token processed or output by the model. The prediction module can employ either an LSTM or Transformer architecture, without limitation here.

[0099] The training samples in the training set can be composed of audio-text pairs {x, y}, where x can represent the audio sample and y can represent the target text label sequence (i.e., the annotated text) corresponding to the audio sample. For example, if the target text label sequence y is "How is the weather today?", the text format input to the prediction module during training can be: <sos>How is the weather today? <sos>The Start Of Sentence is used to initialize the hidden state of the prediction module, indicating the beginning of text generation. The target text label sequence used to calculate the loss can be in the format of: How is the weather today? <eos>.in, <eos>It is the end of sentence symbol (End Of Sentence). <eos>As part of the training objective, this helps the model learn when to terminate its output.

[0100] The first fusion module can fuse the outputs of the encoding module and the prediction module. For example, the output of the first fusion module can be: The first fusion module may include a linear layer and an activation layer (eg, softmax).

[0101] like Figure 2 As shown in FIG, after the first fusion module, an activation layer (such as softmax) can be connected to generate the probability distribution P(y|t,u) of all possible word units (such as characters, words, symbols).

[0102] In this implementation, the pre-trained speech recognition model, through its recurrent neural network transcriber architecture, effectively achieves real-time dynamic alignment of audio and text sequences, demonstrating significant advantages in streaming speech recognition tasks. This architecture, through an encoding module to extract acoustic features, a prediction module to model text context dependencies, and a first fusion module to fuse multimodal information, offers powerful feature representation capabilities, significantly improving the model's recognition accuracy and real-time performance in complex scenarios.

[0103] In the embodiments of the present disclosure, a pre-trained large language model (LLM) may refer to a model that is pre-trained on large-scale text data. In some application scenarios, a large language model may also be referred to as a large language model, a large text model, etc., which is not limited here. In one possible implementation, the pre-trained large language model may be an open source pre-trained large language model, such as Llama, Qwen, etc. These open source models have been learned and trained on massive amounts of text, and have accumulated rich language knowledge and semantic understanding capabilities. In another possible implementation, the pre-trained large language model may be a model trained by the developer himself.

[0104] Large language models, pre-trained on massive amounts of text data, demonstrate powerful language understanding and generation capabilities. Compared to traditional models that rely on large amounts of data, large language models offer advantages in data-scarce languages ​​and complex speech tasks. Their large number of parameters also makes them more robust and adaptable in noisy environments. Furthermore, large language models pre-trained in multiple languages ​​can reduce the need for language-specific data and ease model management complexity.

[0105] In an embodiment of the present disclosure, the large language model may include a large language model body (LLM-body) and a large language model head (LLM-head) that are interconnected. The large language model body may be composed of a multi-layer neural network (such as a Transformer module), which may be responsible for deep feature extraction and semantic understanding of the input text, and obtain general language representation capabilities through pre-training of massive text data; the large language model head, as the output layer of the large language model, may be composed of a linear layer and an activation layer, and may be responsible for mapping the high-dimensional features extracted by the large language model body to the target output space (such as the probability distribution of the word table). In some application scenarios, the large language model head may also be called the output layer of the large language model.

[0106] In an embodiment of the present disclosure, an audio understanding model can be constructed based on a pre-trained speech recognition model and a pre-trained large language model. The audio understanding model includes an encoding module in the pre-trained speech recognition model, a large language model body in the pre-trained large language model, and a second fusion module. In the audio understanding model, the second fusion module is connected to the encoding module and the large language model body, respectively, and can be used to fuse the output of the encoding module with the output of the large language model body.

[0107] In one possible implementation, an audio understanding model can be built on top of a pre-trained speech recognition model. Specifically, the encoding module in the pre-trained speech recognition model can be retained, the prediction module in the pre-trained speech recognition model can be replaced with a large language model body, and the first fusion module in the pre-trained speech recognition model can be reconstructed into a second fusion module.

[0108] Among them, the encoding module in the audio understanding model can be used to process audio input and output acoustic features h enc The large language model body in the audio understanding model can input text and output semantic features h LLM In the audio understanding model, the text context modeling capability can be enhanced by replacing the prediction module in the speech recognition model with a large language model body. The input of the second fusion module in the audio understanding model can include the acoustic features h output by the encoding module. enc and the semantic features h output by the large language model body LLM The second fusion module can dynamically align acoustic features and semantic features.

[0109] In a possible implementation, the second fusion module includes the first fusion module and a first header, wherein the first header is initialized according to the large language model header, and the first header adds a dimension corresponding to the blank symbol blank on the basis of the large language model header.

[0110] In this implementation, the second fusion module may be composed of the first fusion module and the first head.

[0111] Among them, the first fusion module retains the structure of the fusion module in the pre-trained speech recognition model (such as linear layer + activation layer), which can be used to preliminarily fuse acoustic features and semantic features.

[0112] The first header can be initialized based on the large language model header. That is, the parameters of the first header can be inherited from the pre-trained large language model header to utilize the existing language modeling capabilities of the pre-trained large language model. In addition, in this implementation, the dimension corresponding to blank (blank symbol) is added to the large language model header to obtain the first header. This is to adapt to the streaming alignment requirements of speech recognition models (such as recurrent neural network transcribers), because the speech recognition model needs to process blank (indicating no output or waiting for subsequent input) during decoding. For example, if the output dimension of the large language model header is the word table size V, the output dimension of the modified first header is V+1, and the newly added dimension is dedicated to the probability calculation of blank.

[0113] This implementation retains the language generation capabilities of large language models by reusing their header parameters, while also supporting the alignment mechanisms unique to speech recognition models (such as recurrent neural network transcribers) by expanding the dimensionality. The introduction of blanks enables the audio understanding model to dynamically align audio and text (e.g., a path moving right indicates a blank, while an upward path indicates a generated character), thus enabling streaming decoding.

[0114] In a possible implementation, the method further includes: during the pre-training process, controlling the output dimension of the first fusion module to be the same as the input dimension of the large language model head.

[0115] In this implementation, during the pre-training process, the output dimension of the first fusion module is controlled to be the same as the input dimension of the large language model head, thereby ensuring that the two modules can be seamlessly connected in structure, thereby achieving effective feature fusion. The first fusion module is responsible for jointly processing the acoustic features extracted by the encoding module and the semantic features output by the prediction module (or the replaced large language model body). Its output needs to match the input dimension of the large language model head to avoid information loss or calculation errors caused by dimensional mismatch. This alignment operation enables the second fusion module to directly use the parameters of the pre-trained large language model head without additional adjustment, thereby retaining the original language modeling capabilities of the large language model while being compatible with the feature input of the audio modality.

[0116] This implementation enables multimodal features (acoustic and semantic) to be correctly processed by the large language model head after fusion. For example, if the input dimension of the large language model head is D, the output of the first fusion module must also be adjusted to D to ensure that the joint features of its output can be directly input into the large language model head for probability distribution calculation. This consistency not only simplifies the structure of the audio understanding model, but also improves the training efficiency, allowing the audio understanding model to converge quickly during the fine-tuning stage. In addition, this dimensionality control also lays the foundation for subsequent streaming processing (such as adding blank dimensions), ensuring that the audio understanding model maintains stability and robustness when dynamically aligning audio and text.

[0117] In the disclosed embodiment, during the training of the audio understanding model, for each audio sample in the audio training set, the encoding module processes the raw audio signal frame by frame, outputting a sequence of acoustic features. The large language model then autoregressively generates a corresponding sequence of semantic features based on historical text labels. These acoustic and semantic feature sequences provide multimodal input for the subsequent second fusion module.

[0118] In an embodiment of the present disclosure, a second fusion module can determine all valid alignment paths that can generate a target text label sequence corresponding to the audio sample based on the acoustic feature sequence and the semantic feature sequence, and calculate the total probability of all valid alignment paths.

[0119] The alignment path can refer to the mapping relationship between the acoustic feature sequence (time step t) and the semantic feature sequence (text step u). For example:

[0120] Horizontal shift (→): time step t+1, text step u remains unchanged (corresponding to output blank, indicating that no text has been generated yet).

[0121] Vertical movement (↑): time step t remains unchanged, text step u+1 (corresponding to generating an actual text label).

[0122] A valid alignment path refers to all legal paths from the starting point (t = 0, u = 0) to the end point (t = T, u = U), and the generated text sequence is completely consistent with the target text label sequence. Where T is the total time step of the audio sample and U is the length of the target text label sequence.

[0123] The second fusion module can receive the acoustic feature sequence h from the encoding module t and a sequence of semantic features s from a large language model corpus u , output the joint probability distribution p(y through the linear layer and activation layer (such as softmax) t,u |h t ,s u ). Among them, y t,u It can represent the predicted label of the current step.

[0124] The probability of each aligned path is the product of the probabilities of all steps on the path. For example, a path is Path k :(t1,u1)→(t2,u2)→...→(t N ,u N ), then its probability is Sum the probabilities of all valid paths:

[0125] For example, if the target text is "AB" and the audio has 3 time steps, the two possible valid paths are:

[0126] Path 1: (t=1)→(t=2)→(t=3↑)→(t=3↑)

[0127] (Output: blank→blank→A→B)

[0128] Path 2: (t=1↑)→(t=2)→(t=3↑)

[0129] (Output: A→blank→B)

[0130] Total probability P total =P(Path1)+P(Path2).

[0131] In one possible implementation, updating the parameters of the audio understanding model based on the total probability includes: taking the negative logarithm of the total probability to obtain the value of a first loss function corresponding to the audio understanding model; and updating the parameters of the audio understanding model based on the value of the first loss function.

[0132] In this implementation, during the training of the audio understanding model, after calculating the total probability of all valid alignment paths, the total probability can be maximized (i.e., the negative log probability can be minimized) by optimizing the model parameters. total Take the negative logarithm to get the value of the first loss function L RNNT =-log(P total The probability range is [0, 1]. After taking the negative logarithm, the higher the probability (closer to 1), the lower the loss (closer to 0); conversely, the lower the probability (closer to 0), the higher the loss (approaching infinity). This transformation transforms the probability maximization problem into a loss minimization problem, which conforms to the gradient descent optimization framework.

[0133] For example, if an audio sample has two valid alignment paths with probabilities of 0.6 and 0.4 respectively, the total probability P total =0.6+0.4=1, loss value L RNNT =-log(1)=0; If the path probability distribution is not ideal (such as 0.1 and 0.1), the loss value L RNNT =-log(0.2)≈1.61, the audio understanding model will adjust parameters through gradient descent to increase the possibility of generating high-probability paths.

[0134] In this implementation, a gradient can be calculated based on the value of the first loss function, and backpropagation can be used to update the parameters of the audio understanding model. For example, the parameters of at least some modules in the second fusion module, the large language model body, and the encoding module can be updated.

[0135] In this implementation, the first loss function directly optimizes the alignment path between acoustic and semantic features, ensuring that the text sequences generated by the audio understanding model dynamically match the audio input. By minimizing the first loss function, the second fusion module learns to effectively combine the acoustic features of the encoding module with the semantic features of the large language model body.

[0136] In one possible implementation, the audio understanding model also includes a second head, which is connected to the large language model body and initialized according to the large language model head; the method also includes: outputting the text prediction result corresponding to the audio sample through the second head; determining the value of the second loss function corresponding to the audio understanding model based on the text prediction result corresponding to the audio sample and the target text label sequence; and updating the parameters of the audio understanding model based on the value of the second loss function.

[0137] In this implementation, the second head can directly reuse the original head structure of a pre-trained large language model, including, for example, linear and softmax layers, to inherit the powerful language generation capabilities of the large language model. The parameters of the second head can be copied from the large language model head, but it is independent of the first head, forming a dual-task output branch. Furthermore, the second head can be connected to the main body of the large language model to receive its output semantic features.

[0138] The second head can directly generate text prediction results (such as the probability distribution of characters or words) corresponding to the audio sample based on the semantic features output by the large language model body. By calculating a second loss function (such as cross-entropy loss) between the text prediction results and the true labels (i.e., the target text label sequence), it provides additional supervision signals and strengthens language modeling capabilities.

[0139] In a possible implementation, during the training of the audio understanding model, parameters of the second head remain fixed.

[0140] In this implementation, during the training of the audio understanding model, the parameters of the second head can be kept fixed (i.e., it does not participate in gradient updates). This prevents its text generation ability from being interfered with by the speech alignment task while introducing the pre-trained knowledge of the large language model. The second head directly reuses the parameters of the original head of the large language model and, by freezing its weights, ensures that the audio understanding model retains the powerful language modeling capabilities of the large language model. This design not only utilizes the prior knowledge of the large language model to improve the quality of text prediction, but also prevents parameter conflicts during multi-task training, allowing the audio understanding model to converge more stably.

[0141] Figure 3 Schematic diagram of the audio understanding model provided by the embodiment of the present disclosure is shown. Figure 3 As shown, the audio understanding model may include an encoding module, an LLM body, a second fusion module, and a second head. The encoding module may receive audio input, and the LLM body may receive text input. Based on the output of the second fusion module, the value of a first loss function (e.g., RNNT loss) may be calculated. Based on the output of the second head, the value of a second loss function (e.g., CE loss) may be calculated.

[0142] In one possible implementation, the audio understanding model performs first-stage training and second-stage training based on the audio training set; wherein, in the first-stage training, the parameters of the encoding module remain fixed; and in the second-stage training, the parameters of the encoding module are adjusted.

[0143] In this implementation, a two-stage training strategy can be used to optimize the audio understanding model.

[0144] During the first phase of training, the parameters of the encoding module can remain fixed to leverage the acoustic feature extraction capabilities already learned by the pre-trained encoding module, while parameters of other components of the audio understanding model (such as the main large language model and the second fusion module) are adjusted. The encoding module in the pre-trained speech recognition model already possesses the ability to extract effective acoustic features from audio signals. Acoustic features, such as spectral information and prosodic patterns, can provide a foundation for subsequent text generation. By keeping the encoding module parameters fixed during the first phase, excessive adjustments to the encoding module during the initial training phase are avoided, thereby preserving the general acoustic feature extraction capabilities learned during pre-training. This ensures that the audio understanding model can still effectively extract acoustic features when processing new audio data. In the first phase, training focuses on combining the main pre-trained large language model with the encoding module, aligning and fusing acoustic and semantic features through the second fusion module. At this stage, the parameters of the main large language model and the second fusion module are primarily updated to enable the audio understanding model to better understand and generate text.

[0145] In one example, the audio understanding model trained in the first stage can be called a non-streaming audio understanding model.

[0146] In the second phase of training, the encoding module parameters can be unfrozen to optimize its ability to extract local acoustic features. In other words, in the second phase, the audio understanding model can simultaneously update the parameters of the encoding module, the main large language model, and the second fusion module to achieve more precise alignment and fusion of audio and text.

[0147] In one example, the audio understanding model trained in the second stage can be called a streaming audio understanding model.

[0148] This implementation avoids gradient conflicts caused by optimizing all modules simultaneously through phased training. Through progressive training, the encoding module transitions from global feature extraction to local real-time processing, making it suitable for streaming audio understanding scenarios.

[0149] In one possible implementation, the method further includes: in the second stage training, dividing the audio sample into multiple audio segments according to a preset duration, and inputting the multiple audio segments into the encoding module in the audio understanding model in sequence; and limiting the audio understanding model through an attention mask so that when processing the current audio segment, it can only obtain information about the current audio segment and historical audio segments before the current audio segment.

[0150] In this implementation, in the second stage of training, in order to achieve streaming audio processing capabilities, the audio understanding model can adopt a block training strategy combined with attention mask control. Specifically, the complete audio sample can be divided into continuous audio segments according to a preset duration (such as 320ms / block), and input into the encoding module in sequence. For example, a 3-second audio (3000ms) can be divided into about 9 320ms audio segments (the insufficient part at the end is padded with zeros) and processed in sequence. In addition, in the self-attention layer of the audio understanding model, a one-way mask can be applied so that the audio understanding model can only access the information of the current audio segment and historical audio segments, while the information of future audio segments is completely blocked. In this way, the causality of the streaming scenario can be simulated, so that the audio understanding model only relies on the received audio data to output results in real time.

[0151] The encoding module can learn to extract effective features from localized audio samples (i.e., audio segments), rather than relying on the full context. Through segmented training, the encoding module can gradually optimize its ability to capture short-term acoustic patterns (such as phonemes and syllables).

[0152] In a possible implementation, during the training of the audio understanding model, the first head performs full parameter training, and the parameters of the large language model body are updated by fine-tuning.

[0153] In this implementation, the first head in the audio understanding model is directly responsible for text generation (such as generating word-unit probability distributions) and must quickly adapt to the special requirements of speech recognition tasks (such as increasing the blank dimension). By fully training the first head—that is, completely updating its weights via gradient descent—the output layer of the audio understanding model can flexibly learn the language patterns after audio-text alignment.

[0154] Large language models have been pre-trained on massive amounts of text to develop general semantic understanding capabilities. Direct full-parameter training can easily lead to overfitting (especially when audio data is limited). In this implementation, efficient parameter fine-tuning techniques (such as LoRA and Adapter) can be used to train only a small number of newly added low-rank matrices or adaptation layers, freezing the original parameters of the large language model. This allows for subtle adjustments to feature representations to adapt to the audio context while retaining pre-trained knowledge.

[0155] In this implementation, full parameter training of the first head allows for precise optimization of the output distribution and improved recognition accuracy. Fine-tuning the parameters of the large language model body avoids damaging the core capabilities of the pre-trained language model and reduces training risk. Furthermore, compared to full model training, this significantly reduces computational overhead (especially when the large language model body has a large number of parameters).

[0156] The embodiment of the present disclosure also provides an audio understanding method, including: obtaining an audio understanding model trained by the audio understanding model training method; inputting the audio to be processed into the audio understanding model, and outputting a text prediction result corresponding to the audio to be processed through the audio understanding model.

[0157] In an embodiment of the present disclosure, an optimized audio understanding model can be obtained by the training method of the audio understanding model described above. The audio understanding model integrates the acoustic feature extraction capability of the encoding module in the pre-trained speech recognition model and the semantic understanding capability of the large language model. In actual application, the audio to be processed (such as a real-time voice stream or a complete recording) can be input into the audio understanding model. The audio understanding model extracts acoustic features through the encoding module and combines the semantic reasoning of the large language model body. Finally, the second fusion module dynamically aligns the audio and text sequences and outputs the corresponding text prediction results. The audio understanding model provided by the embodiment of the present disclosure supports streaming processing, can generate verbatim text (such as subtitles) in real time, and can also process non-streaming audio. It is suitable for scenarios such as smart assistants and conference transcription.

[0158] In a possible implementation, before inputting the audio to be processed into the audio understanding model, the method further includes: deleting the second head in the audio understanding model.

[0159] In this implementation, the second head is used only during training to assist in optimizing the text generation task (e.g., through cross-entropy loss). Removing the second head reduces redundant computation and lowers inference latency, which is particularly important for real-time streaming processing (such as voice assistants), while ensuring that the audio understanding model retains only the necessary output paths.

[0160] In a possible implementation, the audio to be processed is streaming audio or non-streaming audio.

[0161] In this implementation, the audio understanding model can be used for reasoning about streaming audio or non-streaming audio, depending on the actual business scenario.

[0162] In one possible implementation, the audio understanding model can use beam search for decoding to find the optimal output path. Beam search is a heuristic search algorithm that maintains a set of candidate solutions at each step. These candidate solutions are ranked according to a scoring mechanism, and the top-scoring solutions are typically selected as candidates for the next step.

[0163] Figure 4 A schematic diagram illustrating the dynamic alignment process of the audio understanding model provided by an embodiment of the present disclosure during the decoding phase. Figure 4 It can demonstrate the streaming alignment mechanism between the acoustic feature sequence (time step t) and the semantic feature sequence (text step u).

[0164] exist Figure 4 In the example, the horizontal axis (t) can represent the time step of the audio input (such as the audio frame sequence), and the acoustic feature A output by the encoding module is t (black circle). The vertical axis (u) can represent the number of steps of text generation (such as words or characters that have been output), and the semantic features T output by the large language model body u (white circle). The node y(t,u) can represent the joint decoding result at time step t and text step u, which may be the actual text label (such as Chinese characters, words) or blank. Among them, blank means that no text is generated in the current time step and only the time step is advanced.

[0165] Figure 4 The red arrow in the figure is an example of a decoding path. A horizontal shift (→) indicates time step t+1, and the text step u) remains unchanged (output blank, as shown in ). Vertical movement (↑) means that the time step t remains unchanged and the text step u+1 (the actual text is output, such as y(3,2)).

[0166] In this way, the audio understanding model can align acoustic features with semantic features in real time when processing streaming audio input and generate corresponding text results. This alignment capability not only improves the interactive experience but also significantly enhances the performance of the streaming speech recognition system.

[0167] The audio understanding model training method and audio understanding method provided in the embodiments of the present disclosure can be applied to technical fields such as artificial intelligence, speech recognition, large audio understanding models, and streaming speech recognition, without limitation herein. Furthermore, the audio understanding model trained using the audio understanding model training method provided in the embodiments of the present disclosure can be applied to scenarios such as real-time subtitles, voice search, and smart homes, without limitation herein.

[0168] It is understood that the above-mentioned various method embodiments mentioned in this disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to space limitations, this disclosure will not go into details. It is understood by those skilled in the art that in the above-mentioned methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.

[0169] In addition, the present disclosure also provides a training device for an audio understanding model, a non-volatile computer-readable storage medium for an audio understanding device, and a computer program product. The above can all be used to implement the training of any audio understanding model or audio understanding method provided by the present disclosure. The corresponding technical solutions and technical effects can be found in the corresponding records in the method section and will not be repeated here.

[0170] Figure 5 FIG. 1 is a block diagram of a training device for an audio understanding model provided by an embodiment of the present disclosure. Figure 5 As shown, the training device of the audio understanding model includes:

[0171] An acquisition module 51 is configured to obtain a pre-trained speech recognition model and a pre-trained large language model; wherein the speech recognition model includes an encoding module, a prediction module, and a first fusion module, the first fusion module being connected to the encoding module and the prediction module, respectively; and the large language model includes a large language model body and a large language model head connected to each other.

[0172] a construction module 52 for constructing an audio understanding model based on the speech recognition model and the large language model, wherein the audio understanding model includes the encoding module, the large language model body, and a second fusion module, wherein the second fusion module is connected to the encoding module and the large language model body, respectively;

[0173] An extraction module 53 is configured to extract, for any audio sample in the audio training set, an acoustic feature sequence corresponding to the audio sample using the encoding module in the audio understanding model, and a semantic feature sequence corresponding to the audio sample using the large language model body in the audio understanding model;

[0174] a calculation module 54, configured to determine, through the second fusion module, based on the acoustic feature sequence and the semantic feature sequence, all valid alignment paths that can generate a target text label sequence corresponding to the audio sample, and calculate a total probability of all the valid alignment paths;

[0175] The first updating module 55 is configured to update the parameters of the audio understanding model according to the total probability.

[0176] In a possible implementation, the second fusion module includes the first fusion module and a first header, wherein the first header is initialized according to the large language model header, and the first header adds a dimension corresponding to the blank symbol blank on the basis of the large language model header.

[0177] In a possible implementation, the apparatus further includes:

[0178] A control module is used to control the output dimension of the first fusion module to be the same as the input dimension of the head of the large language model during pre-training.

[0179] In a possible implementation, the first updating module 55 is configured to:

[0180] Taking the negative logarithm of the total probability to obtain a value of a first loss function corresponding to the audio understanding model;

[0181] Update the parameters of the audio understanding model according to the value of the first loss function.

[0182] In one possible implementation, the audio understanding model further includes a second head, the second head is connected to the large language model body, and the second head is initialized according to the large language model head;

[0183] The device further comprises:

[0184] A second output module, configured to output a text prediction result corresponding to the audio sample through the second header;

[0185] A determination module, configured to determine a value of a second loss function corresponding to the audio understanding model based on a text prediction result corresponding to the audio sample and the target text label sequence;

[0186] A second updating module is used to update the parameters of the audio understanding model according to the value of the second loss function.

[0187] In a possible implementation, during the training of the audio understanding model, parameters of the second head remain fixed.

[0188] In a possible implementation, the audio understanding model performs first-stage training and second-stage training based on the audio training set;

[0189] In the first stage of training, the parameters of the encoding module remain fixed; in the second stage of training, the parameters of the encoding module are adjusted.

[0190] In a possible implementation, the apparatus further includes:

[0191] a segmentation module, configured to segment the audio sample into a plurality of audio segments according to a preset duration during the second stage of training, and sequentially input the plurality of audio segments into the encoding module in the audio understanding model;

[0192] A restriction module is used to restrict the audio understanding model through an attention mask to only obtain information of the current audio segment and historical audio segments before the current audio segment when processing the current audio segment.

[0193] In a possible implementation, during the training of the audio understanding model, the first head performs full parameter training, and the parameters of the large language model body are updated by fine-tuning.

[0194] In one possible implementation, the speech recognition model uses a recurrent neural network transcriber.

[0195] According to another aspect of the present disclosure, there is provided an audio understanding apparatus, comprising:

[0196] An acquisition module, configured to acquire an audio understanding model trained using the audio understanding model training method;

[0197] The first output module is used to input the audio to be processed into the audio understanding model, and output the text prediction result corresponding to the audio to be processed through the audio understanding model.

[0198] In a possible implementation, the apparatus further includes:

[0199] A deletion module is used to delete the second head in the audio understanding model.

[0200] In a possible implementation, the audio to be processed is streaming audio or non-streaming audio.

[0201] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. Its specific implementation and technical effects can refer to the description of the above method embodiments. For the sake of brevity, they will not be repeated here.

[0202] According to another aspect of the present disclosure, a training device for an audio understanding model is provided, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.

[0203] According to another aspect of the present disclosure, an audio understanding device is provided, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.

[0204] An embodiment of the present disclosure further provides a non-volatile computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the above method when executed by a processor.

[0205] An embodiment of the present disclosure further provides a computer program product, including a computer program, or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program implements the steps of the above method when executed by a processor.

[0206] Figure 6 FIG1 is a block diagram of an audio understanding model training device or an audio understanding device 1900 according to an exemplary embodiment. For example, the device 1900 can be provided as a server or a terminal device. Figure 6 The apparatus 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions, such as an application, that can be executed by the processing component 1922. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described method.

[0207] The device 1900 may also include a power supply component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to a network, and an input / output interface 1958 (I / O interface). The device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server 2003. TM , MacOS X TM , Unix TM ,Linux TM , FreeBSD TM or similar.

[0208] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by the processing component 1922 of the apparatus 1900 to perform the above-described method.

[0209] A computer-readable storage medium can be a tangible device that can hold and store programs / instructions used by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.

[0210] The computer programs (or computer-readable program instructions) described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0211] The computer program (or computer program instructions) for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, by utilizing state information of computer-readable program instructions to personalize and customize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.

[0212] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0213] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0214] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0215] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0216] The computer program product may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).

[0217] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.

[0218] If the technical solutions of the embodiments of this disclosure involve personal information, the products that apply the technical solutions of the embodiments of this disclosure have clearly informed the individual of the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solutions of the embodiments of this disclosure involve sensitive personal information, the products that apply the technical solutions of the embodiments of this disclosure have obtained the individual's separate consent before processing the sensitive personal information and simultaneously meet the "explicit consent" requirement. For example, on personal information collection devices such as cameras, a clear and prominent sign is set to inform the individual that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that they agree to the collection of their personal information; or on the personal information processing device, while using obvious signs / information to inform the individual of the personal information processing rules, the individual's authorization is obtained through pop-up messages or by asking the individual to upload their personal information. The personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

[0219] While various embodiments of the present disclosure have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.< / eos> < / eos> < / eos> < / sos> < / sos>

Claims

1. A method for training an audio comprehension model, characterized in that: include: Obtaining a pre-trained speech recognition model and a pre-trained large language model; wherein the speech recognition model includes an encoding module, a prediction module, and a first fusion module, the first fusion module being connected to the encoding module and the prediction module respectively; and the large language model includes a large language model body and a large language model head that are connected to each other; Constructing an audio understanding model based on the speech recognition model and the large language model, wherein the audio understanding model includes the encoding module, the large language model body, and a second fusion module, and the second fusion module is connected to the encoding module and the large language model body respectively; For any audio sample in the audio training set, extract the acoustic feature sequence corresponding to the audio sample through the encoding module in the audio understanding model, and extract the semantic feature sequence corresponding to the audio sample through the large language model body in the audio understanding model; Determining, by the second fusion module, all valid alignment paths that can generate a target text label sequence corresponding to the audio sample based on the acoustic feature sequence and the semantic feature sequence, and calculating a total probability of all valid alignment paths; According to the total probability, the parameters of the audio understanding model are updated.

2. The method according to claim 1, characterized in that The second fusion module includes the first fusion module and a first header, wherein the first header is initialized according to the large language model header, and the first header adds a dimension corresponding to the blank symbol blank on the basis of the large language model header.

3. The method according to claim 2, characterized in that The method further comprises: During the pre-training process, the output dimension of the first fusion module is controlled to be the same as the input dimension of the large language model head.

4. The method according to any one of claims 1 to 3, characterized in that The updating of the parameters of the audio understanding model according to the total probability includes: Taking the negative logarithm of the total probability to obtain a value of a first loss function corresponding to the audio understanding model; Update the parameters of the audio understanding model according to the value of the first loss function.

5. The method according to any one of claims 1 to 3, characterized in that The audio understanding model further includes a second head, the second head being connected to the large language model body and the second head being initialized based on the large language model head; The method further comprises: Outputting a text prediction result corresponding to the audio sample through the second head; Determining a value of a second loss function corresponding to the audio understanding model according to a text prediction result corresponding to the audio sample and the target text label sequence; Update the parameters of the audio understanding model according to the value of the second loss function.

6. The method according to claim 5, characterized in that During the training of the audio understanding model, the parameters of the second head remain fixed.

7. The method according to any one of claims 1 to 3, characterized in that The audio understanding model performs first-stage training and second-stage training based on the audio training set; In the first stage of training, the parameters of the encoding module remain fixed; in the second stage of training, the parameters of the encoding module are adjusted.

8. The method according to claim 7, characterized in that The method further comprises: In the second stage of training, the audio sample is divided into multiple audio segments according to a preset duration, and the multiple audio segments are sequentially input into the encoding module of the audio understanding model; The attention mask is used to restrict the audio understanding model from only obtaining information about the current audio segment and historical audio segments before the current audio segment when processing the current audio segment.

9. The method according to claim 2 or 3, characterized in that During the training process of the audio understanding model, the first head performs full parameter training, and the parameters of the large language model body are updated by fine-tuning.

10. The method according to any one of claims 1 to 3, characterized in that The speech recognition model uses a recurrent neural network transcriber.

11. An audio understanding method, characterized in that: include: Obtaining an audio understanding model trained by the audio understanding model training method according to any one of claims 1 to 10; The audio to be processed is input into the audio understanding model, and the audio understanding model outputs the text prediction result corresponding to the audio to be processed.

12. The method according to claim 11, characterized in that Before inputting the audio to be processed into the audio understanding model, the method further includes: The second head in the audio understanding model is deleted.

13. The method according to claim 11 or 12, characterized in that The audio to be processed is streaming audio or non-streaming audio.

14. A training device for an audio comprehension model, characterized in that: include: An acquisition module, configured to obtain a pre-trained speech recognition model and a pre-trained large language model; wherein the speech recognition model includes an encoding module, a prediction module, and a first fusion module, the first fusion module being connected to the encoding module and the prediction module, respectively; and the large language model includes a large language model body and a large language model head connected to each other; A construction module, configured to construct an audio understanding model based on the speech recognition model and the large language model, wherein the audio understanding model includes the encoding module, the large language model body, and a second fusion module, wherein the second fusion module is connected to the encoding module and the large language model body respectively; An extraction module is configured to extract, for any audio sample in the audio training set, an acoustic feature sequence corresponding to the audio sample using the encoding module in the audio understanding model, and to extract a semantic feature sequence corresponding to the audio sample using the large language model body in the audio understanding model; a calculation module, configured to determine, through the second fusion module, based on the acoustic feature sequence and the semantic feature sequence, all valid alignment paths that can generate a target text label sequence corresponding to the audio sample, and calculate a total probability of all the valid alignment paths; A first updating module is used to update the parameters of the audio understanding model according to the total probability.

15. An audio understanding device, characterized in that: include: an acquisition module, configured to acquire an audio understanding model trained by the audio understanding model training method according to claim 14; The first output module is used to input the audio to be processed into the audio understanding model, and output the text prediction result corresponding to the audio to be processed through the audio understanding model.

16. A training device for an audio comprehension model, comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 10.

17. An audio understanding device comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 11 to 13.

18. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.

19. A computer program product comprising a computer program, or a non-volatile computer-readable storage medium carrying a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.

Citation Information

Patent Citations

  • End-to-end far-field speech recognition system training method and device, and computer equipment

    CN115527526A

  • Speech recognition model training method and speech recognition method

    CN119107940A

  • Method for training speech recognition model, method and system for speech recognition

    US11580957B1

  • Generation of optimized spoken language understanding model through joint training with integrated knowledge-language module

    US20220230628A1

  • Methods and systems for streamable multimodal language understanding

    US20230223018A1

Cited By

  • Joint optimization method and system for end-to-end streaming speech recognition and natural language understanding

    CN121214927A

  • Model training method and device, electronic equipment, storage medium and program product

    CN121354542A

  • Training method and device of voice large model, equipment and medium

    CN121438813A