Speech Large Model Modal Alignment Method and Device Based on Two-Stage Decoupling Method

Through the modal alignment method of the speech large model with two-stage decoupling method, the problems of information loss and performance degradation in the existing technology are solved, and more efficient speech information decoupling and re-extraction are achieved, improving the performance of downstream tasks.

CN119670718BActive Publication Date: 2025-06-10HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN) +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510185747.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-10
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

Existing speech models have insufficient decoupling problems in modal alignment, resulting in information loss and performance degradation.

Method used

The speech large model modal alignment method adopts a two-stage decoupling method, and the first-stage decoupling is performed through two different encoders and length reduction methods, and the second-stage decoupling is performed using sequence-level splicing and deengaging modules during the information integration process.

Benefits of technology

It effectively reduces the loss of original information during information integration and improves the performance of comprehensive speech understanding and analysis tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119670718B_ABST
    Figure CN119670718B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for modal alignment of a speech large model based on a two-stage decoupling method, which relates to the technical field of natural language processing. The method includes: obtaining a pre-trained speech data set and a pre-trained task instruction text; constructing an initial speech large model, and pre-training the initial speech large model by using a two-stage decoupling method according to the pre-trained speech data set and the pre-trained task instruction text to obtain a pre-trained speech large model; performing instruction fine-tuning on the pre-trained speech large model by using the LoRA fine-tuning technology to obtain a trained speech large model; inputting the speech data to be processed and the instruction corresponding to the speech data into the trained speech large model for processing, and outputting a text that matches the instruction requirement corresponding to the speech data. By using the present invention, the problem of information loss caused by feature decoupling can be solved, and the performance of the speech large model in task analysis can be improved by using the present invention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly to a method and device for speech large model modality alignment based on a two-stage decoupling method. Background Art

[0002] With the rapid development of computer technology and artificial intelligence technology, speech processing technology has gradually transitioned from traditional signal processing methods to the mainstream stage dominated by methods based on machine learning and deep learning algorithms. Traditional speech processing methods include: methods based on sound signal processing, methods based on feature extraction, and methods based on deep learning models. In the multi-modal field, modality alignment is a highly concerned issue. In the multi-modal field, modality alignment methods include: methods based on statistical learning, explicit alignment methods based on deep learning, implicit alignment methods based on deep learning, and other alignment methods based on deep learning. The purpose of modality alignment of speech large models is to enable the large language model trained based on a large-scale text corpus at the backend to fully understand the content of the input speech and complete various speech-related downstream tasks in combination with the speech content. According to the differences in the selection of target downstream tasks, the specific technologies are mainly divided into the following two categories: the alignment method when the output modality only contains text, and the alignment method when the output modality contains text and speech.

[0003] Currently, the main problems to be solved in modality alignment on speech large models are two aspects: (1) The length of speech representation is much longer than the length of the text representation sequence. The longer sequence length will not only affect the accuracy of large model inference but also increase the inference latency of the model. Therefore, it is necessary to reduce the length of the input speech representation; (2) There are significant differences between speech representation and text representation in the representation space. Usually, the speech representation is mapped to the input space of the large model through a trainable linear mapping module, but a high-quality and sufficiently large training dataset is required as support. Currently, the method for modality alignment in speech large models is the SALMONN model jointly proposed by Tsinghua University and ByteDance. It uses two encoders for semantic features and acoustic features to encode speech, introduces speech-text modality alignment through the Q-Former module for image-text alignment, extracts the required features from the speech features through a preset trainable multiple Token sequence, and the length is consistent with the preset sequence length. Finally, good performance has been achieved in multiple downstream tasks through the above method.

[0004] Problems existing in existing speech large models include: insufficient decoupling. Generally, speech information can be divided into semantic information and acoustic information, and different downstream tasks have different degrees of attention to these two aspects of information. However, in existing technologies, only a single encoder is used, which only focuses on one of these aspects of information; while another part, like the SALMONN model, although uses two encoders to encode the features of the two aspects respectively, it is then coupled in an unclear way in the follow-up, resulting in the mixing of information. Summary of the Invention

[0005] In order to solve the technical problem in the existing technology that due to often only considering single-aspect information or having feature coupling resulting in information loss, and thus insufficient decoupling and re-extraction of speech information are not achieved, the embodiments of the present invention provide a speech large model modality alignment method and device based on a two-stage decoupling method. The technical solution is as follows:

[0006] On the one hand, a speech large model modality alignment method based on a two-stage decoupling method is provided. This method is implemented by a speech large model modality alignment device based on a two-stage decoupling method, and this method includes:

[0007] S1. Obtain a pre-trained speech data set and a pre-trained task instruction text;

[0008] S2. Use the Librosa method to clean the speech data set to obtain a cleaned speech data set;

[0009] S3. Construct an initial speech large model; the initial speech large model includes: a large language model and an alignment module;

[0010] S4. Input the cleaned speech data set and the pre-trained task instruction text into the initial speech large model. Through the alignment module, decouple the cleaned speech data set to obtain final speech features; use the tokenizer of the large language model to process the pre-trained task instruction text to obtain pre-trained text features; process the final speech features and the pre-trained text features through a sequence-level splicing method to obtain a final feature sequence; according to the final feature sequence, pre-train the initial speech large model to obtain a pre-trained speech large model;

[0011] S5. Obtain an instruction fine-tuning training set; according to the instruction fine-tuning training set, use the LoRA fine-tuning technology to perform instruction fine-tuning on the pre-trained speech large model to obtain a trained speech large model;

[0012] S6. Obtain the voice data to be processed and the instructions corresponding to the voice data; input the voice data to be processed and the instructions corresponding to the voice data into the trained large voice model for processing, and output the text that matches the instruction requirements corresponding to the voice data.

[0013] Optionally, the initial large voice model further includes: a voice encoder and an audio encoder;

[0014] Among them, the voice encoder is used to process the input data to obtain semantic features;

[0015] Among them, the audio encoder is used to process the input data to obtain audio features.

[0016] Optionally, the alignment module includes: a continuous integration excitation module, a convolutional projection layer, a required content encoder, a non-required content encoder, and a disentanglement module;

[0017] Among them, the continuous integration excitation module is used to perform length reduction processing on the semantic features to obtain the reduced semantic features;

[0018] Among them, the convolutional projection layer is used to perform length reduction processing on the audio features to obtain the reduced audio features;

[0019] Among them, the required content encoder is used to encode the language content information from the voice input;

[0020] Among them, the non-required content encoder is used to model non-linguistic language features;

[0021] Among them, the disentanglement module is used to separate different modality information.

[0022] Optionally, in step S4, inputting the cleaned voice data set and the pre-trained task instruction text into the initial large voice model, and performing decoupling processing on the cleaned voice data set through the alignment module to obtain the final voice features includes:

[0023] S41. Input the cleaned voice data set and the pre-trained task instruction text into the initial large voice model, process the cleaned voice data set through the voice encoder to obtain semantic features; process the cleaned voice data set through the audio encoder to obtain audio features;

[0024] S42. Input the semantic features into the continuous integration excitation module for length reduction processing to obtain the reduced semantic features; input the audio features into the convolutional projection layer for length reduction processing to obtain the reduced audio features;

[0025] S43. Concatenate the reduced semantic features and the reduced audio features at the sequence level to obtain the concatenated features; input the concatenated features into the required content encoder for processing to obtain the first concatenated features; input the concatenated features into the non-required content encoder for processing to obtain the second concatenated features.

[0026] S44. Input the first concatenated features and the second concatenated features into the disentanglement module for decoupling processing to obtain the final speech features.

[0027] Optionally, the step S4 of pre-training the initial speech large model according to the final feature sequence to obtain the pre-trained speech large model includes:

[0028] Input the final feature sequence into the large language model for decoding to output the decoding result; pre-train the initial speech large model according to the decoding result and the cross-entropy loss function to obtain the pre-trained speech large model.

[0029] Optionally, the instruction fine-tuning training set in step S5 includes: an instruction fine-tuning speech data set and the task instruction text corresponding to the speech data set.

[0030] Optionally, the step S5 of performing instruction fine-tuning on the pre-trained speech large model by using the LoRA fine-tuning technique according to the instruction fine-tuning training set to obtain the trained speech large model includes:

[0031] S51. Input the instruction fine-tuning speech data set and the task instruction text corresponding to the speech data set into the pre-trained speech large model, and process the task instruction text corresponding to the speech data set through the tokenizer of the large language model to obtain the instruction fine-tuning text features.

[0032] S52. Process the instruction fine-tuning speech data set in a decoupled manner to obtain the instruction fine-tuning speech features.

[0033] S53. Concatenate the instruction fine-tuning text features and the instruction fine-tuning speech features at the sequence level to obtain the concatenated feature sequence; input the feature sequence into the large language model, and perform instruction fine-tuning on the large language model by using the LoRA fine-tuning technique to obtain the trained speech large model.

[0034] On the other hand, a speech large model modality alignment device based on a two-stage decoupling method is provided. This device is applied to the speech large model modality alignment method based on the two-stage decoupling method, and the device includes:

[0035] A first acquisition unit for acquiring a pre-trained speech data set and a pre-trained task instruction text.

[0036] A second acquisition unit, configured to clean the speech dataset by using the Librosa method to obtain a cleaned speech dataset;

[0037] A construction unit, configured to construct an initial speech large model; the initial speech large model includes: a large language model and an alignment module;

[0038] A pre-training unit, configured to input the cleaned speech dataset and the pre-trained task instruction text into the initial speech large model, decouple the cleaned speech dataset through the alignment module to obtain final speech features; process the pre-trained task instruction text by using the tokenizer of the large language model to obtain pre-trained text features; process the final speech features and the pre-trained text features through sequence-level concatenation to obtain a final feature sequence; pre-train the initial speech large model according to the final feature sequence to obtain a pre-trained speech large model;

[0039] An instruction fine-tuning unit, configured to obtain an instruction fine-tuning training set; fine-tune the pre-trained speech large model by using the LoRA fine-tuning technique according to the instruction fine-tuning training set to obtain a trained speech large model;

[0040] An output unit, configured to obtain speech data to be processed and an instruction corresponding to the speech data; input the speech data to be processed and the instruction corresponding to the speech data into the trained speech large model for processing, and output a text that matches the instruction requirement corresponding to the speech data.

[0041] Optionally, the initial speech large model further includes: a speech encoder and an audio encoder;

[0042] Wherein, the speech encoder is configured to process input data to obtain semantic features;

[0043] Wherein, the audio encoder is configured to process input data to obtain audio features.

[0044] Optionally, the alignment module includes: a continuous integration excitation module, a convolutional projection layer, a required content encoder, a non-required content encoder, and a disentanglement module;

[0045] Wherein, the continuous integration excitation module is configured to perform length reduction processing on the semantic features to obtain reduced semantic features;

[0046] Wherein, the convolutional projection layer is configured to perform length reduction processing on the audio features to obtain reduced audio features;

[0047] Among them, the required content encoder is used to encode language content information from the speech input;

[0048] Among them, the non-required content encoder is used to model non-linguistic language features;

[0049] Among them, the disentanglement module is used to separate different modality information.

[0050] Optionally, inputting the cleaned speech data set and the pre-trained task instruction text into the initial speech large model, and decoupling the cleaned speech data set through the alignment module to obtain the final speech features, including:

[0051] Inputting the cleaned speech data set and the pre-trained task instruction text into the initial speech large model, processing the cleaned speech data set through a speech encoder to obtain semantic features; processing the cleaned speech data set through an audio encoder to obtain audio features;

[0052] Inputting the semantic features into a continuous integration excitation module for length reduction processing to obtain reduced semantic features; inputting the audio features into a convolutional projection layer for length reduction processing to obtain reduced audio features;

[0053] Performing sequence-level concatenation on the reduced semantic features and the reduced audio features to obtain concatenated features; inputting the concatenated features into the required content encoder for processing to obtain first concatenated features; inputting the concatenated features into the non-required content encoder for processing to obtain second concatenated features;

[0054] Inputting the first concatenated features and the second concatenated features into the disentanglement module for decoupling processing to obtain the final speech features.

[0055] Optionally, pre-training the initial speech large model according to the final feature sequence to obtain a pre-trained speech large model, including:

[0056] Inputting the final feature sequence into a large language model for decoding to output a decoding result; pre-training the initial speech large model according to the decoding result and the cross-entropy loss function to obtain a pre-trained speech large model.

[0057] Optionally, the instruction fine-tuning training set includes: an instruction fine-tuning speech data set and the task instruction text corresponding to the speech data set.

[0058] Optionally, the instruction fine-tuning unit is used for:

[0059] Input the voice dataset for instruction fine-tuning and the task instruction text corresponding to the voice dataset into the pre-trained large speech model. Process the task instruction text corresponding to the voice dataset through the tokenizer of the large language model to obtain the text features for instruction fine-tuning.

[0060] Process the voice dataset for instruction fine-tuning in a decoupled manner to obtain the voice features for instruction fine-tuning.

[0061] Concatenate the text features for instruction fine-tuning and the voice features for instruction fine-tuning at the sequence level to obtain the concatenated feature sequence. Input the feature sequence into the large language model and perform instruction fine-tuning on the large language model through the LoRA fine-tuning technique to obtain the trained large speech model.

[0062] On the other hand, provide a voice large model modality alignment device based on a two-stage decoupling method. The voice large model modality alignment device based on the two-stage decoupling method includes: a processor; a memory, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor, any of the methods in the above voice large model modality alignment method based on the two-stage decoupling method is implemented.

[0063] On the other hand, provide a computer-readable storage medium, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement any of the methods in the above voice large model modality alignment method based on the two-stage decoupling method.

[0064] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:

[0065] In the embodiments of the present invention, first, a pre-trained speech dataset and pre-trained task instruction texts are obtained; the Librosa method is used to clean the speech dataset to obtain a cleaned speech dataset; secondly, an initial speech large model is constructed; the initial speech large model includes: a large language model and an alignment module; the cleaned speech dataset and the pre-trained task instruction texts are input into the initial speech large model, and the cleaned speech dataset is decoupled through the alignment module to obtain final speech features; the pre-trained task instruction texts are processed using the tokenizer of the large language model to obtain pre-trained text features; the final speech features and the pre-trained text features are processed through sequence-level concatenation to obtain a final feature sequence; according to the final feature sequence, the initial speech large model is pre-trained to obtain a pre-trained speech large model; an instruction fine-tuning training set is obtained; according to the instruction fine-tuning training set, the LoRA fine-tuning technique is used to perform instruction fine-tuning on the pre-trained speech large model to obtain a trained speech large model; finally, the speech data to be processed and the instruction corresponding to the speech data are obtained; the speech data to be processed and the instruction corresponding to the speech data are input into the trained speech large model for processing, and a text matching the instruction requirement corresponding to the speech data is output.

[0066] In the modal alignment of the embodiments of the present invention, two different encoders and different length reduction methods are used for the first-stage decoupling, fully extracting the required information from the input speech and decoupling with different degrees of length reduction. In the information integration process, sequence-level concatenation and disentanglement modules are used for the second-stage decoupling, reducing the loss of original information in the information integration process and performing screening and filtering. Finally, an input representation meeting the requirements of downstream tasks is obtained. Using the present invention can solve the problem of information loss caused by feature decoupling and improve the performance in speech comprehensive understanding and analysis tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0068] Figure 1 It is a schematic structural diagram of a method for modal alignment of a speech large model based on a two-stage decoupling method provided by the embodiments of the present invention;

[0069] Figure 2 It is a flowchart of a method for modal alignment of a speech large model based on a two-stage decoupling method provided by the embodiments of the present invention;

[0070] Figure 3 It is a block diagram of a speech large model modality alignment device based on a two-stage decoupling method provided by an embodiment of the present invention;

[0071] Figure 4 It is a schematic structural diagram of a speech large model modality alignment device based on a two-stage decoupling method provided by an embodiment of the present invention. Specific embodiments

[0072] Next, the technical solutions in the present invention will be described with reference to the accompanying drawings.

[0073] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or more advantageous than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two can be selected.

[0074] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same.

[0075] In the embodiments of the present invention, sometimes subscripts such as W 1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meanings they express are the same.

[0076] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0077] The embodiments of the present invention provide a speech large model modality alignment method based on a two-stage decoupling method. This method can be implemented by a speech large model modality alignment device based on a two-stage decoupling method. The speech large model modality alignment device based on a two-stage decoupling method can be a terminal or a server. As Figure 1It is a schematic structural diagram of a voice large model modal alignment method based on a two-stage decoupling method provided by an embodiment of the present invention; in a feasible implementation manner, the obtained voice data and the task instruction text corresponding to the voice data are input into the initially constructed voice large model. The voice encoder processes the voice data to obtain semantic features; the audio encoder processes the voice data to obtain audio features; the semantic features are input into the continuous integration excitation module for length reduction processing to obtain reduced semantic features; the audio features are input into the convolutional projection layer for length reduction processing to obtain reduced audio features; the reduced semantic features and the reduced audio features are serially concatenated to obtain a concatenated result; the concatenated result is input into the required content encoder for encoding processing to obtain a first encoding result; the concatenated result is input into the non-required content encoder for encoding processing to obtain a second encoding result; according to the first encoding result and the second encoding result, they are input into the disentanglement module to output the final voice features; the input task instruction text is processed by the tokenizer of the backend large model to obtain task instruction text features; wherein, the backend large model is a large language model; according to the final voice features and the task instruction text features, they are input into the backend large model for training, and the initially constructed voice large model is further trained to obtain a pre-trained voice large model; according to the pre-trained voice large model, the LoRA fine-tuning technology is used to perform instruction fine-tuning training on the pre-trained voice large model to obtain a trained voice large model.

[0078] As Figure 2 shown in the flowchart of the voice large model modal alignment method based on the two-stage decoupling method, the processing flow of this method can include the following steps:

[0079] S1. Obtain a pre-trained voice dataset and a pre-trained task instruction text.

[0080] In a feasible implementation manner, the pre-trained voice dataset can be a voice downstream task dataset including: datasets in voice processing related application scenarios such as speech recognition, speech translation, and speech emotion recognition.

[0081] S2. Use the Librosa method to clean the voice dataset to obtain a cleaned voice dataset.

[0082] In a feasible implementation manner, the present application can use the Librosa method, the Soundfile method, and the Montreal Forced Aligner method to clean the pre-trained voice dataset.

[0083] S3. Construct an initial voice large model; the initial voice large model includes: a large language model and an alignment module.

[0084] Among them, the large language model refers to a deep learning model trained using a large amount of text data, which can generate natural language text or understand the meaning of language text.

[0085] Optionally, the initial speech large model further includes: a speech encoder and an audio encoder;

[0086] Among them, the speech encoder is used to process the input data to obtain semantic features;

[0087] Among them, the audio encoder is used to process the input data to obtain audio features.

[0088] Among them, the alignment module is used to establish a corresponding relationship for different modal data.

[0089] In a feasible implementation, by separately using the speech encoder and the audio encoder to encode the input speech, the semantic information and acoustic information of the speech can be considered simultaneously, and the decoupling at the encoding end can be effectively achieved, which helps the model to extract more effective information from the speech.

[0090] Optionally, the alignment module includes: a continuous integration excitation module, a convolutional projection layer, a desired content encoder, an undesired content encoder, and a disentanglement module;

[0091] Among them, the continuous integration excitation module is used to perform length reduction processing on the semantic features to obtain the reduced semantic features;

[0092] Among them, the convolutional projection layer is used to perform length reduction processing on the audio features to obtain the reduced audio features;

[0093] Among them, the desired content encoder is used to encode the language content information from the speech input;

[0094] Among them, the undesired content encoder is used to model non-linguistic language features;

[0095] Among them, the disentanglement module is used to separate different modal information.

[0096] Among them, the disentanglement module can reduce the correlation between features and improve the performance and robustness of the speech large model.

[0097] In a feasible implementation, for different encoded information, two length reduction methods, namely the continuous integration excitation module and the convolutional projection layer, are used to reduce the length of the feature sequence, effectively shortening the sequence length of the final input large language model and improving the accuracy and efficiency of model inference.

[0098] In a feasible implementation, during information integration, different feature representation sequences are concatenated in terms of dimensions, reducing the loss of the extracted original information. Then, a disentanglement module is used to complete the second-stage decoupling. The integrated information is further screened and filtered for downstream tasks, retaining the information most suitable for the current downstream task, thereby improving the performance of the downstream task.

[0099] S4. Input the cleaned speech data set and the pre-trained task instruction text into the initial speech large model. Process the cleaned speech data set through a decoupling method to obtain the final speech features; use the tokenizer of the large language model to process the pre-trained task instruction text to obtain the pre-trained text features; process the final speech features and the pre-trained text features through a sequence-level concatenation method to obtain the final feature sequence; pre-train the initial speech large model according to the final feature sequence to obtain the pre-trained speech large model.

[0100] Among them, the pre-training loss function is the cross-entropy loss function.

[0101] In a feasible implementation, the cross-entropy loss function and the AdamW optimizer are used to train and optimize the model parameters.

[0102] Among them, this application uses a two-stage decoupling method to process the speech input. The first-stage decoupling includes: using a speech encoder and an audio encoder to encode the input speech data to obtain semantic features and audio features, using a continuous integration excitation module and a convolutional projection layer to respectively perform length compression processing on the semantic features and the audio features, and projecting them onto the same dimension for sequence-level concatenation to obtain a concatenation result.

[0103] Among them, the second-stage decoupling includes: using an alignment module to process the concatenation result to obtain the final speech features.

[0104] Optionally, in S4, inputting the cleaned speech data set and the pre-trained task instruction text into the initial speech large model, and using the alignment module to perform decoupling processing on the cleaned speech data set to obtain the final speech features includes:

[0105] S41. Input the cleaned speech data set and the pre-trained task instruction text into the initial speech large model. Use the speech encoder to process the cleaned speech data set to obtain semantic features; use the audio encoder to process the cleaned speech data set to obtain audio features.

[0106] S42. Input the semantic features into the continuous integration excitation module for length reduction processing to obtain the reduced semantic features; input the audio features into the convolutional projection layer for length reduction processing to obtain the reduced audio features.

[0107] S43. Concatenate the reduced semantic features and the reduced audio features at the sequence level to obtain the concatenated features; input the concatenated features into the required content encoder for processing to obtain the first concatenated features; input the concatenated features into the non-required content encoder for processing to obtain the second concatenated features.

[0108] S44. Input the first concatenated features and the second concatenated features into the disentanglement module for decoupling processing to obtain the final speech features.

[0109] Optionally, the pre-training of the initial speech large model according to the final feature sequence to obtain the pre-trained speech large model includes:

[0110] Input the final feature sequence into the large language model for decoding and output the decoding result; pre-train the initial speech large model according to the decoding result and the cross-entropy loss function to obtain the pre-trained speech large model.

[0111] Among them, only the alignment module is trained during the pre-training process, and the speech large model is frozen, aiming to provide good initial parameters for the alignment module.

[0112] S5. Obtain the instruction fine-tuning training set; according to the instruction fine-tuning training set, use the LoRA fine-tuning technique to perform instruction fine-tuning on the pre-trained speech large model to obtain the trained speech large model.

[0113] Optionally, the instruction fine-tuning training set in S5 includes: the speech dataset for instruction fine-tuning and the task instruction text corresponding to the speech dataset.

[0114] Optionally, the specific implementation process of S5 includes S51 - S53:

[0115] S51. Input the speech dataset for instruction fine-tuning and the task instruction text corresponding to the speech dataset into the pre-trained speech large model, and process the task instruction text corresponding to the speech dataset through the tokenizer of the large language model to obtain the text features for instruction fine-tuning.

[0116] S52. Process the speech dataset for instruction fine-tuning in a decoupled manner to obtain the speech features for instruction fine-tuning.

[0117] S53. Concatenate the text features of instruction fine-tuning and the speech features of instruction fine-tuning at the sequence level to obtain a concatenated feature sequence; input the feature sequence into the large language model, and fine-tune the large language model through the LoRA fine-tuning technique to obtain a trained speech large model.

[0118] In a feasible implementation, obtain the speech training set for instruction fine-tuning training and the task instruction text corresponding to the speech training set; after performing the same speech encoding and two-stage decoupling process on the task instruction text and the speech data as in the pre-training process, input them into the large language model, output the result that meets the instruction requirements, and optimize the model parameters using the same cross-entropy loss and optimizer as in the pre-training. Among them, during the instruction fine-tuning training process, the large language model is no longer frozen, and the model parameters are fine-tuned through the LoRA fine-tuning technique of the Peft library to obtain a trained speech large model.

[0119] S6. Obtain the speech data to be processed and the instruction corresponding to the speech data; input the speech data to be processed and the instruction corresponding to the speech data into the trained speech large model for processing, and output the text that matches the instruction requirements corresponding to the speech data.

[0120] In a feasible implementation, obtain the speech data to be processed and the task instruction text corresponding to the speech data; input the speech data to be processed and the task instruction text corresponding to the speech data into the trained speech large model for processing, and output the text that matches the instruction requirements. For example, input a piece of speech and the text of the speech translation instruction into the trained speech large model, and output the text result after translating the actual content of the speech into the target language.

[0121] In the embodiments of the present invention, first, a pre-trained speech dataset and pre-trained task instruction texts are obtained; the Librosa method is used to clean the speech dataset to obtain a cleaned speech dataset; secondly, an initial speech large model is constructed; the initial speech large model includes: a large language model and an alignment module; the cleaned speech dataset and the pre-trained task instruction texts are input into the initial speech large model, and the cleaned speech dataset is decoupled by the alignment module to obtain final speech features; the pre-trained task instruction texts are processed by the tokenizer of the large language model to obtain pre-trained text features; the final speech features and the pre-trained text features are processed by sequence-level concatenation to obtain a final feature sequence; according to the final feature sequence, the initial speech large model is pre-trained to obtain a pre-trained speech large model; an instruction fine-tuning training set is obtained; according to the instruction fine-tuning training set, the LoRA fine-tuning technique is used to perform instruction fine-tuning on the pre-trained speech large model to obtain a trained speech large model; finally, the speech data to be processed and the instruction corresponding to the speech data are obtained; the speech data to be processed and the instruction corresponding to the speech data are input into the trained speech large model for processing, and a text matching the instruction requirement corresponding to the speech data is output.

[0122] In the modal alignment of the embodiments of the present invention, two different encoders and different length reduction methods are used for the first-stage decoupling, and the required information is fully mined from the input speech and decoupled with different degrees of length reduction. In the information integration process, sequence-level concatenation and disentanglement modules are used for the second-stage decoupling, reducing the loss of original information in the information integration process and performing screening and filtering. Finally, an input representation that meets the requirements of downstream tasks is obtained. The present invention can solve the problem of information loss caused by feature decoupling and improve the performance in speech comprehensive understanding and analysis tasks.

[0123] Figure 3 It is a block diagram of a speech large model modal alignment device based on a two-stage decoupling method shown according to an exemplary embodiment. The device is used for the speech large model modal alignment method based on the two-stage decoupling method. Refer to Figure 3 , the device includes a first acquisition unit 310, a second acquisition unit 320, a construction unit 330, a pre-training unit 340, an instruction fine-tuning unit 350, and an output unit 360. Among them:

[0124] The first acquisition unit 310 is used to acquire a pre-trained speech dataset and pre-trained task instruction texts;

[0125] The second acquisition unit 320 is used to clean the speech dataset by using the Librosa method to obtain a cleaned speech dataset;

[0126] Building unit 330, for building an initial speech large model; the initial speech large model includes: a large language model and an alignment module;

[0127] Pretraining unit 340, for inputting the cleaned speech dataset and the pre-trained task instruction text into the initial speech large model, decoupling the cleaned speech dataset through the alignment module to obtain the final speech features; using the tokenizer of the large language model to process the pre-trained task instruction text to obtain pre-trained text features; processing the final speech features and the pre-trained text features through sequence-level concatenation to obtain the final feature sequence; pre-training the initial speech large model according to the final feature sequence to obtain a pre-trained speech large model;

[0128] Instruction fine-tuning unit 350, for obtaining an instruction fine-tuning training set; according to the instruction fine-tuning training set, using the LoRA fine-tuning technique to perform instruction fine-tuning on the pre-trained speech large model to obtain a trained speech large model;

[0129] Output unit 360, for obtaining the speech data to be processed and the instruction corresponding to the speech data; inputting the speech data to be processed and the instruction corresponding to the speech data into the trained speech large model for processing, and outputting the text that matches the instruction requirement corresponding to the speech data.

[0130] Optionally, the initial speech large model further includes: a speech encoder and an audio encoder;

[0131] Among them, the speech encoder is used to process the input data to obtain semantic features;

[0132] Among them, the audio encoder is used to process the input data to obtain audio features.

[0133] Optionally, the alignment module includes: a continuous integration excitation module, a convolutional projection layer, a required content encoder, a non-required content encoder, and a disentanglement module;

[0134] Among them, the continuous integration excitation module is used to perform length reduction processing on the semantic features to obtain the reduced semantic features;

[0135] Among them, the convolutional projection layer is used to perform length reduction processing on the audio features to obtain the reduced audio features;

[0136] Among them, the required content encoder is used to encode language content information from the speech input;

[0137] Among them, the non-desired content encoder is used to model non-linguistic language features;

[0138] Among them, the disentanglement module is used to separate different modality information.

[0139] Optionally, inputting the cleaned speech data set and the pre-trained task instruction text into the initial speech large model, and decoupling the cleaned speech data set through the alignment module to obtain the final speech features, including:

[0140] Inputting the cleaned speech data set and the pre-trained task instruction text into the initial speech large model, processing the cleaned speech data set through a speech encoder to obtain semantic features; processing the cleaned speech data set through an audio encoder to obtain audio features;

[0141] Inputting the semantic features into a continuous integration excitation module for length reduction processing to obtain reduced semantic features; inputting the audio features into a convolutional projection layer for length reduction processing to obtain reduced audio features;

[0142] Sequentially concatenating the reduced semantic features and the reduced audio features to obtain concatenated features; inputting the concatenated features into the desired content encoder for processing to obtain first concatenated features; inputting the concatenated features into the non-desired content encoder for processing to obtain second concatenated features;

[0143] Inputting the first concatenated features and the second concatenated features into the disentanglement module for decoupling processing to obtain the final speech features.

[0144] Optionally, pre-training the initial speech large model according to the final feature sequence to obtain a pre-trained speech large model, including:

[0145] Inputting the final feature sequence into a large language model for decoding to output a decoding result; pre-training the initial speech large model according to the decoding result and the cross-entropy loss function to obtain a pre-trained speech large model.

[0146] Optionally, the instruction fine-tuning training set includes: an instruction fine-tuning speech data set and a task instruction text corresponding to the speech data set.

[0147] Optionally, the instruction fine-tuning unit 350 is used for:

[0148] Input the speech dataset for instruction fine-tuning and the task instruction text corresponding to the speech dataset into the pre-trained large speech model. Process the task instruction text corresponding to the speech dataset through the tokenizer of the large language model to obtain the text features for instruction fine-tuning;

[0149] Process the speech dataset for instruction fine-tuning in a decoupled manner to obtain the speech features for instruction fine-tuning;

[0150] Perform sequence-level concatenation on the text features for instruction fine-tuning and the speech features for instruction fine-tuning to obtain the concatenated feature sequence; input the feature sequence into the large language model, and perform instruction fine-tuning on the large language model through the LoRA fine-tuning technique to obtain the trained large speech model.

[0151] In the embodiment of the present invention, first, obtain the pre-trained speech dataset and the pre-trained task instruction text; use the Librosa method to clean the speech dataset to obtain the cleaned speech dataset; secondly, construct an initial large speech model; the initial large speech model includes: a large language model and an alignment module; input the cleaned speech dataset and the pre-trained task instruction text into the initial large speech model, and perform decoupling processing on the cleaned speech dataset through the alignment module to obtain the final speech features; use the tokenizer of the large language model to process the pre-trained task instruction text to obtain the pre-trained text features; perform processing on the final speech features and the pre-trained text features through sequence-level concatenation to obtain the final feature sequence; according to the final feature sequence, perform pre-training on the initial large speech model to obtain the pre-trained large speech model; obtain the instruction fine-tuning training set; according to the instruction fine-tuning training set, perform instruction fine-tuning on the pre-trained large speech model through the LoRA fine-tuning technique to obtain the trained large speech model; finally, obtain the speech data to be processed and the instruction corresponding to the speech data; input the speech data to be processed and the instruction corresponding to the speech data into the trained large speech model for processing, and output the text that matches the instruction requirements corresponding to the speech data.

[0152] In the modal alignment of the embodiment of the present invention, two different encoders and different length reduction methods are used for the first-stage decoupling, fully mining the required information from the input speech and performing different degrees of length reduction after decoupling. In the information integration process, sequence-level concatenation and disentanglement modules are used for the second-stage decoupling, reducing the loss of original information in the information integration process and performing screening and filtering. Finally, the input representation that meets the requirements of downstream tasks is obtained. Using the present invention can solve the problem of information loss caused by feature decoupling and improve the performance in speech comprehensive understanding and analysis tasks.

[0153] Figure 4It is a schematic structural diagram of a voice large model modality alignment device based on a two-stage decoupling method provided by an embodiment of the present invention. As Figure 4 shown, the voice large model modality alignment device based on the two-stage decoupling method may include the above Figure 3 shown voice large model modality alignment device based on the two-stage decoupling method. Optionally, the voice large model modality alignment device 410 based on the two-stage decoupling method may include a first processor 2001.

[0154] Optionally, the voice large model modality alignment device 410 based on the two-stage decoupling method may further include a memory 2002 and a transceiver 2003.

[0155] Among them, the first processor 2001, the memory 2002, and the transceiver 2003, such as, may be connected through a communication bus.

[0156] Next, in combination with Figure 4 each component of the voice large model modality alignment device 410 based on the two-stage decoupling method will be specifically introduced:

[0157] Among them, the first processor 2001 is the control center of the voice large model modality alignment device 410 based on the two-stage decoupling method, which may be a processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or may be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).

[0158] Optionally, the first processor 2001 may execute various functions of the voice large model modality alignment device 410 based on the two-stage decoupling method by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.

[0159] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 4 the CPU0 and CPU1 shown in

[0160] In a specific implementation, as an example, the speech large model modality alignment device 410 based on the two-stage decoupling method may also include multiple processors, such as Figure 4 the first processor 2001 and the second processor 2004 shown in

[0161] Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0162] Among them, the memory 2002 is used to store the software program for implementing the solution of the present invention and is controlled by the first processor 2001 for execution. The specific implementation method can refer to the above method embodiment and will not be elaborated here. Figure 4 Optionally, the memory 2002 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic storage media such as magnetic disk storage, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 can be integrated with the first processor 2001 or exist independently and is coupled to the first processor 2001 through the interface circuit ( Figure 4 not shown in

[0163] The transceiver 2003 is used to communicate with a network device or with a terminal device.

[0164] Optionally, the transceiver 2003 can include a receiver and a transmitter ( Figure 4 not shown separately in

[0165] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or exist independently, and is coupled to the first processor 2001 through the interface circuit ( Figure 4 not shown) of the voice large model modality alignment device 410 based on the two-stage decoupling method. The embodiments of the present invention do not make specific limitations on this.

[0166] It should be noted that Figure 4 the structure of the voice large model modality alignment device 410 based on the two-stage decoupling method shown in

[0167] does not constitute a limitation to the router. The actual knowledge structure recognition device may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0168] It should be understood that the first processor 2001 in the embodiments of the present invention can be a central processing unit (CPU), and this processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or this processor can also be any conventional processor, etc.

[0169] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0170] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0171] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be understood specifically with reference to the context.

[0172] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0173] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0174] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0175] Those skilled in the art can clearly understand that for the sake of convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0176] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0177] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0178] In addition, the functional units in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0179] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0180] As described above, the above are only specific implementation manners of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A speech large model modal alignment method based on a two-stage decoupling method, characterized in that: The method comprises: S1, obtaining a pre-trained speech data set and a pre-trained task instruction text; S2, using the Librosa method to clean the speech data set to obtain a cleaned speech data set; S3, constructing an initial large speech model; the initial large speech model includes: a large language model and an alignment module; Wherein, the initial large speech model further includes: a speech encoder and an audio encoder; The speech encoder is used to process the input data to obtain semantic features; The audio encoder is used to process the input data to obtain audio features; The alignment module includes: a continuous integration excitation module, a convolutional projection layer, a desired content encoder, an undesired content encoder, and a disentanglement module; The continuous integration excitation module is used to reduce the length of the semantic features to obtain the reduced semantic features; The convolutional projection layer is used to reduce the length of the audio features to obtain the reduced audio features; Wherein, the required content encoder is used to encode language content information from speech input; Wherein, the undesired content encoder is used to model non-linguistic language features; Wherein, the de-entanglement module is used to separate different modal information; S4, inputting the cleaned speech data set and the pre-trained task instruction text into the initial speech large model, decoupling the cleaned speech data set through the alignment module to obtain the final speech features; using the word segmenter of the large language model to process the pre-trained task instruction text to obtain the pre-trained text features; processing the final speech features and the pre-trained text features through sequence-level splicing to obtain a final feature sequence; pre-training the initial speech large model according to the final feature sequence to obtain a pre-trained speech large model; The step S4 pre-trains the initial speech model according to the final feature sequence to obtain a pre-trained speech model, including: Input the final feature sequence into the large language model for decoding, and output the decoding result; pre-train the initial large speech model according to the decoding result and the cross entropy loss function to obtain a pre-trained large speech model; The step S4 inputs the cleaned speech data set and the pre-trained task instruction text into the initial speech large model, and decouples the cleaned speech data set through the alignment module to obtain the final speech features, including: S41, inputting the cleaned speech data set and the pre-trained task instruction text into the initial speech large model, processing the cleaned speech data set through a speech encoder to obtain semantic features; processing the cleaned speech data set through an audio encoder to obtain audio features; S42, inputting the semantic features into the continuous integration excitation module for length reduction processing to obtain reduced semantic features; inputting the audio features into the convolutional projection layer for length reduction processing to obtain reduced audio features; S43, concatenating the reduced semantic features and the reduced audio features at a sequence level to obtain concatenated features; inputting the concatenated features into a required content encoder for processing to obtain a first concatenated feature; inputting the concatenated features into a non-required content encoder for processing to obtain a second concatenated feature; S44, inputting the first concatenated feature and the second concatenated feature into a de-entanglement module for decoupling processing to obtain a final speech feature; S5. Obtain a command fine-tuning training set; according to the command fine-tuning training set, use LoRA fine-tuning technology to perform command fine-tuning on the pre-trained voice model to obtain a trained voice model; The instruction fine-tuning training set of S5 includes: a speech data set for instruction fine-tuning and a task instruction text corresponding to the speech data set; Among them, the step S5 fine-tunes the pre-trained speech model according to the instruction fine-tuning training set, adopts the LoRA fine-tuning technology to perform instruction fine-tuning on the pre-trained speech model, and obtains the trained speech model, including: S51, inputting the speech data set of instruction fine-tuning and the task instruction text corresponding to the speech data set into the pre-trained speech large model, processing the task instruction text corresponding to the speech data set through the word segmenter of the large language model, and obtaining the text features of the instruction fine-tuning; S52, processing the voice data set for instruction fine-tuning in a decoupling manner to obtain voice features for instruction fine-tuning; S53, concatenating the text features of the instruction fine-tuning and the voice features of the instruction fine-tuning at the sequence level to obtain a concatenated feature sequence; inputting the feature sequence into the large language model, and performing instruction fine-tuning on the large language model through the LoRA fine-tuning technology to obtain a trained voice large model; S6. Obtain the speech data to be processed and the instructions corresponding to the speech data; input the speech data to be processed and the instructions corresponding to the speech data into the trained speech model for processing, and output text that matches the instruction requirements corresponding to the speech data.

2. A speech large model modal alignment device based on a two-stage decoupling method, the speech large model modal alignment device based on a two-stage decoupling method is used to implement the speech large model modal alignment method based on a two-stage decoupling method as claimed in claim 1, characterized in that: The device comprises: A first acquisition unit, used to acquire a pre-trained speech data set and a pre-trained task instruction text; A second acquisition unit is used to clean the speech data set using the Librosa method to obtain a cleaned speech data set; A construction unit, used to construct an initial large speech model; the initial large speech model includes: a large language model and an alignment module; A pre-training unit is used to input the cleaned speech data set and the pre-trained task instruction text into the initial speech large model, decouple the cleaned speech data set through the alignment module to obtain the final speech features; use the word segmenter of the large language model to process the pre-trained task instruction text to obtain the pre-trained text features; process the final speech features and the pre-trained text features through sequence-level splicing to obtain a final feature sequence; pre-train the initial speech large model according to the final feature sequence to obtain a pre-trained speech large model; An instruction fine-tuning unit is used to obtain an instruction fine-tuning training set; according to the instruction fine-tuning training set, the pre-trained speech model is fine-tuned by using the LoRA fine-tuning technology to obtain a trained speech model; The output unit is used to obtain the voice data to be processed and the instructions corresponding to the voice data; the voice data to be processed and the instructions corresponding to the voice data are input into the trained voice model for processing, and the text matching the instruction requirements corresponding to the voice data is output.

3. A speech large model modal alignment device based on a two-stage decoupling method, characterized in that: The speech large model modal alignment device based on the two-stage decoupling method includes: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to claim 1 is implemented.

4. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to claim 1.

Citation Information

Patent Citations

  • Voice annotation method and device for voice feature description

    CN118571229A