Low-resource language adaptive speech recognition method based on AdLoRA Plus

By introducing AdLoRA Plus fine-tuning module into the Whisper model, the problem of low-resource language speech recognition accuracy is solved, and efficient speech recognition in Indonesian language takeaway customer service scenarios is achieved, reducing computing overhead and improving the adaptability of the model.

CN119943052AActive Publication Date: 2025-05-06XIAN UNIV OF TECH

Patent Information

Application Number
CN202510037717.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-06
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

Existing speech recognition systems perform poorly when dealing with low-resource languages ​​(such as Indonesian and their dialects), especially in takeaway customer service scenarios, there are problems such as insufficient training data, inefficient traditional fine-tuning methods, and poor performance in specific fields.

Method used

Adaptive speech recognition method based on AdLoRA Plus is adopted. By obtaining speech data and using the Whisper model for preliminary transcription, commonly used scene dialogue texts are screened, and the AdLoRA Plus fine-tuning module is added to the Whisper model, and training is used for training, and optimization models are generated to improve speech recognition accuracy.

Benefits of technology

It significantly improves the speech recognition accuracy of the Whisper model in a low-resource language environment, reduces computing overhead and training time, improves the adaptability and scalability of the model, and ensures outstanding performance in practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943052A_ABST
    Figure CN119943052A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of speech recognition. The invention provides a low-resource language self-adaptive speech recognition method based on AdLoRA Plus. The low-resource language self-adaptive speech recognition method comprises the following steps of: obtaining a speech According to the embodiment of the invention, the AdLoRA Plus fine tuning technology is introduced, so that the speech recognition precision of the Whisper model is remarkably improved in a low-resource language environment. Through a low-rank adapter layer and dynamic learning rate adjustment, efficient parameter updating is realized under limited training data. According to the method, a strategy of freezing basic parameters is adopted, only an AdLoRA Plus adapter layer is trained, calculation overhead and training time are remarkably reduced, pre-training knowledge is kept, and overfitting is prevented. The method has wide adaptability and expansibility, and is ensured to be outstanding in practical application. The voice recognition precision, the reasoning efficiency and the training efficiency are effectively improved, multiple challenges in application in small languages and specific fields in the prior art are solved, and the method has important practical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The disclosed embodiments relate to the field of speech recognition technology, and in particular to a low-resource language adaptive speech recognition method based on AdLoRA Plus. Background Art

[0002] With the development of artificial intelligence technology, voice interaction is becoming increasingly important in all walks of life. In a language-diverse country like Indonesia, the food delivery industry is undergoing a voice-driven transformation. More and more consumers communicate with customer service through voice to obtain order status, delivery details or refund procedures. However, traditional manual customer service is difficult to efficiently respond to massive customer needs, especially when voice recognition is inaccurate, which affects the response speed and quality.

[0003] Existing speech recognition systems perform poorly when dealing with low-resource languages ​​such as Indonesian and its dialects because they are primarily designed for resource-rich languages ​​such as English or Chinese. Indonesian has complex sound change rules and diverse dialects, and the background noise and changing conversation patterns in food delivery customer service scenarios further increase the difficulty of recognition.

[0004] The following are some of the specific defects in related technologies: 1. Insufficient training data: This makes it difficult for the model to fully learn language features. 2. Traditional fine-tuning methods are inefficient: High performance cannot be achieved on limited data. 3. Poor performance in specific areas: Under complex business logic and speech requirements such as takeaway customer service, the model's performance is not ideal.

[0005] Therefore, it is necessary to improve one or more problems existing in the above-mentioned related technical solutions.

[0006] It should be noted that this section is intended to provide background or context for the technical solutions of the present disclosure stated in the claims. The description herein is not admitted to be prior art by virtue of being included in this section. Summary of the invention

[0007] The purpose of the embodiments of the present disclosure is to provide a low-resource language adaptive speech recognition method based on AdLoRA Plus, thereby overcoming one or more problems caused by the limitations and defects of the relevant technology at least to a certain extent.

[0008] According to an embodiment of the present disclosure, a low-resource language adaptive speech recognition method based on AdLoRA Plus is provided, the method comprising: Acquire voice data, and perform preliminary transcription on the voice data using a Whisper model to obtain a text data set; Processing the text dataset using a large language model to filter out dialogue texts in common scenarios; Accurately transcribing the common scene dialogue text to obtain transcribed voice data, and combining the common scene dialogue text to obtain a labeled voice dataset; Preprocessing the labeled speech data set, and dividing the preprocessed labeled speech data set into a training data set and a test data set according to a preset ratio; Adding the AdLoRA Plus fine-tuning module to the Whisper model, training using the training data set and the test data set, and combining the weights of the Whisper model to obtain a trained optimization model; The speech to be recognized is input into the trained optimization model to obtain predicted text data.

[0009] Further, the step of obtaining voice data and preliminarily transcribing the voice data using the Whisper model to obtain a text data set includes: Determining a target application scenario, data standard, and acquisition path of the voice data to obtain the voice data, and performing normalization processing on the voice data; Performing preliminary transcription of the speech data using the pre-trained Whisper model to obtain the text data set; The text data set is associated with the voice data, and the consistency between the text data set and the voice data is checked.

[0010] Furthermore, the step of processing the text dataset using a large language model to filter out dialogue texts in common scenarios includes: Inputting the text data set and the directive prompt into the large language model for analysis; The large language model outputs the common scenario dialogue text.

[0011] Further, the step of accurately transcribing the common scene dialogue text to obtain transcribed voice data, and combining the common scene dialogue text to obtain a labeled voice data set includes: Recording each of the common scenario dialogue texts and recording the duration of each recording to obtain the transcribed voice data; All the commonly used scene dialogue texts and their corresponding transcribed voice data are sorted to obtain the labeled voice data set.

[0012] Further, the step of preprocessing the labeled speech data set and dividing the preprocessed labeled speech data set into a training data set and a test data set according to a preset ratio includes: Using a data filler to perform data segmentation and filling on the labeled speech dataset, and performing data allocation for multi-card training on the labeled speech dataset; Processing the labeled speech data set into a preset input format, performing feature extraction on the labeled speech data set, and standardizing the extracted features to obtain standardized features; The standardized features are filled with data and batch sorted, and divided into the training data set and the test data set according to a preset ratio.

[0013] Further, the step of adding an AdLoRA Plus fine-tuning module to the Whisper model, training using the training data set and the test data set, and combining the weights of the Whisper model to obtain a trained optimization model includes: Configure the task type and language of the Whisper model, and set the forced_decoder_ids or suppress_tokens of the Whisper model; wherein the Whisper model includes a feature extraction module, an encoder, and a decoder; Adding the AdLoRAPlus fine-tuning module in the query layer and value layer of the encoder and the query layer and value layer of the decoder; Use loraplus_lr_ratio to control the learning rate ratio of the low-rank matrix A and the low-rank matrix B in the AdLoRA Plus fine-tuning module, use loraplus_lr_embedding to control the learning rate of the embedding layer, and use RankAllocator to dynamically adjust the ranks of the low-rank matrix A and the low-rank matrix B; Freezing the original parameters of the Whisper model; Training the Whisper model using the training data set to obtain an AdLoRA Plus model; Merging the weights of the AdLoRA Plus model and the Whisper model to obtain an optimized model; The optimization model is tested using the test data set to obtain a trained optimization model.

[0014] Further, the step of training the Whisper model using the training data set to obtain the AdLoRAPlus model includes: Setting hyperparameters according to the configuration; wherein the hyperparameters include at least learning rate, batch size and gradient accumulation steps; Enable half-precision training; if fp16 is set to True, use the compiler to optimize the Whisper model; Use the training data set to train, check the performance of the model regularly according to the set number of evaluation steps, and output the log to tensorboard; Resume or continue training based on the saved checkpoint, saving the final state of the Whisper model and the optimal model to the output directory.

[0015] Further, the step of merging the weights of the AdLoRA Plus model and the Whisper model to obtain an optimized model includes: Check whether the AdLoRA Plus model path exists. If so, load the configuration parameters of the AdLoRA Plus model from the path. Use the Whisper model path obtained from the configuration to load the Whisper model and load it into the device. If the configuration specifies that the model is only loaded locally, set local_files_only=True. The merge_and_unload() method is used to merge the parameters of the AdLoRA Plus model and the weights of the Whisper model into an overall model, and the AdLoRA Plus model itself is cleared to ensure that the parameters of the AdLoRAPlus model are correctly merged into the Whisper model to form the optimized model; wherein, after the merging is completed, the optimized model is set to evaluation mode.

[0016] Furthermore, the step of inputting the speech to be recognized into the trained optimization model to obtain predicted text data includes: Preprocessing the speech to be recognized; The pre-processed speech to be recognized is input into the feature extraction module of the trained optimization model, and after converting the one-dimensional audio waveform data into a two-dimensional spectrogram by short-time Fourier transform, a Mel filter is used to generate a Mel spectrogram; a convolution layer is used to perform dimensionality reduction processing on the Mel spectrogram to extract local features in time and frequency to generate audio features; The audio features are input into the encoder of the trained optimization model, position coding is added to the audio features, and the audio features are passed through a multi-layer encoder to generate global semantic information including speech features; wherein each layer of the encoder comprises a multi-head self-attention module, a residual connection and normalization module, and a multi-layer perceptron module, wherein the multi-head self-attention module is used to capture long-distance dependencies by calculating the global correlation of features; the residual connection and the normalization module are used to ensure the stability of feature distribution; and the multi-layer perceptron module is used to learn high-order nonlinear relationships of features; Input the historical text data and the global semantic information into the multi-layer decoder of the trained optimization model, add learnable position encoding to the historical text data and the global semantic information, model the position information of the text sequence, and output a probability distribution to represent the tag sequence of the current time step; wherein the decoder includes a self-attention module, a cross-attention module and a multi-layer perceptron module, the self-attention module is used to capture the global correlation of the decoder input and model the dependency between historical generated texts; the cross-attention module is used to combine the information of the encoder output and the decoder input, and realize the alignment of audio features and text semantics through the cross-attention mechanism; The tag sequence is decoded into text, restored to human-readable sentences using a predefined vocabulary or subword decomposition method, spliced ​​according to chronological order, punctuation and format correction are added, and the text results are verified using grammar and spelling check tools to generate predicted text data.

[0017] The technical solution provided by the embodiments of the present disclosure may have the following beneficial effects: In the embodiments of the present disclosure, through the above-mentioned low-resource language adaptive speech recognition method based on AdLoRA Plus, on the one hand, by introducing the AdLoRA Plus fine-tuning technology, the speech recognition accuracy of the Whisper model is significantly improved in a low-resource language environment. Through the low-rank adapter layer and dynamic learning rate adjustment, efficient parameter updates are achieved with limited training data. By adopting the strategy of freezing the basic parameters, only the AdLoRA Plus adapter layer is trained, which significantly reduces the computational overhead and training time, while maintaining pre-training knowledge and preventing overfitting. On the other hand, the method has wide adaptability and scalability to ensure excellent performance in practical applications. It effectively improves the speech recognition accuracy, reasoning efficiency and training efficiency, solves many challenges of the prior art in small languages ​​and specific field applications, and has important practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.

[0019] Figure 1 A step diagram showing a low-resource language adaptive speech recognition method based on AdLoRA Plus in an exemplary embodiment of the present disclosure; Figure 2 A specific flow chart of a low-resource language adaptive speech recognition method based on AdLoRA Plus in an exemplary embodiment of the present disclosure is shown; Figure 3 A schematic diagram showing an attention mechanism in an exemplary embodiment of the present disclosure; Figure 4 A flowchart of inference of a trained optimization model in an exemplary embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0020] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the disclosure will be more comprehensive and complete and to fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0021] In addition, the accompanying drawings are only schematic illustrations of the embodiments of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated descriptions will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities.

[0022] This example implementation provides a low-resource language adaptive speech recognition method based on AdLoRA Plus. Figure 2 As shown in , the low-resource language adaptive speech recognition method based on AdLoRA Plus may include: steps S101 to S106.

[0023] Step S101: Acquire speech data, and use the Whisper model to preliminarily transcribe the speech data to obtain a text data set; Step S102: Processing the text data set using a large language model to filter out dialogue texts in common scenarios; Step S103: accurately transcribing the common scene dialogue text to obtain transcribed voice data, and combining the common scene dialogue text to obtain a labeled voice data set; Step S104: preprocessing the labeled speech data set, and dividing the preprocessed labeled speech data set into a training data set and a test data set according to a preset ratio; Step S105: adding an AdLoRA Plus fine-tuning module to the Whisper model, training using the training data set and the test data set, and combining the weights of the Whisper model to obtain a trained optimization model; Step S106: inputting the speech to be recognized into the trained optimization model to obtain predicted text data.

[0024] Through the above-mentioned low-resource language adaptive speech recognition method based on AdLoRA Plus, on the one hand, by introducing the AdLoRA Plus fine-tuning technology, the speech recognition accuracy of the Whisper model is significantly improved in a low-resource language environment. Through the low-rank adapter layer and dynamic learning rate adjustment, efficient parameter updates are achieved with limited training data. The strategy of freezing the basic parameters is adopted, and only the AdLoRA Plus adapter layer is trained, which significantly reduces the computational overhead and training time, while maintaining pre-training knowledge and preventing overfitting. On the other hand, this method has wide adaptability and scalability to ensure excellent performance in practical applications. It effectively improves the speech recognition accuracy, reasoning efficiency and training efficiency, solves many challenges of existing technologies in small languages ​​and specific field applications, and has important practical application value.

[0025] Next, we will refer to Figures 1 to 4 The various steps of the above-mentioned low-resource language adaptive speech recognition method based on AdLoRA Plus in this example implementation are described in more detail.

[0026] In step S101 and step S102, voice data is acquired and initially transcribed using the Whisper model to obtain a text data set. The text data set is processed using a large language model to filter out common scene dialogue texts.

[0027] For example, multi-platform collection. Customer service call records are collected from multiple Indonesian food delivery platforms (such as GoFood, GrabFood, etc.) to ensure that the data covers common scenarios such as order inquiry, delivery status inquiry, refund processing, complaint processing, etc. Through multi-platform data collection, the diversity and representativeness of the data are guaranteed.

[0028] Preprocessing and transcription: Use the pre-trained Whisper model to convert raw speech data into high-quality text data, and then pass it to the Large Language Model (LLM) for further processing to generate multi-round long dialogue scenarios, providing rich annotation information for subsequent fine-tuning.

[0029] Accurate transcription. The obtained text data is transcribed with high precision to ensure that the speech and text content are completely consistent, providing a reliable foundation for model training.

[0030] Temporal alignment. Labeling the duration of each sentence ensures that the model can leverage the temporal alignment information for deeper learning, especially in multi-turn conversations.

[0031] In step S103 and step S104, the common scenario dialogue text is accurately transcribed to obtain transcribed voice data, and combined with the common scenario dialogue text to obtain a labeled voice data set.

[0032] For example, the labeled speech data set is preprocessed, and the preprocessed labeled speech data set is divided into a training data set and a test data set according to a preset ratio. Load the Whisper model. Download and load a suitable Whisper pre-trained model (such as the medium version) to ensure that the model can efficiently process speech data in multiple languages ​​and scenarios.

[0033] Preprocess the voice data. Perform preprocessing such as format conversion and audio normalization on the collected voice data to ensure that the data meets the input requirements of the Whisper model, including noise removal, sampling rate adjustment, and volume standardization.

[0034] Extract high-level features. Input the preprocessed audio data into the loaded Whisper model to generate feature representations such as mel-spectrograms to improve the recognition accuracy of the model.

[0035] Text label conversion: Convert the annotated text labels into token IDs to ensure the consistency of text data and audio data, which facilitates joint training of the model.

[0036] Data alignment and padding: Use data padding to unify audio sequences and text sequences of different lengths into a fixed length to ensure the efficiency and stability of the model during training.

[0037] In step S105, an AdLoRA Plus fine-tuning module is added to the Whisper model, and the training data set and the test data set are used for training, and the weight of the Whisper model is combined to obtain a trained optimization model.

[0038] For example, the AdLoRA+ module (i.e., the AdLoRA Plus fine-tuning module) is defined: the model weights are adjusted by introducing a low-rank matrix to reduce the number of parameters and computational overhead, while improving the adaptability and generalization ability of the model.

[0039] Insert AdLoRA+ module: Load the pre-trained Whisper model and embed it into the AdLoRA+ module, ready for fine-tuning. Make sure the model can run on the CPU or GPU and supports local file loading to avoid unnecessary network downloads.

[0040] Set the optimizer and learning rate scheduler: Select the AdamW optimizer and the linear learning rate scheduler to ensure that the model converges smoothly during training.

[0041] Freeze basic parameters. During the fine-tuning phase, freeze the original parameters of the Whisper model and only update the parameters of the AdLoRA+ adapter layer to speed up training and prevent overfitting.

[0042] Monitor key indicators such as the loss function in real time, regularly evaluate the performance of the model on the validation set, and save the best model to ensure the smooth progress of the training process.

[0043] Merge the adapter layer with the base model. After fine-tuning, merge the AdLoRA+ adapter layer with the base Whisper model to generate a complete optimized model, simplifying the structure and improving the inference speed.

[0044] Save the final model. Save the optimized model and its auxiliary components (such as feature extractors, tokenizers, and processors) to a specified directory to ensure that the model can be easily loaded and applied to the production environment.

[0045] The optimized model is evaluated using the test set, covering scenarios such as order inquiry, delivery status inquiry, refund processing, and complaint processing, to verify the robustness and generalization ability of the model.

[0046] Iterative optimization: Adjust the model’s hyperparameters or retrain based on the evaluation results to ensure that the model achieves optimal performance in all scenarios.

[0047] In step S106, the speech to be recognized is input into the trained optimization model to obtain predicted text data.

[0048] For example, the process of inferring the speech to be recognized by the trained optimization model includes core steps such as audio data preprocessing, feature extraction, encoder processing, decoder generation, and text post-processing.

[0049] In a specific embodiment, a fine-tuning method based on the AdLoRA+ and Whisper models of the present application is described in detail below with reference to examples and drawings.

[0050] This application relates to the field of speech recognition, and aims to improve the accuracy and robustness of speech recognition in low-resource languages ​​(such as Indonesian) and specific fields (such as the food delivery industry). Through the AdLoRA+ fine-tuning method, this application uses a small amount of domain data to optimize the Whisper model, combined with low-rank adapters and multi-learning rate adjustment strategies, to give full play to the advantages of large-scale pre-trained models, so that the model can still achieve high-precision speech recognition in a low-resource environment. Especially in the food delivery customer service scenario, the optimized model can accurately recognize voice data in different dialects, accents and background noise, thereby improving service quality. By freezing the basic model parameters and updating only the adapter layer, this application speeds up the fine-tuning process, prevents overfitting, and ensures the stability and generalization ability of the model.

[0051] like Figure 2 As shown, Figure 2 The specific steps of collecting domain-specific speech data include: Clarify target application scenarios and data standards.

[0052] Identify the target application scenarios (such as Indonesian takeaway customer service voice recognition) and determine the standards for the required voice data: sampling rate 16kHz, mono, WAV format, duration 10 to 30 seconds, covering common customer service scenario conversations such as order inquiries, delivery status, refund processing, and complaints.

[0053] Select the channel to obtain data.

[0054] Cooperate with food delivery platforms or customer service companies to obtain anonymized real conversation data to ensure practicality and representativeness; use public voice datasets (such as Mozilla Common Voice or OpenSLR) to screen Indonesian data that meets the requirements and increase data volume and diversity.

[0055] Data normalization processing.

[0056] Use Python's librosa or ffmpeg library to unify the audio format, remove samples with severe distortion or excessive noise, and manually check to ensure the audio is clear. Cover different regional accents, gender differences, and a variety of noise backgrounds. Store by scene category for easy subsequent annotation and processing.

[0057] Figure 2 The specific steps of using the Whisper model to perform preliminary transcription of the collected voice data include: Prepare the Whisper model and choose the appropriate version (such as base, small, or medium) to balance accuracy and efficiency. Use Python's whisper library to load the model and ensure that necessary tools such as ffmpeg are installed. Convert the audio files to 16kHz, mono WAV format and transcribe them one by one using the Whisper model to get the text results. Select multi-language support or specific language parameters as needed. After transcription, associate the text with the original audio and check the consistency of the text and audio to provide a basis for data annotation.

[0058] Figure 2 The large language model is used to process the transcribed text and filter out the dialogue texts in common scenarios. The specific steps include: Import and build tips.

[0059] Import the resulting transcribed text into the interactive interface of a large language model (such as GPT-4 or similar models). In order to process these texts efficiently, they can be fed into the model in batches. The input includes the original transcribed text and a prompt that clearly tells the model the task, for example: "Please analyze the following text to determine whether it belongs to a common scenario conversation and extract the important content of the conversation. Common scenarios include takeaway order confirmation, delivery time discussion, and complaint handling, etc., and output concise and clear results." Batch processing and result acquisition.

[0060] In operation, all texts are passed to the large language model for processing by building a loop script or calling the model's API. The model returns structured data containing the screening results, such as texts marked as "common scenarios" and specific scenario type labels. If some text cannot be classified as common scenarios, it can be marked as "other" or stored for subsequent analysis.

[0061] Figure 2 The conversation texts of common scenarios collected in the transcriber are accurately transcribed into labeled speech data. The specific steps include: Text preparation.

[0062] Prepare the dialogue texts that have been screened and annotated for common scenarios, and ensure that each text has been annotated with relevant scenario labels (such as takeaway order confirmation, delivery time discussion, complaint handling, etc.). These texts will serve as the basis for recording, ensuring that each voice data is highly relevant to a specific scenario.

[0063] Precise recording by professionals.

[0064] During the recording process, ensure that the voice recording of each conversation meets certain quality standards to avoid background noise or unclear pronunciation affecting the data quality. In order to ensure the accuracy of the voice, the pronunciation needs to be strictly in accordance with the text content during recording to ensure high consistency with the transcribed text. The duration of each recording should be recorded and associated with the voice data. In this step, professional recording equipment is used for clear voice collection, and the duration of each recording is automatically recorded by software.

[0065] In addition, in order to ensure the timing and semantic consistency of the conversation, the recording staff should try to simulate the conversation rhythm of the real scene when recording, and ensure that the duration of each voice data is consistent with the actual duration of the conversation. During the recording process, the recording software will automatically generate the start and end timestamps of each recording to accurately record the duration of the recording.

[0066] Data organization and storage.

[0067] Ultimately, all recorded speech data will be organized into labeled speech datasets together with corresponding text and scene labels. These data will be stored in a structured format (such as WAV audio files and corresponding text labels) and provide high-quality training data for subsequent feature extraction and model fine-tuning.

[0068] Figure 2 Load the high-quality dataset in . The specific steps include: Data loading.

[0069] When loading the dataset, use the custom CustomDataset class, which reads the file containing the speech data path and matches it according to the annotation file. The initialization of the CustomDataset class requires passing in multiple parameters, such as the speech data file path, text processor (such as WhisperProcessor), audio length and language settings, etc.; if necessary, data enhancement configuration can also be performed. With these parameters, the dataset is automatically processed and ready for training.

[0070] Data segmentation and filling.

[0071] During the loading process, ensure that the speech data is segmented and padded as required, usually by using a data filler (such as DataCollatorSpeechSeq2SeqWithPadding). The filler ensures that speech data of different lengths can be unified in size during batch processing to avoid errors caused by inconsistent lengths. The dataset will be filtered according to the preset minimum and maximum audio lengths to remove samples that do not meet the requirements to ensure data quality.

[0072] Data allocation for multi-card training.

[0073] When training on multiple GPUs, CustomDataset will automatically handle data distribution based on the training configuration (such as local_rank) to ensure that each device is loaded into the corresponding data shard. This step ensures the efficiency and stability of distributed training, fully utilizes multi-GPU resources, and accelerates the training process.

[0074] Passed to the trainer.

[0075] The loaded training and test datasets will be passed to the Seq2SeqTrainer in the training process for model training and evaluation. These high-quality data lay the foundation for subsequent model fine-tuning and performance improvement, ensuring that the model can perform well in practical applications.

[0076] Figure 2 Preprocess the loaded data set in . The specific steps include: Load the processor and perform data preprocessing.

[0077] First, load the processor from WhisperProcessor to convert the audio data and text labels into an input format acceptable to the model. The audio data will be converted to a 16kHz sampling rate and the text labels will be tokenized by the tokenizer of WhisperProcessor. If the timestamp function is enabled, the processor will automatically extract and add the timestamp information in the audio data.

[0078] During preprocessing, the CustomDataset class is responsible for processing each sample one by one to ensure that the audio data conforms to the model input format. Audio samples that do not meet the required length will be automatically filtered. If data enhancement configuration is provided, the audio data will be enhanced, such as adding background noise or audio cropping, to improve the generalization ability of the model.

[0079] Feature extraction.

[0080] In addition, the preprocessing step also includes extracting high-level features of the audio, such as Mel-spectrogram and MFCC (Mel-frequency cepstral coefficients), and standardizing the extracted features to ensure that they meet the input requirements of the model. These high-level features help the model better understand the content of the speech signal.

[0081] Data filling and batch organization.

[0082] The processed data will be padded by DataCollatorSpeechSeq2SeqWithPadding class to ensure that the audio and text lengths of each batch are consistent. During the batch sorting process, any samples that do not meet the standard length will be padded to a uniform length to ensure efficient training.

[0083] Figure 2 Load the pre-trained Whisper model in . The specific steps include: Load the base model.

[0084] In the specific implementation process, first load the Whisper model according to the specified basic model path (such as openai / whisper-large). Use the WhisperFor Conditional Generation.from_pretrained() method to load the model and set relevant configurations according to requirements, such as whether to use 8-bit quantization (use_8bit) and whether to load local model files (local_files_only).

[0085] Configuring task type and language When loading a model, you can automatically configure the model to suit a specific task. WhisperFor ConditionalGeneration supports specifying the task type through the task parameter (for example, transcribe for transcription, translate for translation). In addition, you can also specify the language used during training through the language parameter to ensure that the model can correctly recognize and generate text in the corresponding language when processing input audio.

[0086] Decoder settings.

[0087] After the model is loaded, further configuration adjustments can be made to the model, such as setting forced_decoder_ids or suppress_tokens to ensure that the decoding process meets the task requirements. If fine-tuning or enhancement is required, the adapter layer (AdLoRA+) can be loaded for adjustment to further improve the model's performance. This will be explained in detail in the following steps.

[0088] Figure 2In the original model (Whisper model), the AdLoRA+ low-rank matrix is ​​added. The core idea is to enhance the adaptability of the Whisper model in specific low-resource speech recognition tasks by introducing low-rank matrices, especially when dealing with low-resource languages ​​such as Indonesian, which significantly improves the performance of the model. In the specific implementation, the AdLoRA+ method dynamically adjusts the rank of the low-rank matrix and adopts different learning rate adjustment mechanisms (such as loraplus_lr_ratio and loraplus_lr_embedding) to enable the adapter to accurately optimize specific tasks, thereby improving task performance. The specific steps include: Low-rank matrix is ​​introduced.

[0089] The AdLoRA+ adapter is inserted into key layers of the Whisper model, specifically between the self-attention mechanism and the feedforward network of the Transformer module. By adding low-rank matrix layers to these modules, AdLoRA+ is able to inject task-specific information and improve the model's accuracy on low-resource speech data without changing most of the model's parameters. The introduction of low-rank matrices allows the model to flexibly adapt to the needs of new tasks while maintaining pre-trained knowledge.

[0090] Adjustment of learning rate.

[0091] The learning rate adjustment mechanism of AdLoRA+ is implemented by two main parameters: loraplus_lr_ratio and loraplus_lr_embedding.

[0092] loraplus_lr_ratio controls the learning rate ratio of the low-rank matrices A and B in the LoRA adapter. Matrix A is used to inject task-specific features, while matrix B is used to enhance the representation of these features. By adjusting the learning rates of A and B, loraplus_lr_ratio makes the learning rate of matrix B larger and able to quickly adapt to task requirements, while matrix A maintains the representation of the pre-trained model at a lower learning rate. This ratio is usually set at 2 3 To 2 4 To ensure the reasonable distribution of learning rates of different matrices and improve the adaptability and stability of the model.

[0093] loraplus_lr_embedding is used to control the learning rate of the embedding layer. The embedding layer contains a large number of parameters and is crucial for natural language processing tasks. During fine-tuning, in order to avoid drastic changes in the parameters of the embedding layer, loraplus_lr_embedding is usually set to a smaller value (such as 1e-6) to ensure its stability and avoid hindering the model's learning of task-specific features. Based on these parameters, the optimizer divides the parameters of the model into multiple groups and assigns a different learning rate to each group. For example, the parameters of the main model use the basic learning rate (lr), the matrices A and B in the LoRA adapter use the adjusted learning rate, and the embedding layer uses a specially set learning rate. In this way, the model flexibly adjusts the learning rate of each module while maintaining pre-training capabilities, avoiding overfitting or underfitting, and improving the performance of low-resource speech recognition tasks.

[0094] Dynamically adjust the rank.

[0095] RankAllocator is an innovation in the AdLoRA+ method, which is mainly used to dynamically adjust the rank of low-rank matrices. This mechanism determines the rank of low-rank matrices in real time based on training error, gradient changes, and task complexity. When the training error is high, RankAllocator increases the rank of the low-rank matrix to improve the expressiveness of the model; when the error tends to be stable, the rank is reduced to reduce computational overhead and prevent overfitting. When the gradient of the low-rank matrix changes greatly, RankAllocator increases the rank to help the model better capture complex features; when the gradient changes smoothly, the rank is reduced to avoid waste of resources. RankAllocator also adjusts the rank according to the complexity of the task. In simple tasks, the rank is reduced to prevent overfitting; in complex tasks, the rank is appropriately increased to ensure sufficient expressiveness.

[0096] In the early stages of training, the model needs more capacity to learn the complexity of the task, so RankAllocator will increase the rank of the low-rank matrix to help the model converge quickly. As training progresses, the model gradually approaches convergence, and RankAllocator will gradually reduce the rank of the low-rank matrix to improve computational efficiency and avoid overfitting.

[0097] The dynamic adjustment mechanism of RankAllocator works together with the learning rate adjustment strategy to ensure that the rank and learning rate of the low-rank matrix can be flexibly adjusted according to the task requirements. Through this collaborative optimization, the AdLoRA+ method can effectively improve the performance in low-resource speech recognition tasks, especially in the processing of low-resource languages ​​such as Indonesian, taking full advantage of the large-scale pre-trained model.

[0098] The Whisper model is an end-to-end speech recognition model based on the Transformer architecture. Its core components include encoders and decoders, which interact with each other to extract features from audio signals and generate text output. Whisper's encoder uses a multi-layer self-attention mechanism, in which each layer contains multi-head self-attention, feedforward networks, residual connections, and layer normalization. The decoder relies on a similar structure to the encoder to generate language output, and the role of the attention mechanism is very critical in the model.

[0099] In the Transformer architecture, such as Figure 3 The following is a schematic diagram of the attention mechanism. The core self-attention mechanism (Self-Attention) calculates the relationship between different positions in the input sequence through three representations: query, key, and value. Specifically, each input unit participates in the calculation of the self-attention weight by generating the corresponding query, key, and value. In this process, the query layer (Q), key layer (K), and value layer (V) are combined with the representation of the input data through their respective linear transformations to calculate the attention weight for each input. Finally, the representation of each input is obtained by weighted average.

[0100] In the Whisper model, the Q, K, and V layers are usually generated from the input features through linear transformations in each layer of the multi-head self-attention mechanism. These transformations are completed through the parameter matrix learned through training. Usually Q, K, and V are mapped through different matrices, and their sizes determine the expressiveness and computational complexity of the model. In each self-attention layer, the Q, K, and V vectors are used to calculate the self-attention score, which determines how each position pays attention to the information at other positions. Specifically, Q and K are calculated through dot product operations (i.e., attention scores), and then normalized by Softmax to generate weighted coefficients, which are then multiplied with the V vector and finally merged into the output.

[0101] When fine-tuning the Whisper model, the AdLoRAPlus method introduces a low-rank matrix adjustment strategy to effectively reduce the scale of parameter updates while retaining sufficient learning ability. During the fine-tuning process of Whisper, AdLoRAPlus optimizes the model by only adjusting the adapter layer in the network without directly changing the original weights. This method can significantly reduce training time and storage requirements, especially in low-resource environments.

[0102] Specifically for the fine-tuning of Whisper, the main location where AdLoRAPlus is inserted is the Q and V matrices in the self-attention mechanism. The traditional approach is to directly adjust the weight parameters of these matrices, while AdLoRAPlus performs fine-tuning by inserting low-rank adapter layers into these matrices. Specifically, AdLoRAPlus adds a low-rank matrix after the Q and V matrices in each self-attention layer. The advantage of this is that through the training of the low-rank matrix, the adaptive adjustment of the model can be achieved without significantly increasing the computational overhead. The learning ability of the low-rank matrix is ​​sufficient to capture the specific patterns required for different tasks, while avoiding the need for large-scale updates to the entire model.

[0103] In terms of implementation, the adapter layer of AdLoRAPlus is implemented by adding a pair of matrices (usually W A and W B ). For each matrix in Q and V, AdLoRAPlus adds a product of a low-rank matrix on the basis of its original weight matrix. This low-rank matrix does not directly affect the parameters of the original model. During training, only the parameters of the adapter layer are updated, while the original Q and V matrices remain unchanged. In this way, AdLoRAPlus can fine-tune specific data sets with a small number of parameters to achieve better results.

[0104] In summary, the AdLoRA Plus method avoids the need to fully train the entire Whisper model by inserting low-rank matrices into the Q and V matrices, and only adjusts the parameters of the adapter layer. This method effectively reduces the computational complexity while still retaining sufficient flexibility to adapt to new task requirements.

[0105] Figure 2 Freeze the original parameters of the Whisper model in . During the specific implementation process, first ensure that the original parameters in the Whisper model are frozen to prevent them from being updated during the training process. These frozen parameters usually include pre-trained parameters such as the encoder and decoder weights, position embeddings, etc. By setting the requires_grad attribute of these parameters to False, their gradient calculation and update can be prevented, thereby avoiding modification. This can effectively reduce video memory usage and computational overhead, and focus computing resources on the parts that need to be optimized. After freezing the original parameters, ensure that the parameters of the newly added AdLoRA+ module are in a trainable state, allowing them to be updated during training. Finally, check the model to confirm that the original parameters have been frozen and the parameters of the newly added modules are still trainable.

[0106] Figure 2 Model training. Specifically includes the following steps: Hyperparameter configuration.

[0107] First, the hyperparameters are set according to the configuration, including learning rate, batch size, number of gradient accumulation steps, etc. During training, the model will save checkpoints and perform regular evaluations according to the set number of steps to monitor the performance of the training process.

[0108] Enable half-precision training and compiler optimizations.

[0109] In order to speed up training and reduce video memory consumption, half-precision training (fp16) is usually enabled, which can reduce memory usage and improve computational efficiency. If fp16 is set to True, the Pytorch 2.0 compiler will also be used, which will help further improve training efficiency, especially when dealing with large-scale models.

[0110] Save and evaluate regularly.

[0111] At certain intervals, the training process saves the state of the current model and evaluates the model performance on the validation set. If the model performs well in the evaluation, the optimal model is saved to ensure that the best training results are retained.

[0112] During the training process, Seq2SeqTrainer will use the provided training set and evaluation set for training, and periodically check the performance of the model according to the set number of evaluation steps. As the training progresses, the model will continuously adjust parameters, learn task-specific features, and gradually optimize its weights. To ensure the stability of the training, the logs during training will be output to tensorboard, which is convenient for subsequent viewing of training progress and debugging.

[0113] After training, the model will be restored or continue training based on the saved checkpoint, and the final state of the model and the optimal model will be saved in the specified output directory for subsequent use or further tuning.

[0114] Figure 2 The trained AdLoRA Plus model is merged with the original model weights. The specific steps include: Load the fine-tuned AdLoRA Plus model.

[0115] Check if the AdLoRA Plus model path exists and load the AdLoRA Plus model configuration parameters from that path. Then, use the base model path obtained from the configuration to load Whisper's base model and load it into the device (such as CPU or GPU). If the configuration specifies that the model should only be loaded locally, set local_files_only=True.

[0116] Merge the AdLoRA Plus model with the base model.

[0117] Use the merge_and_unload() method to merge the parameters of the AdLoRA Plus model with the weights of the base model into an overall model and clear the LoRA module itself to reduce memory usage. This step ensures that the parameters of the LoRA module are correctly merged into the base model to form the final optimized model.

[0118] Sets the evaluation mode.

[0119] After the merge is complete, the model is set to evaluation mode using the model.eval() method to ensure that it does not undergo any further training after the merge. This step ensures that the model is in the best state during inference and avoids unnecessary gradient calculations.

[0120] Save the merged model.

[0121] Save the combined model, feature extractor, tokenizer, and processor to the specified directory. The save directory is dynamically generated based on the base model path in the LoRA configuration. If the base model path ends with a slash, it will be stripped. The saved folder name will include the base model name and a "finetune" suffix to indicate that this is a fine-tuned model. To ensure that large models can be stored correctly, the save operation will split the file into 4GB blocks.

[0122] Evaluate the combined model, After the model merge is complete, the final model is evaluated to verify its performance. The evaluation process includes loading the merged model and processor, switching it to evaluation mode, and performing inference on the test set. By comparing the model's prediction results with the true labels of the test set, evaluation metrics (such as character error rate CER or word error rate WER, etc.) are calculated to measure the performance of the model in the actual task. The evaluation results will help confirm whether the merged model retains the ability to fine-tune the model while maintaining the generalization performance of the base model. This step is critical to ensuring model quality and verifying the correctness of the merge operation.

[0123] In a specific embodiment, Figure 4 As shown, this is the inference flowchart of the trained optimization model.

[0124] The inference audio data flow of the Whisper model includes core steps such as audio data preprocessing, feature extraction, encoder processing, decoder generation, and text post-processing. Specifically, it includes the following steps: Audio data preprocessing.

[0125] Specifically, the raw audio data input needs to be uniformly preprocessed. Specifically, all audio is resampled to 16kHz to meet the model input requirements, and denoising is performed to reduce the interference of background noise on the model. For longer audio data, it is divided into segments of fixed length (such as 30 seconds) to facilitate the model to process segment by segment while maintaining semantic continuity. In addition, the audio is amplitude normalized to keep its signal strength consistent, thereby enhancing the inference stability of the model.

[0126] Feature extraction.

[0127] After the speech data is preprocessed, the audio is converted into a logarithmic Mel spectrogram for feature extraction.

[0128] Specifically, after the audio preprocessing is completed, the audio feature extraction operation is performed. Through the short-time Fourier transform (STFT), the one-dimensional audio waveform data is converted into a two-dimensional spectrogram, and the Mel filter is further used to generate the Mel spectrogram. If the robustness of the feature needs to be enhanced, data enhancement techniques such as time masking and frequency masking can be used on the Mel spectrogram to simulate the voice changes in real scenes. On this basis, the convolution layer is used to reduce the dimension of the Mel spectrum, extract local features in time and frequency, and generate a high-dimensional dense feature representation that meets the model input requirements.

[0129] Encoder processing.

[0130] The audio features are then input into the Transformer encoder for processing. In this process, position encoding is first added to the features so that the model can capture the time step information in the audio sequence. The audio features pass through multiple layers of Transformer encoders, each of which contains a multi-head self-attention module, a residual connection and normalization module, and a multi-layer perceptron module. The multi-head self-attention module captures long-distance dependencies by calculating the global correlation of features, while the residual connection and normalization module ensures the stability of the feature distribution. The multi-layer perceptron module further learns the high-order nonlinear relationships of the features. Finally, the encoder generates a high-dimensional representation that contains the global semantic information of the speech features, providing contextual information for the decoder.

[0131] Decoder generation.

[0132] In the decoding stage, the Transformer decoder receives both the audio feature representation and historical text data from the encoder as input. The historical text data includes the text tags generated at the previous moment, and is semantically fused with the audio features through the cross-attention mechanism. The decoder first adds a learnable position encoding to the input data to model the position information of the text sequence. The multi-layer decoder consists of a self-attention module, a cross-attention module, and a multi-layer perceptron module. The self-attention module is used to capture the global correlation of the decoder input and model the dependency between historically generated texts; the cross-attention module combines the information of the encoder output and the decoder input to align the audio features with the text semantics through the cross-attention mechanism. The final output of the decoder is a probability distribution, which is used to represent the tag sequence of the current time step.

[0133] Text post-processing.

[0134] The token sequence generated by the decoder is restored to the complete text through a post-processing step.

[0135] First, the token sequence is decoded into text and restored to human-readable sentences using a predefined vocabulary or subword decomposition method (such as Byte Pair Encoding, BPE). For multiple paragraphs of text generated from long audio, they are concatenated according to chronological order, and punctuation and format corrections are added. In addition, grammar and spelling check tools are used to verify the text results and correct possible spelling or grammatical errors. Finally, the generated text results are formatted into the required output form (such as JSON or plain text) and saved to the specified location, while supporting real-time display or export. This complete process gradually converts audio data from raw signals into high-precision text representations and ensures the accuracy and usability of the generated results.

[0136] Through the above-mentioned low-resource language adaptive speech recognition method based on AdLoRA Plus, by introducing AdLoRA+ fine-tuning technology, the speech recognition accuracy of the Whisper model is significantly improved in a low-resource language environment, especially in the Indonesian takeaway customer service scenario, showing better robustness and accuracy. Traditional models face the problems of insufficient data and poor generalization ability when dealing with low-resource languages, while this application achieves efficient parameter updates with limited training data through low-rank adapter layers and dynamic learning rate adjustment. Adopting the strategy of freezing basic parameters, only the AdLoRA+ adapter layer is trained, which significantly reduces the computational overhead and training time, while maintaining pre-training knowledge and preventing overfitting. This method is not only suitable for Indonesian takeaway customer service, but also has wide adaptability and scalability, ensuring excellent performance in practical applications. In summary, this application effectively improves speech recognition accuracy, reasoning efficiency and training efficiency, solves many challenges of existing technologies in small languages ​​and specific field applications, and has important practical application value.

[0137] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification.

[0138] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any modification, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the appended claims.

Claims

1. A low-resource language adaptive speech recognition method based on AdLoRA Plus, characterized in that: The method includes: Acquire voice data, and perform preliminary transcription on the voice data using a Whisper model to obtain a text data set; Processing the text dataset using a large language model to filter out dialogue texts in common scenarios; Accurately transcribing the common scene dialogue text to obtain transcribed voice data, and combining the common scene dialogue text to obtain a labeled voice dataset; Preprocessing the labeled speech data set, and dividing the preprocessed labeled speech data set into a training data set and a test data set according to a preset ratio; Adding the AdLoRA Plus fine-tuning module to the Whisper model, training using the training data set and the test data set, and combining the weights of the Whisper model to obtain a trained optimization model; The speech to be recognized is input into the trained optimization model to obtain predicted text data.

2. The low-resource language adaptive speech recognition method based on AdLoRA Plus according to claim 1, characterized in that: The step of obtaining voice data and preliminarily transcribing the voice data using a Whisper model to obtain a text data set includes: Determining a target application scenario, data standard, and acquisition path of the voice data to obtain the voice data, and performing normalization processing on the voice data; Performing preliminary transcription of the speech data using the pre-trained Whisper model to obtain the text data set; The text data set is associated with the voice data, and the consistency between the text data set and the voice data is checked.

3. The low-resource language adaptive speech recognition method based on AdLoRA Plus according to claim 1, characterized in that: The step of processing the text dataset using a large language model to filter out dialogue texts in common scenarios includes: Inputting the text data set and the directive prompt into the large language model for analysis; The large language model outputs the common scenario dialogue text.

4. The low-resource language adaptive speech recognition method based on AdLoRA Plus according to claim 1, characterized in that: The steps of accurately transcribing the common scene dialogue text to obtain transcribed voice data, and combining the common scene dialogue text to obtain a labeled voice data set include: Recording each of the common scenario dialogue texts and recording the duration of each recording to obtain the transcribed voice data; All the commonly used scene dialogue texts and their corresponding transcribed voice data are sorted to obtain the labeled voice data set.

5. The low-resource language adaptive speech recognition method based on AdLoRA Plus according to claim 1, characterized in that: The step of preprocessing the labeled speech data set and dividing the preprocessed labeled speech data set into a training data set and a test data set according to a preset ratio includes: Using a data filler to perform data segmentation and filling on the labeled speech dataset, and performing data allocation for multi-card training on the labeled speech dataset; Processing the labeled speech data set into a preset input format, performing feature extraction on the labeled speech data set, and standardizing the extracted features to obtain standardized features; The standardized features are filled with data and batch sorted, and divided into the training data set and the test data set according to a preset ratio.

6. The low-resource language adaptive speech recognition method based on AdLoRA Plus according to claim 1, characterized in that: The step of adding the AdLoRA Plus fine-tuning module to the Whisper model, training using the training data set and the test data set, and combining the weights of the Whisper model to obtain a trained optimization model includes: Configure the task type and language of the Whisper model, and set the forced_decoder_ids or suppress_tokens of the Whisper model; wherein the Whisper model includes a feature extraction module, an encoder, and a decoder; Add the AdLoRA Plus fine-tuning module in the query layer and value layer of the encoder and the query layer and value layer of the decoder; Use loraplus_lr_ratio to control the learning rate ratio of the low-rank matrix A and the low-rank matrix B in the AdLoRA Plus fine-tuning module, use loraplus_lr_embedding to control the learning rate of the embedding layer, and use RankAllocator to dynamically adjust the ranks of the low-rank matrix A and the low-rank matrix B; Freezing the original parameters of the Whisper model; Training the Whisper model using the training data set to obtain an AdLoRA Plus model; Merging the weights of the AdLoRA Plus model and the Whisper model to obtain an optimized model; The optimization model is tested using the test data set to obtain a trained optimization model.

7. The low-resource language adaptive speech recognition method based on AdLoRA Plus according to claim 6, characterized in that: The step of training the Whisper model using the training data set to obtain the AdLoRA Plus model includes: Setting hyperparameters according to the configuration; wherein the hyperparameters include at least learning rate, batch size and gradient accumulation steps; Enable half-precision training; if fp16 is set to True, the Whisper model is optimized using the compiler; Use the training data set to train, check the performance of the model regularly according to the set number of evaluation steps, and output the log to tensorboard; Resume or continue training based on the saved checkpoint, saving the final state of the Whisper model and the optimal model to the output directory.

8. The low-resource language adaptive speech recognition method based on AdLoRA Plus according to claim 7, characterized in that: The step of merging the weights of the AdLoRA Plus model and the Whisper model to obtain an optimized model includes: Check whether the AdLoRA Plus model path exists. If so, load the configuration parameters of the AdLoRA Plus model from the path. Use the Whisper model path obtained from the configuration to load the Whisper model and load it into the device. If the configuration specifies that the model is only loaded locally, set local_files_only=True. The merge_and_unload() method is used to merge the parameters of the AdLoRA Plus model and the weights of the Whisper model into an overall model, and the AdLoRA Plus model itself is cleared to ensure that the parameters of the AdLoRA Plus model are correctly merged into the Whisper model to form the optimized model; wherein, after the merging is completed, the optimized model is set to evaluation mode.

9. The low-resource language adaptive speech recognition method based on AdLoRA Plus according to claim 8, characterized in that: The step of inputting the speech to be recognized into the trained optimization model to obtain predicted text data includes: Preprocessing the speech to be recognized; The pre-processed speech to be recognized is input into the feature extraction module of the trained optimization model, and after converting the one-dimensional audio waveform data into a two-dimensional spectrogram by short-time Fourier transform, a Mel filter is used to generate a Mel spectrogram; a convolution layer is used to perform dimensionality reduction processing on the Mel spectrogram to extract local features in time and frequency to generate audio features; The audio features are input into the encoder of the trained optimization model, position coding is added to the audio features, and the audio features are passed through a multi-layer encoder to generate global semantic information including speech features; wherein each layer of the encoder comprises a multi-head self-attention module, a residual connection and normalization module, and a multi-layer perceptron module, wherein the multi-head self-attention module is used to capture long-distance dependencies by calculating the global correlation of features; the residual connection and the normalization module are used to ensure the stability of feature distribution; and the multi-layer perceptron module is used to learn high-order nonlinear relationships of features; Input the historical text data and the global semantic information into the multi-layer decoder of the trained optimization model, add learnable position encoding to the historical text data and the global semantic information, model the position information of the text sequence, and output a probability distribution to represent the tag sequence of the current time step; wherein the decoder includes a self-attention module, a cross-attention module and a multi-layer perceptron module, the self-attention module is used to capture the global correlation of the decoder input and model the dependency between historical generated texts; the cross-attention module is used to combine the information of the encoder output and the decoder input, and realize the alignment of audio features and text semantics through the cross-attention mechanism; The tag sequence is decoded into text, restored to human-readable sentences using a predefined vocabulary or subword decomposition method, spliced ​​according to chronological order, punctuation and format correction are added, and the text results are verified using grammar and spelling check tools to generate predicted text data.

Citation Information

Patent Citations

  • Speech recognition pre-training model fine tuning method and system

    CN117275463A

  • Intelligent scoring method for clinical oral tests of resident physicians based on Whisper model

    CN118098212A

  • Log anomaly detection method based on efficient fine tuning of adaptive low-rank parameters

    CN118260689A

  • Multi-language speech recognition model based on multiple experts and training method thereof

    CN118675508A

  • Multi-modal opinion expression recognition system and method driven by large language model

    CN118917407A

Cited By

  • Adaptive traffic domain service voice generation method and system based on thinking chain fine-tuning large model

    CN120126484A

  • Lao acoustic representation method based on Lao pronunciation similarity

    CN120636372A

  • Large model adaptive training method and system based on task semantic perception

    CN120688563A

  • Manual intervention word segmentation method for improving ASR recognition effect

    CN120833785A

  • Ship radiation noise identification method and system based on double low-rank adjustment network

    CN121034340A