Speech recognition method and system for constructing small language based on whispertoken

The dynamic vocabulary and joint training framework for small languages were constructed through Whispertokenizer, which solved the problems of inaccurate vocabulary and low model training efficiency in small language speech recognition, improved the accuracy of speech recognition and model generalization capabilities, and simplified the construction process.

CN120340494AActive Publication Date: 2025-07-18BEIJING RUI KELUN INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510470690.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-18
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The problems of inaccurate construction of vocabulary lists in the prior art, low model training efficiency and poor speech recognition effect in the speech-text alignment task are especially fuzzy, resulting in poor speech recognition effect in the speech-text alignment task.

Method used

The initial candidate set is constructed by Whispertokenizer, high-frequency tokens are filtered through frequency statistics and supplemented with low-frequency tokens, and a dynamic vocabulary is constructed. The model training is carried out by combining encoder, decoder and CTC decoder, and speech recognition is used by Fbank feature extraction and AttentionDecoder. The final text recognition results are generated through the joint optimization of BeamSearch and AttentionDecoder.

Benefits of technology

It significantly improves the accuracy and model generalization ability of small language speech recognition, reduces the word error rate by 8%, reduces the number of model parameters and calculation costs, and simplifies the model construction process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340494A_ABST
    Figure CN120340494A_ABST
Patent Text Reader

Abstract

The invention provides a voice recognition method and system for constructing a small language based on whispertoken, and relates to the technical field of natural language processing and voice recognition, and the method comprises the steps: extracting all tokens related to a target small language in a whispertoken, and forming an initial candidate set; matching and analyzing the tokens in the initial candidate set and the collected training text corpus of the target small language, and counting the occurrence frequency of the tokens in the corpus; and screening high-frequency tokens according to a frequency statistical result, and supplementing low-frequency tokens to construct a dynamic vocabulary. According to the method, the vocabulary quality is improved, the model training efficiency is optimized, the speech recognition accuracy is enhanced, the model generalization ability is improved, and the model construction process is simplified, so that an efficient, accurate and easy-to-implement solution is provided for the field of minority language speech recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of natural language processing (NLP) and speech recognition, and particularly to a speech recognition method and system for minority languages based on whisper tokens. Background Art

[0002] In the existing natural language processing (NLP) technology, the model training of minority languages (languages with scarce resources) faces significant challenges, and the core problem lies in the construction of the vocabulary (vocab). Traditional methods usually rely on statistical word segmentation (such as the BPE algorithm) or rule-based word segmentation, but for minority languages, the following defects exist: 1. Insufficient corpus resources lead to inaccurate word segmentation results, and low-frequency words are easily mismerged or split. 2. The general vocabulary of existing multilingual models (such as mBART, XLM-R) has limited coverage of minority languages and cannot adapt to their language characteristics. 3. In the speech-text alignment task, due to the fuzzy mapping relationship between speech features and text tokens, the speech recognition effect of minority languages is poor. The closest prior art is a multilingual model based on the Whisper speech recognition system, but its native tokenizer still has problems such as insufficient number of minority language tokens and lack of dynamic expansion ability for the speech features of minority languages. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a speech recognition method and system for minority languages based on whisper tokens, so as to solve the problems of inaccurate construction of minority language vocabulary, low model training efficiency, and poor speech recognition effect in the prior art.

[0004] To solve the above technical problems, the technical solution of the present invention is as follows:

[0005] In a first aspect, a speech recognition method for minority languages based on whisper tokens, the method includes:

[0006] Extract all tokens related to the target minority language in the Whisper tokenizer to form an initial candidate set;

[0007] Match and analyze the tokens in the initial candidate set with the collected target minority language training text corpus, and count the occurrence frequency of the tokens in the corpus;

[0008] Screen high-frequency tokens according to the frequency statistics results, and supplement low-frequency tokens to construct a dynamic vocabulary;

[0009] Construct a speech recognition model, the model includes an encoder, a decoder, a CTC decoder, and a CTC loss function, and train the model through the dynamic vocabulary;

[0010] Extract Fbank features from the input speech signal, preprocess the extracted Fbank features, and input them into the Encoder for forward calculation to generate high-level semantic features;

[0011] Perform CTCDecoder processing on the high-level semantic features, and generate multiple candidate results through the BeamSearch strategy;

[0012] Use AttentionDecoder to re-score these candidate results, select the result with the highest score as the preliminary recognition result, and perform post-processing on the preliminary recognition result to obtain the final text recognition result.

[0013] Furthermore, extract all tokens related to the target minority language in Whispertokenizer to form an initial candidate set, including:

[0014] Obtain a pre-trained Whispertokenizer model to recognize the target minority language;

[0015] Traverse all tokens in Whispertokenizer and extract tokens related to the target minority language; integrate all related tokens into a set to form an initial candidate set.

[0016] Furthermore, match and analyze the tokens in the initial candidate set with the collected training text corpus of the target minority language, and count the occurrence frequencies of the tokens in the corpus, including:

[0017] Collect and organize the training text corpus of the target minority language. The corpus format is a plain text file, and the encoding format is UTF-8. The content contains diverse text data in the target minority language;

[0018] Match each token in the initial candidate set with the training text corpus of the target minority language, identify the tokens that appear in the corpus, and perform analysis;

[0019] Count the occurrence frequencies of each token in the training text corpus of the target minority language to obtain the frequency statistics result.

[0020] Furthermore, screen out high-frequency tokens and supplement low-frequency tokens according to the frequency statistics result to construct a dynamic vocabulary, including:

[0021] Set a threshold T according to the dynamic adjustment method of data distribution;

[0022] According to the set threshold, tokens with higher frequencies are selected from the frequency statistics results, and for tokens with lower frequencies, they are supplemented through specific processing methods;

[0023] The selected high-frequency tokens and the processed low-frequency tokens are merged to construct a dynamic vocabulary.

[0024] Furthermore, a speech recognition model is constructed. The model includes an encoder, a decoder, a CTC decoder, and a CTC loss function. The model is trained using the dynamic vocabulary, including:

[0025] Use a deep learning framework to construct a speech recognition model, which includes an encoder, a decoder, a CTC decoder, and a CTC loss function;

[0026] Use a signal processing library to extract features from the speech signal and convert the transcribed text into a token sequence;

[0027] Collect small-language corpora of speech signals and their corresponding texts, perform preprocessing to form a structured dataset, and divide the dataset into a training set, a validation set, and a test set;

[0028] Initialize the various parameters in the model, write code for forward propagation, loss calculation, backpropagation, and parameter update, set the hyperparameter adjustment strategy, and achieve the alignment of speech features and text tokens by jointly optimizing the CTC loss and the Attention decoder loss;

[0029] Use the test set to evaluate the trained model and calculate various performance metrics.

[0030] Furthermore, perform Fbank feature extraction on the input speech signal, preprocess the extracted Fbank features, and input them to the Encoder for forward calculation to generate high-level semantic features, including:

[0031] Perform Fbank feature extraction on the input speech signal to obtain the spectral feature representation of the speech signal, and preprocess the extracted Fbank features;

[0032] Input the preprocessed Fbank features into the Encoder. The Encoder performs forward calculation to convert the input Fbank features into a high-level semantic feature representation.

[0033] Furthermore, perform CTCDecoder processing on the high-level semantic features, and generate multiple candidate results through the BeamSearch strategy, including:

[0034] Input the high-level semantic features into the CTCDecoder;

[0035] The CTCDecoder processes the high-level semantic features and generates multiple candidate results through the BeamSearch strategy.

[0036] Furthermore, the AttentionDecoder is used to re-score these candidate results, select the result with the highest score as the preliminary recognition result, and post-process the preliminary recognition result to obtain the final text recognition result, including:

[0037] Use the AttentionDecoder to re-score the multiple candidate results generated by the CTCDecoder;

[0038] According to the re-scoring results, select the candidate result with the highest score as the preliminary recognition result;

[0039] Post-process the preliminary recognition result to obtain the final text recognition result.

[0040] In a second aspect, a speech recognition system for a minority language based on whispertoken includes:

[0041] An acquisition module, configured to extract all tokens related to the target minority language in the Whispertokenizer to form an initial candidate set; match and analyze the tokens in the initial candidate set with the collected target minority language training text corpus, and count the occurrence frequency of the tokens in the corpus;

[0042] A construction module, configured to screen high-frequency tokens and supplement low-frequency tokens according to the frequency statistics results to construct a dynamic vocabulary; construct a speech recognition model, the model includes an encoder, a decoder, a CTC decoder, and a CTC loss function, and train the model through the dynamic vocabulary;

[0043] A processing module, configured to extract Fbank features from the input speech signal, preprocess the extracted Fbank features, input them to the Encoder for forward calculation to generate high-level semantic features; perform CTCDecoder processing on the high-level semantic features, and generate multiple candidate results through the BeamSearch strategy; use the AttentionDecoder to re-score these candidate results, select the result with the highest score as the preliminary recognition result, and post-process the preliminary recognition result to obtain the final text recognition result.

[0044] In a third aspect, a computing device includes:

[0045] One or more processors;

[0046] A storage device for storing one or more programs which, when executed by one or more processors, cause the one or more processors to implement the method.

[0047] In a fourth aspect, a computer-readable storage medium stores a program which, when executed by a processor, implements the method.

[0048] The above solution of the present invention has at least the following beneficial effects:

[0049] By reusing the Whisper tokenizer, the method of the present invention makes full use of its multilingual phoneme-text mapping relationship, reduces the dependence on manual rules, and reduces the error rate of token splitting for minority languages by 30%-50% compared with the traditional BPE method, significantly improving the accuracy and reliability of the vocabulary. A dynamic vocabulary is constructed using a dynamic frequency screening strategy, reducing redundant tokens, reducing the number of model parameters, and improving the convergence speed and training efficiency. Combining a speech-text joint training framework enhances the adaptability of the decoder to speech features, and on test sets of low-resource languages such as Tibetan and Vietnamese, the word error rate (WER) is reduced by an average of 8%, improving the speech recognition accuracy. The reuse and extension mechanism of the Whisper tokenizer enables the model to flexibly adapt to different minority language speech features, improving the generalization ability. The method of the present invention provides an efficient, accurate and easily implementable solution for minority language speech recognition by improving the vocabulary quality, optimizing the training efficiency, enhancing the accuracy, improving the generalization ability and simplifying the construction process. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is a schematic flowchart of a method for constructing a speech recognition method for minority languages based on Whisper token provided by an embodiment of the present invention.

[0051] Figure 2 is a schematic diagram of a speech recognition system for minority languages based on Whisper token provided by an embodiment of the present invention.

[0052] Figure 3 is a structural diagram of an encoder of a speech recognition system for minority languages based on Whisper token provided by an embodiment of the present invention.

[0053] Figure 4 is a structural diagram of a convolutional module of a speech recognition system for minority languages based on Whisper token provided by an embodiment of the present invention.

[0054] Figure 5 is a structural diagram of a feed-forward module of a speech recognition system for minority languages based on Whisper token provided by an embodiment of the present invention.

[0055] Figure 6 It is a structural diagram of a training set division model for a speech recognition system for minority languages constructed based on whispertoken provided by an embodiment of the present invention. Specific implementation manners

[0056] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0057] As Figure 1 shown, an embodiment of the present invention proposes a speech recognition method for minority languages constructed based on whispertoken, and the method includes the following steps:

[0058] Step 11: Extract all tokens related to the target minority language in the Whispertokenizer to form an initial candidate set;

[0059] Step 12: Match and analyze the tokens in the initial candidate set with the collected training text corpus of the target minority language, and count the occurrence frequencies of the tokens in the corpus;

[0060] Step 13: Screen high-frequency tokens according to the frequency statistics results and supplement low-frequency tokens to construct a dynamic vocabulary;

[0061] Step 14: Construct a speech recognition model, the model includes an encoder, a decoder, a CTC decoder and a CTC loss function, and train the model through the dynamic vocabulary;

[0062] Step 15: Extract Fbank features from the input speech signal, preprocess the extracted Fbank features, and input them to the Encoder for forward calculation to generate high-level semantic features;

[0063] Step 16: Perform CTCDecoder processing on the high-level semantic features, and generate multiple candidate results through the BeamSearch strategy;

[0064] Step 17: Use the AttentionDecoder to re-score these candidate results, select the result with the highest score as the preliminary recognition result, and post-process the preliminary recognition result to obtain the final text recognition result.

[0065] In the embodiments of the present invention, tokens related to the target minority language are extracted from the Whisper tokenizer to form an initial set, covering unique vocabulary and grammatical structures. By combining corpus frequency to screen high-frequency tokens and supplement low-frequency tokens, the recognition ability of long-tail words is improved, redundant tokens are reduced, the model parameters and computational costs are lowered, and the training efficiency is enhanced. The CTC loss and the Attention decoder loss are jointly optimized, taking into account both the global alignment constraint and the local flexible alignment, effectively handling the monotonic and complex alignment scenarios of minority languages. The dynamic vocabulary is directly used for training, avoiding the problem of phoneme-token mismatch in the general vocabulary, and enhancing the model's learning ability for the semantic and pronunciation features of minority languages. The Fbank features are used to retain the speech spectrum envelope information, better representing the suprasegmental features such as tones and stresses of minority languages, and being more adaptable to the phonological characteristics of minority languages than MFCC. The CTC decoder generates candidate paths, and the Beam Search is combined to improve the recall rate of long speech. The Attention mechanism re-evaluates the candidate results to correct the alignment errors. Post-processing is performed on the grammar rules of minority languages to reduce grammar errors and non-compliant outputs. The full-process optimization enables the model to reduce the word error rate (WER) by 15%-30% compared with the baseline system, significantly improving the recognition accuracy and user experience.

[0066] In a preferred embodiment of the present invention, step 11 may include:

[0067] Step 111, obtaining a pre-trained Whisper tokenizer model to recognize the target minority language;

[0068] Step 112, traversing all tokens in the Whisper tokenizer and extracting tokens related to the target minority language;

[0069] Step 113, integrating all related tokens into a set to form an initial candidate set.

[0070] In the embodiments of the present invention, the target minority language is recognized through a pre-trained Whisper tokenizer model, and related tokens are extracted to ensure that the initial candidate set contains the unique vocabulary and grammatical structures of the minority language. Integrating the related tokens into the initial candidate set avoids using a large and redundant general vocabulary, significantly reducing the number of model parameters and computational costs.

[0071] In the embodiments of the present invention, the specific steps include:

[0072] Step 111, using Python code to load the WhisperTokenizer class in the HuggingFace Transformers library and initializing the tokenizer object by specifying the pre-trained model path.

[0073] Step 112: Preprocess the input training text corpus, including removing special characters, redundant spaces, punctuation marks, etc. in the text, only retaining the text content, splitting the preprocessed text into multiple text segments by sentences or paragraphs to form a text segment list. Traverse the text segment list and use the encode method of the tokenizer to encode each text segment to obtain the corresponding token sequence.

[0074] Step 113: Merge the token sequences of all text segments to form a set containing all the tokens that appear. Initialize an empty dictionary to count the token frequencies. Traverse all the extracted tokens. For each token, if it is already in the dictionary, increment the corresponding value by 1; otherwise, add it to the dictionary and set the value to 1. Finally, obtain a dictionary containing all the tokens and their occurrence frequencies. According to the token frequency dictionary, sort them from high to low frequency and filter out the tokens that meet a specific frequency threshold to form an initial candidate set.

[0075] In a preferred embodiment of the present invention, the above step 12 may include:

[0076] Step 121: Collect and organize the training text corpus of the target minority language. The corpus format is a plain text file with an encoding format of UTF-8, and the content contains diverse text data in the target minority language;

[0077] Step 122: Match each token in the initial candidate set with the training text corpus of the target minority language, identify the tokens that appear in the corpus, and perform analysis;

[0078] Step 123: Count the occurrence frequencies of each token in the training text corpus of the target minority language to obtain the frequency statistics result.

[0079] In the embodiment of the present invention, by collecting the training text corpus of the target minority language, matching the initial candidate tokens, and counting the frequencies, this solution realizes multiple advantages of accurately adapting to language characteristics, improving the quality of the vocabulary, enhancing training reliability, reducing resource costs, and having technical scalability, provides a high-quality data basis for subsequent model training, and significantly improves the performance and efficiency of minority language speech recognition.

[0080] In the embodiment of the present invention, the specific steps include:

[0081] Step 121: Collect the corpus of the target minority language from public databases (such as Wikipedia, United Nations documents), local cultural institutions, social media, books, and spoken language transcriptions. Convert all the corpus into plain text files (.txt) with the encoding format set to UTF-8 to avoid character garbling. Remove irrelevant content (such as HTML tags, advertising texts), and retain the core language data.

[0082] Step 122: Import the initial candidate token set to form a benchmark list for matching. For each token in the initial candidate set, search for its occurrence position sentence by sentence in the corpus, record the matching results, and observe the context of the matching token in the corpus to verify whether its syntactic and semantic functions conform to the rules of the target minority language.

[0083] Step 123: Count the occurrence frequency of each token in the corpus of the target minority language training text to obtain the frequency statistics result.

[0084] In a preferred embodiment of the present invention, the above step 13 may include:

[0085] Step 131: Set a threshold T according to the dynamic adjustment method of data distribution;

[0086] Step 132: Screen out the tokens with higher occurrence frequencies from the frequency statistics result according to the set threshold. For the tokens with lower occurrence frequencies, supplement them through specific processing methods;

[0087] Step 133: Merge the screened high-frequency tokens and the processed low-frequency tokens to construct a dynamic vocabulary.

[0088] In the embodiment of the present invention, by setting the dynamic threshold T, it is ensured that the screened high-frequency tokens cover the core vocabulary of the target minority language, improving the model's understanding ability of common expressions. Targeted processing of low-frequency tokens avoids waste of computing resources caused by excessive expansion of the vocabulary, while retaining cultural or domain-specific vocabulary. By merging similar tokens, the sparsity problem caused by the small vocabulary in the minority language corpus is alleviated, making it easier for the model to learn language rules.

[0089] In the embodiment of the present invention, the specific steps include:

[0090] Step 131, the threshold T is set using a dynamic adjustment method based on data distribution. Calculate the mean (μ) and standard deviation (σ) of all token frequencies. Determine the threshold according to the formula T = μ + kσ, where k is an adjustable coefficient, usually taking values between 0 and 2, and the specific value is determined according to the distribution characteristics of the corpus and the strictness of the definition of high-frequency and low-frequency tokens. For example, when k = 1, T = μ + σ, which means that tokens with frequencies higher than the mean plus one standard deviation are regarded as high-frequency tokens; when k = 0.5, T = μ + 0.5σ, indicating that the screening of high-frequency tokens is relatively more strict.

[0091] Step 132, select tokens with occurrences ≥ T from the frequency statistics results to form a high-frequency token set. For example, if T = 100, then all tokens with frequencies ≥ 100 (such as "the", "and") are screened as high-frequency tokens. Mark tokens with frequencies < T as low-frequency tokens. For example, rare words with frequencies lower than 100 (such as "xyz", "abcde") enter the low-frequency processing process. According to Whisper's sub-word splitting algorithm, decompose low-frequency tokens into smaller sub-units. For example, "abcde" can be split into "a", "bc", "de". Recombine the split sub-tokens according to rules, including: frequency priority, merge adjacent and higher-frequency sub-tokens (such as "frequ" + "ent" → "frequent"); semantic relevance, merge sub-tokens according to semantic relevance (such as "auto" + "mobile" → "automobile"); grammatical structure, combine based on roots or affixes (such as "un" + "happy" → "unhappy").

[0092] Step 133, merge the screened high-frequency token set with the processed low-frequency tokens (including the split sub-tokens and merged compound tokens). For example, the high-frequency token "the" and the processed low-frequency tokens "a", "bc", "de" together form a candidate set. Remove duplicates from the merged token set to ensure no duplicates in the vocabulary. The tokens can be sorted according to frequency or alphabetical order to optimize storage and retrieval efficiency. Save the final set as a dynamic vocabulary to support subsequent model training or text processing tasks. The vocabulary can be iteratively updated according to changes in new corpora or rules to ensure adaptation to different language characteristics or domain requirements.

[0093] In a preferred embodiment of the present invention, the above step 14 may include:

[0094] Step 141: Build a speech recognition model using a deep learning framework. The model includes an encoder, a decoder, a CTC decoder, and a CTC loss function.

[0095] Step 142: Use a signal processing library to extract features from the speech signal and convert the transcribed text into a token sequence.

[0096] Step 143: Collect a small language corpus of speech signals and their corresponding texts, perform preprocessing to form a structured dataset, and divide the dataset into a training set, a validation set, and a test set.

[0097] Step 144: Initialize the various parameters in the model, write code for forward propagation, loss calculation, backpropagation, and parameter update, set the hyperparameter adjustment strategy, and achieve the alignment of speech features and text tokens by jointly optimizing the CTC loss and the Attention decoder loss.

[0098] Step 145: Use the test set to evaluate the trained model and calculate various performance metrics.

[0099] In the embodiment of the present invention, the complete process from data collection to model evaluation forms a "data - model - evaluation" closed loop, supporting rapid iterative optimization and adapting to the challenges of low resources and high variability in small language speech recognition. The customized design for small languages can reduce the dependence of speech recognition technology on large-scale labeled data and promote the popularization of the technology to resource-scarce languages. The standardized steps can be used as a reference for engineering practice and also provide an experimental framework for research directions such as alignment mechanisms and low-resource learning in speech recognition.

[0100] In the embodiment of the present invention, the specific steps include:

[0101] Step 141: Build a speech recognition model using a deep learning framework. The encoder (Encoder) adopts a Conformer structure, which is composed of multiple stacked ConformerBlocks. Each Block contains core components such as a convolution module (ConvolutionModule), a multi-head self-attention mechanism (RelPositionMultiHeadedAttention), and a feed-forward network (FeedForwardModule). The convolution module implements pointwise convolution, a GLU activation layer, 1-D depth convolution, BatchNorm, and a Swish activation layer to enhance the local feature extraction ability.

[0102] The feed-forward module uses a linear layer with an expansion factor of 4, combined with the Swish activation and pre-normalization (Pre-Norm) residual units to enhance the non-linear modeling ability. For downsampling and positional encoding, the input features are downsampled by the Conv2dSubsampling4 module, and positional encoding (PositionalEncoding) is added to preserve the sequence order information. The decoder is based on the Transformer architecture and contains multiple layers of DecoderLayer. Each layer integrates self-attention mechanism, cross-attention mechanism (interacts with the encoder output), and feed-forward network to generate the text sequence. The CTC module integrates the CTC decoder and the CTC loss function, directly performs linear transformation and Softmax calculation on the encoder output to obtain the character probability distribution, and solves the problem of input-output sequence alignment.

[0103] Step 142: Use a signal processing library (such as Librosa, Kaldi) to preprocess the original speech signal, extract features such as Mel-frequency cepstral coefficients (MFCC), filter bank energy (Fbank), etc., to form a speech feature sequence. Normalize or standardize the features to ensure the consistency of the input data distribution. The transcription text corresponding to the speech signal is in the format of a plain text file, encoded as UTF-8. Each line contains the transcription content of a speech file, and this transcription content has been converted into the corresponding token sequence according to the generated dynamic vocabulary, corresponding one-to-one with the speech file.

[0104] Step 143: Collect small-language corpora containing speech signals and corresponding texts, covering different accents, speech rates, and scenarios to ensure data diversity. Prioritize open-source datasets or collect data through crowdsourcing to avoid data bias. Clean, annotate, and organize the data to form a structured dataset, which is divided into a training set, a validation set, and a test set in the ratio of 8:1:1 to ensure the consistency of the data distribution in each subset and avoid model overfitting or evaluation bias.

[0105] Step 144: Use the Xavier or Kaiming initialization method to randomly initialize the weights and biases of the encoder, decoder, and CTC module to avoid gradient vanishing or explosion. The forward pass of the encoder takes the input speech feature sequence and outputs a high-level semantic feature representation. The forward pass of the decoder combines the encoder output and the generated token sequence, and predicts the probability distribution of the next character through self-attention and cross-attention mechanisms. CTC forward: Perform a linear transformation and Softmax on the encoder output to obtain the character probability distribution. The CTC loss calculates the alignment loss between the encoder output and the true token sequence using torch.nn.CTCLoss. The Attention decoder loss is based on the cross-entropy loss (CrossEntropy) to maximize the autoregressive probability, and the TeacherForcing strategy is adopted to accelerate convergence. The total loss weights and sums the CTC loss and the Attention loss (such as 0.5:0.5) to form a joint loss function. Calculate the gradient through the backpropagation algorithm, and use the Adam or SGD optimizer to update the model parameters. Regularly monitor the validation set loss and stop training when the loss is stable or reaches a preset threshold. Dynamically adjust the learning rate according to the training curve (such as learning rate decay or Warmup strategy), and adjust hyperparameters such as batch size and Dropout rate to balance the training speed and the model generalization ability.

[0106] Step 145: Use the test set to evaluate the trained model and calculate various performance metrics.

[0107] In a preferred embodiment of the present invention, the above step 15 may include:

[0108] Step 151: Extract Fbank features from the input speech signal to obtain the spectral feature representation of the speech signal, and preprocess the extracted Fbank features.

[0109] Step 152: Input the preprocessed Fbank features into the Encoder. Through forward calculation, the Encoder converts the input Fbank features into a high-level semantic feature representation.

[0110] In the embodiments of the present invention, Fbank feature extraction and preprocessing provide high-quality input for the Encoder. Its normalization and length alignment operations ensure that the Encoder can stably process different speech samples. The Encoder converts the spectral features into high-level semantic representations through forward calculation, providing a basis for the subsequent decoder to generate text sequences. Through the combination of Fbank feature extraction and Encoder semantic encoding, the model has been significantly improved in terms of robustness, context modeling, low-resource adaptation, training efficiency, and technical scalability, laying a solid foundation for building a high-performance and generalized speech recognition system.

[0111] In the embodiments of the present invention, the specific steps include:

[0112] Step 151: Sample the original voice signal to the standard sampling rate (such as 16 kHz) and convert it to the mono format to ensure the consistency of the input data. Remove environmental noise or low-frequency interference through filtering techniques (such as band-pass filters) to improve the signal clarity. Frame the voice signal with a fixed frame length (such as 25 ms) and frame shift (such as 10 ms), and apply a Hamming window to each frame to reduce spectral leakage. Perform FFT on each frame of the signal to convert the time-domain signal into a frequency-domain signal and obtain a spectrogram. Filter the spectrogram through a Mel-scale triangular filter bank, calculate the energy within each filter, obtain the Mel-frequency energy (Mel-Filterbank Energies), take the logarithm of the filter bank energy to enhance the dynamic range compression ability, and form the final Fbank features. Perform mean-variance normalization (CMVN) on the Fbank features frame-by-frame or globally to eliminate channel or speaker differences. Unify the feature sequence to a fixed length through zero-padding or truncation to adapt to the requirements of the Encoder input. Scale the feature values (such as scaling to the range [-1, 1]) to avoid gradient explosion caused by overly large numerical values. Visualize a partial Fbank feature spectrogram to check whether the spectral energy distribution is reasonable and verify the effectiveness of feature extraction.

[0113] Step 152: Ensure that the dimension of the preprocessed Fbank features is consistent with the configuration of the Encoder input layer (such as the feature dimension is 80 and the sequence length is T). Add positional encoding (Positional Encoding) to the Fbank feature sequence to inject sequence order information and make up for the lack of positional information in the convolutional or self-attention mechanism. Perform downsampling on the Fbank features through a convolutional layer (such as Conv2dSubsampling) to reduce the sequence length while extracting local features. Perform depth convolution on the downsampled features to capture local spectral correlations, calculate the global dependencies between features through multi-head self-attention (Multi-Head Self-Attention), and model long-distance context. Use a feed-forward fully connected layer for further non-linear transformation to enhance the feature expression ability. Apply layer normalization (Layer Normalization) and residual connections after each sub-module to stabilize the training and accelerate convergence. Through the stacking of multiple ConformerBlocks, gradually convert the Fbank features from low-level acoustic features (such as spectral energy) to high-level semantic features (such as phonemes, word-level representations).

[0114] In a preferred embodiment of the present invention, the above step 16 may include:

[0115] Step 161: Input the high-level semantic features into the CTCDecoder;

[0116] Step 162: The CTCDecoder processes the high-level semantic features and generates multiple candidate results through the BeamSearch strategy.

[0117] In the embodiments of the present invention, through the processing of high-level semantic features by the CTCDecoder and the application of the BeamSearch strategy, the model has been significantly improved in terms of decoding efficiency, alignment ability, candidate result quality, real-time performance, and low-resource adaptability, providing key technical support for building an efficient and robust end-to-end speech recognition system.

[0118] In the embodiments of the present invention, the specific steps include:

[0119] Step 161: Confirm that the high-level semantic features output by the Encoder (such as with dimensions TxD, where T is the time step and D is the feature dimension) are compatible with the input layer of the CTCDecoder. If the Encoder output contains redundant frames (such as due to insufficient downsampling), additional downsampling or truncation operations can be performed to ensure that the length of the feature sequence is suitable for processing by the CTCDecoder. Map the high-level semantic features to the output space of the size of the character set (including the blank label) through a fully connected layer (LinearLayer), with the output dimension being Tx(V + 1), where V is the size of the character set. Apply Softmax or Log-Softmax activation to the output of the linear layer to convert the features into a probability distribution, representing the probability of generating each character at each time step. Explicitly add a blank label (Blank) on the basis of the character set to handle consecutive repeated characters or alignment ambiguities, ensuring that the CTCDecoder can correctly model the length difference between the input and output sequences.

[0120] Step 162: Construct a dynamic programming table α[t][s] to store the cumulative probability of generating the first s characters of the character sequence at time step t. Based on the Softmax output probability, recursively calculate the probabilities of all possible paths to form a complete path probability graph. Initialize an empty path set B, set the Beam width K (e.g., K = 5), which defines the number of the highest probability paths to be retained at each time step. Add the initial blank path (containing only the blank label) to the set B with a probability of 1. For each time step t, for each path in the current candidate path set B, attempt to expand by one character (including the blank label) and calculate the probability of the expanded path. Sort all the expanded paths by the cumulative probability and retain the top K paths to form a new candidate set B'. According to the CTC rule, merge consecutive repeated characters (e.g., aa is merged into a) and blank labels. When all time steps are processed, or the highest probability path in the candidate paths reaches a preset threshold, stop the search. Extract the N paths with the highest path probabilities (e.g., N = 3) from the final candidate set B', remove the blank labels and merge the repeated characters to generate multiple candidate results.

[0121] In a preferred embodiment of the present invention, the above step 17 may include:

[0122] Step 171: Use the AttentionDecoder to re-score the multiple candidate results generated by the CTCDecoder;

[0123] Step 172: According to the re-scoring results, select the candidate result with the highest score as the preliminary recognition result;

[0124] Step 173: Post-process the preliminary recognition result to obtain the final text recognition result.

[0125] In the embodiments of the present invention, the AttentionDecoder performs refined scoring on multiple candidate results generated by the CTCDecoder through global context modeling, effectively correcting the errors caused by local ambiguity of CTC (such as similar pronunciations), and significantly improving the recognition accuracy. Its re-scoring mechanism based on multiple candidate paths avoids the limitations of a single path. When there is noise or non-standard pronunciation, the robustness can be enhanced through high-confidence paths. At the same time, the AttentionDecoder implicitly learns the characteristics of the language model, preferentially selects results with correct grammar and coherent semantics, and reduces problems such as character repetition, omission, or unreasonable combinations. By using the self-attention mechanism to capture long-distance dependencies and combining context information to optimize candidate results, more natural text is generated. The post-processing rules can be flexibly adapted to different scenarios, supporting manual or automatic correction. In addition, the re-scoring mechanism decouples the fast decoding of CTC and the fine-grained modeling of Attention, facilitating modular optimization (such as upgrading the Encoder or replacing the Attention structure), and supporting multi-modal feature inputs (such as speech + text + image), further improving the performance in complex scenarios.

[0126] In the embodiments of the present invention, the specific steps include:

[0127] Step 171, from the multiple candidate results output by the CTCDecoder, extract the character sequence of each candidate and its corresponding path probability (such as N candidates, each containing a character sequence and a cumulative probability). Convert the candidate character sequence into a vector representation through the EmbeddingLayer as the input of the AttentionDecoder. If the lengths of the candidate sequences are different, alignment is required through padding or truncation. Use the high-level semantic features (such as TxD) output by the Encoder as the key and value of the AttentionDecoder, and the embedded vector of the candidate sequence as the query, and calculate the context-related feature representation through the multi-head self-attention mechanism. Based on the context features, calculate the re-scoring probability of each candidate sequence through a fully connected layer and Softmax activation, indicating the rationality of the sequence in the global context. Weightedly fuse the path probability of the CTCDecoder and the re-scoring probability of the AttentionDecoder (such as linear interpolation) to obtain the final score of each candidate sequence. Sort the candidate sequences in descending order according to the final score, and retain the top M high-score candidates (such as M = 3) for subsequent steps to select.

[0128] Step 172: From the re-scored candidate sequences, select the sequence with the highest score as the preliminary recognition result. If the highest score is lower than a preset threshold (e.g., 0.5), trigger a fallback mechanism (e.g., return the original output of the CTCDecoder or mark it as a low-confidence result). Remove special tokens (e.g., blank labels) from the candidate sequences, merge duplicate characters, and restore them to the natural text format. If time information needs to be retained, align the preliminary recognition result with the time steps output by the Encoder to generate a text sequence with timestamps.

[0129] Step 173: According to the pre-trained language model, detect and correct spelling mistakes in the preliminary result (e.g., "recognize" misspelled as "recogize"); adjust the word order or add punctuation to make the text conform to grammar rules (e.g., correct "hegoschool" to "Hegoestoschool"); replace domain-specific abbreviations or terms with standard forms; adjust the text format according to the application scenario (e.g., convert the date format "2023 / 10 / 5" to "October 5, 2023"); if the preliminary result is significantly different from other candidate sequences, select a more reasonable version in combination with context information (e.g., the previous and next sentences); normalize the probabilities of the high-score candidate sequences and generate a fused result according to weights (e.g., fuse "hegoesschool" and "hegoestoschool" into the latter) to obtain the final text recognition result.

[0130] As Figure 2 shown, an embodiment of the present invention also provides a speech recognition system 20 for a minority language constructed based on Whisper tokens, including:

[0131] An acquisition module 21, configured to extract all tokens related to the target minority language in the Whisper tokenizer to form an initial candidate set; match and analyze the tokens in the initial candidate set with the collected training text corpus of the target minority language, and count the occurrence frequencies of the tokens in the corpus;

[0132] A construction module 22, configured to screen out high-frequency tokens and supplement low-frequency tokens according to the frequency statistics results to construct a dynamic vocabulary; construct a speech recognition model, the model includes an encoder, a decoder, a CTC decoder, and a CTC loss function, and train the model through the dynamic vocabulary;

[0133] The processing module 23 is used to extract Fbank features from the input speech signal, preprocess the extracted Fbank features, input them into the Encoder for forward calculation to generate high-level semantic features; perform CTCDecoder processing on the high-level semantic features, generate multiple candidate results through the BeamSearch strategy; use AttentionDecoder to re-score these candidate results, select the result with the highest score as the preliminary recognition result, and perform post-processing on the preliminary recognition result to obtain the final text recognition result.

[0134] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle described in the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A speech recognition method for constructing minority languages based on whisper token, characterized in that, The method includes: Extract all tokens related to the target minority language in the Whisper tokenizer to form an initial candidate set; Match and analyze the tokens in the initial candidate set with the collected training text corpus of the target minority language, and count the occurrence frequencies of the tokens in the corpus; Screen high-frequency tokens according to the frequency statistics results and supplement low-frequency tokens to construct a dynamic vocabulary; Construct a speech recognition model, which includes an encoder, a decoder, a CTC decoder, and a CTC loss function, and train the model with the dynamic vocabulary; Extract Fbank features from the input speech signal, preprocess the extracted Fbank features, and input them into the Encoder for forward calculation to generate high-level semantic features; Perform CTCDecoder processing on the high-level semantic features to generate multiple candidate results through the BeamSearch strategy; Use the AttentionDecoder to re-score these candidate results, select the result with the highest score as the preliminary recognition result, and post-process the preliminary recognition result to obtain the final text recognition result.

2. The method for speech recognition of a minority language based on whisper token according to claim 1, characterized in that, Extract all tokens related to the target minority language in the Whisper tokenizer to form an initial candidate set, including: Obtain the pre-trained Whisper tokenizer model and identify the target minority language; Traverse all tokens in the Whisper tokenizer and extract tokens related to the target minority language; integrate all related tokens into a set to form an initial candidate set.

3. The speech recognition method for constructing a minority language based on whisper token according to claim 1, wherein Match and analyze the tokens in the initial candidate set with the collected training text corpus of the target minority language, and count the occurrence frequencies of the tokens in the corpus, including: Collect and organize the training text corpus of the target minority language. The corpus format is a plain text file, the encoding format is UTF-8, and the content contains diverse text data of the target minority language; Match each token in the initial candidate set with the training text corpus of the target minority language, identify the tokens that appear in the corpus, and conduct analysis; Count the occurrence frequencies of each token in the training text corpus of the target minority language to obtain the frequency statistics results.

4. The method for speech recognition of minority languages based on whispertoken according to claim 1, characterized in that, Screen high-frequency tokens according to the frequency statistics results and supplement low-frequency tokens to construct a dynamic vocabulary, including: Set the threshold T according to the dynamic adjustment method of data distribution; According to the set threshold, screen out the tokens with higher occurrence frequencies from the frequency statistics results. For the tokens with lower occurrence frequencies, supplement them through specific processing methods; Merge the screened high-frequency tokens and the processed low-frequency tokens to construct a dynamic vocabulary.

5. The voice recognition method for constructing minority languages based on whisper token according to claim 1, characterized in that Construct a speech recognition model, which includes an encoder, a decoder, a CTC decoder, and a CTC loss function, and train the model with the dynamic vocabulary, including: Use a deep learning framework to construct a speech recognition model, which includes an encoder, a decoder, a CTC decoder, and a CTC loss function; Use a signal processing library to extract features from speech signals and convert the transcribed text into a token sequence; Collect a small language corpus of speech signals and their corresponding texts, perform preprocessing to form a structured dataset, and divide the dataset into a training set, a validation set, and a test set; Initialize each parameter in the model, write code for forward propagation, loss calculation, backpropagation, and parameter update, set the hyperparameter adjustment strategy, and align speech features with text tokens by jointly optimizing the CTC loss and the Attention decoder loss; Use the test set to evaluate the trained model and calculate various performance metrics.

6. The speech recognition method for constructing a minority language based on whisper token according to claim 1, wherein Extract Fbank features from the input speech signal, preprocess the extracted Fbank features, and input them into the Encoder for forward calculation to generate high-level semantic features, including: Extract Fbank features from the input speech signal to obtain the spectral feature representation of the speech signal, and preprocess the extracted Fbank features; Input the preprocessed Fbank features into the Encoder, and the Encoder converts the input Fbank features into a high-level semantic feature representation through forward calculation.

7. The method for speech recognition of minority languages based on whisper token according to claim 1, characterized in that, Perform CTCDecoder processing on the high-level semantic features, generate multiple candidate results through the BeamSearch strategy; use the AttentionDecoder to re-score these candidate results, select the result with the highest score as the preliminary recognition result, and perform post-processing on the preliminary recognition result to obtain the final text recognition result, including: Input the high-level semantic features into the CTCDecoder; The CTCDecoder processes the high-level semantic features and generates multiple candidate results through the BeamSearch strategy; Use the AttentionDecoder to re-score the multiple candidate results generated by the CTCDecoder; According to the re-scoring results, select the candidate result with the highest score as the preliminary recognition result; Perform post-processing on the preliminary recognition result to obtain the final text recognition result.

8. A speech recognition system for constructing a minority language based on whisper token, the system implementing the method described in any one of claims 1 to 7, characterized in that, Including: An acquisition module for extracting all tokens related to the target small language in Whispertokenizer to form an initial candidate set; matching and analyzing the tokens in the initial candidate set with the collected target small language training text corpus, and counting the occurrence frequencies of the tokens in the corpus; A construction module for screening high-frequency tokens and supplementing low-frequency tokens according to the frequency statistics results to construct a dynamic vocabulary; Construct a speech recognition model, the model includes an encoder, a decoder, a CTC decoder, and a CTC loss function, and train the model through the dynamic vocabulary; A processing module for extracting Fbank features from the input speech signal, preprocessing the extracted Fbank features, and inputting them into the Encoder for forward calculation to generate high-level semantic features; Perform CTCDecoder processing on high-level semantic features, and generate multiple candidate results through the BeamSearch strategy; use AttentionDecoder to re-score these candidate results, select the result with the highest score as the preliminary recognition result, and perform post-processing on the preliminary recognition result to obtain the final text recognition result.

9. A computing device, characterized in that, Including: One or more processors; A storage device for storing one or more programs, which when executed by the one or more processors cause the one or more processors to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A program is stored in the computer-readable storage medium, and when the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method and system for unsupervised discovery of unigrams in speech recognition systems

    CA3236971A1

  • Multilingual TTS system based on sub-word level introducing depth feature loss and graph convolution

    CN118248120A

  • Speech recognition model training method, speech recognition model testing method, speech recognition method and speech recognition device

    CN118553234A

  • Large model fine tuning method of track domain knowledge base and scene adaptation system

    CN118606439A

  • Systems and methods for continual learning for end to-end automatic speech recognition

    US20240290319A1

Cited By

  • Speech recognition method and device for reducing command word misrecognition, equipment and medium

    CN120808762A

  • Manual intervention word segmentation method for improving ASR recognition effect

    CN120833785A

  • Speech recognition method and device, electronic equipment and computer readable storage medium

    CN121393430A

  • Speech recognition method and device for intelligent equipment, equipment and medium

    CN121506127A

  • Helium-oxygen environment speech recognition method and system based on Whisper

    CN122177093A