Voice recognition method and system based on AI large model

By combining anti-noise suppression network and spectrum quality scores with dialect thermal gallery, acoustic and language model adaptation matrix is generated, which solves the problems of low accuracy in dialect recognition and high recognition error rate in noise environments, and achieves efficient dialect recognition.

CN120340465AActive Publication Date: 2025-07-18GUANGZHOU WEIJIE INTELLIGENT TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510555177.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-18
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The existing speech recognition system has problems in the field of dialect recognition with low recognition accuracy, inability to accurately capture pronunciation deviations, lack of speech quality evaluation mechanisms and dynamic decoding strategies, especially in noisy environments with high recognition error rates.

Method used

The pre-trained anti-noise suppression network is used for noise reduction and frequency band enhancement, combined with spectrum quality score, a dialect thermal graph library is constructed and a dynamic screening mechanism is designed, acoustic and language model adaptation matrix is generated through the hypernetwork, and a fusion thermal graph and adversarial discriminative network are used for speech recognition.

Benefits of technology

It improves the accuracy of dialect recognition, improves the model's adaptability to different dialects, optimizes the computing efficiency, and maintains stable recognition performance in a noisy environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340465A_ABST
    Figure CN120340465A_ABST
Patent Text Reader

Abstract

The invention provides a voice recognition method and system based on an AI large model, and belongs to the technical field of voice recognizing.The method comprises the steps that firstly, noise reduction and frequency band enhancement in a noise environment are achieved through a pre-trained anti-noise suppression network, and a basis is provided for follow-up processing in combination with multi-dimensional spectrum quality scores; based on similarity matching of element feature vectors and a pre-constructed dialect thermodynamic diagram library and matching weight adjustment based on frequency spectrum quality scores, accurate modeling of specific dialect pronunciation deviation is achieved, moreover, through an acoustic adaptation matrix and a language model adaptation matrix generated through a super network, the adaptability of the model to different dialects is improved, and in addition, the accuracy of the model is improved. According to the method, a fusion thermodynamic diagram and an acoustic adaptation matrix are jointly injected into a pre-trained acoustic model, the recognition accuracy of dialect phonemes is improved through multi-level attention correction, and finally, accurate speech recognition is achieved by adopting a thermodynamic diagram guided cluster search algorithm and combining verification of an adversarial discrimination network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recognition, and in particular to a speech recognition method and system based on an AI large model. Background Art

[0002] With the rapid development of artificial intelligence technology, speech recognition systems based on deep learning have made significant progress in standard Mandarin scenarios, but still face major challenges in the field of dialect recognition. Traditional speech recognition methods usually use a single general model to process all dialects, ignoring the significant pronunciation differences and acoustic feature changes between dialects, resulting in a significant decrease in recognition accuracy. In the prior art, although the noise reduction method based on adversarial training can effectively suppress environmental noise, it lacks a quantitative evaluation mechanism for the quality of the denoised speech and cannot provide a reliable basis for subsequent dialect adaptation. The mainstream multi-dialect recognition solutions mostly adopt simple model fine-tuning or hybrid training strategies, which are difficult to accurately capture the phoneme-level pronunciation deviation characteristics of specific dialects. More critically, the current system often adopts a fixed strategy in the decoding stage and cannot dynamically adjust the recognition process according to the quality of the input speech, resulting in a high recognition error rate for low-quality dialect speech. In addition, traditional heatmap technology is mainly applied in the field of computer vision, and its application in speech recognition is still limited to coarse-grained dialect classification and fails to establish the association between pronunciation deviation and the attention mechanism of the acoustic model. These technical defects make it difficult for existing systems to meet the high-precision requirements for dialect recognition in practical applications, especially in complex dialect scenarios under noisy environments. There is an urgent need for an innovative solution that can integrate speech quality evaluation, fine-grained pronunciation deviation modeling, and dynamic decoding strategies.

[0003] Therefore, it is necessary to provide a speech recognition method and system based on an AI large model to solve the above technical problems. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a speech recognition method and system based on an AI large model, achieving the beneficial effect of high-precision dialect recognition.

[0005] The present invention provides a speech recognition method based on an AI large model. The speech recognition method includes the following steps: S1: Denoise and enhance the frequency band of the original speech signal through a pre-trained adversarial noise suppression network to generate a Mel spectrogram matrix and a spectrum quality score; S2: Extract meta-feature vectors from the Mel spectrogram matrix, screen N dialects from a pre-constructed dialect heatmap library based on a preset screening method to form a preselected dialect set, and calculate the matching weights of each dialect in the preselected dialect set in combination with the spectrum quality score; S3: Load the pronunciation deviation heat maps of all dialects in the preselected dialect set from the pre - built dialect heat map library, and generate a fused heat map in combination with the matching weights. At the same time, input the meta - feature vector into the hyper - network to obtain the acoustic adaptation matrix and the language model adaptation matrix; S4: Input the fused heat map and the acoustic adaptation matrix into the pre - trained acoustic model to correct the attention weights and output the corrected phoneme probability distribution; S5: Input the language model adaptation matrix into the pre - trained language model to obtain the dynamic language model probability, and in combination with the corrected phoneme probability distribution, use a preset algorithm to generate a candidate text list, and select the candidate text with the highest confidence in the candidate text list through the pre - trained adversarial discriminant network as the final speech recognition result.

[0006] Preferably, step S1 includes the following steps: S101: Extract the multi - scale noise distribution features of the original speech signal through the pre - trained adversarial noise suppression network, and generate the denoised intermediate speech representation and the noise suppression intensity; S102: Conduct frequency - band energy analysis on the denoised intermediate speech representation, calculate the signal - to - noise ratio and the average frequency - band signal - to - noise ratio of each frequency band, and perform non - linear gain compensation on the frequency bands with signal - to - noise ratio lower than the preset signal - to - noise ratio threshold to obtain the frequency - band enhanced speech signal, and at the same time calculate the frequency - band energy balance degree; S103: Map the frequency - band enhanced speech signal to the Mel scale to generate the Mel spectrogram matrix. At the same time, construct a multi - dimensional spectral quality score based on the noise suppression intensity, the frequency - band energy balance degree, and the average frequency - band signal - to - noise ratio.

[0007] Preferably, in step S2, the screening step of screening N dialects to form the preselected dialect set includes: Based on the preset basic N value and the maximum signal - to - noise ratio supported by the system, multiply the preset basic N value by the ratio of the average frequency - band signal - to - noise ratio to the maximum signal - to - noise ratio supported by the system to obtain the first corrected N value. If the spectral quality score is higher than the preset score threshold, then use the first corrected N value as the final N value. If the spectral quality score is not higher than the preset score threshold, then perform a secondary correction; The secondary correction obtains the final N value by multiplying the first corrected N value by the spectral quality score, and screens out the dialects with the number of the final N value to form the preselected dialect set.

[0008] Preferably, in step S2, the step of calculating the matching weights of all dialects in the preselected dialect set in combination with the spectral quality score includes: Calculate the cosine similarity between the meta - feature vector and all dialects in the preselected dialect set as the initial matching weight; Based on historical data statistics, obtain the noise sensitivity coefficient, the balance sensitivity coefficient, and the signal - to - noise ratio sensitivity coefficient of each dialect respectively; Perform a linear combination calculation on the three sub - dimensions of the spectrum quality score, namely the noise suppression intensity, the frequency band energy balance degree, and the average frequency band signal - to - noise ratio, and the noise sensitivity coefficient, the balance sensitivity coefficient, and the signal - to - noise ratio sensitivity coefficient of each dialect in the pre - selected dialect set to obtain the weight adjustment coefficient of each dialect in the pre - selected dialect set; Adjust the initial matching weights based on the weight adjustment coefficients of each dialect in the pre - selected dialect set to obtain the matching weights of each dialect in the pre - selected dialect set.

[0009] Preferably, in step S3, the pronunciation deviation heat map includes three dimensions: dialect types, standard phoneme set, and deviation dimension.

[0010] Preferably, in step S3, the generation process of the acoustic adaptation matrix adopts orthogonal constraint processing Preferably, step S4 includes the following steps: S401: Inject the acoustic adaptation parameter matrix into the intermediate Transformer layer of the pre - trained acoustic model to perform matrix multiplication operations with the query vector and the key vector respectively, generate the corrected query projection vector and key projection vector, and calculate the initial attention score matrix; S402: Convert the fusion heat map into an attention bias matrix, and inject it into the intermediate Transformer layer of the pre - trained acoustic model to perform matrix addition operations with the initial attention score matrix to generate the corrected attention weights; S403: Generate the encoder feature representation based on the corrected attention weights, and output the corrected phoneme probability distribution through the phoneme classification layer.

[0011] Preferably, in step S5, the preset algorithm is a beam search algorithm guided by the fusion heat map.

[0012] Preferably, during the execution of the beam search algorithm guided by the fusion heat map, the decoding frame rate of the acoustic model is dynamically adjusted using the spectrum quality score.

[0013] The present invention provides a speech recognition system based on an AI large model, which is applied to a speech recognition method based on an AI large model. The speech recognition system includes: A speech enhancement and evaluation module, which is used to denoise and enhance the frequency band of the original speech signal through a pre - trained anti - noise suppression network to generate a Mel spectrogram matrix and a spectrum quality score; A dialect feature matching module, which is used to extract meta - feature vectors from the Mel spectrogram matrix, screen N dialects from a pre - constructed dialect heat map library based on a preset screening method to form a pre - selected dialect set, and calculate the matching weights of each dialect in the pre - selected dialect set in combination with the spectrum quality score; The multi-modal parameter generation module is used to load the pronunciation deviation heat maps of various dialects in the pre-built dialect heat map library for the pre-selected dialect set, generate a fused heat map in combination with the matching weights. At the same time, the meta-feature vector is input into the hyper-network to obtain the acoustic adaptation matrix and the language model adaptation matrix; The acoustic model correction module is used to input the fused heat map and the acoustic adaptation matrix into the pre-trained acoustic model for attention weight correction, and output the corrected phoneme probability distribution; The speech recognition decision module is used to input the language model adaptation matrix into the pre-trained language model to obtain the dynamic language model probability, and in combination with the corrected phoneme probability distribution, generate a candidate text list using a preset algorithm, and select the candidate text with the highest confidence in the candidate text list as the final speech recognition result through the pre-trained adversarial discriminant network.

[0014] Compared with the related technologies, a speech recognition method and system based on the AI large model provided by the present invention have the following beneficial effects: The present invention first realizes accurate noise reduction and frequency band enhancement in a noisy environment through the pre-trained adversarial noise suppression network, and provides a reliable quality basis for subsequent processing in combination with the multi-dimensional spectrum quality score; secondly, it innovatively constructs a dialect heat map library and designs a dynamic screening mechanism, and based on the similarity matching between the meta-feature vector and the heat map library and the weight adjustment based on the spectrum quality score, realizes the accurate modeling of the pronunciation deviation of specific dialects; furthermore, through the acoustic adaptation matrix and the language model adaptation matrix generated by the hyper-network, combined with orthogonal constraint processing to ensure parameter stability, improves the adaptability of the model to different dialects; in addition, maps the fused heat map to an attention bias matrix and injects it into the acoustic model together with the adaptation parameters, and improves the recognition accuracy of dialect phonemes through a multi-level attention correction mechanism; finally, adopts a heat map-guided beam search algorithm and combines the verification mechanism of the adversarial discriminant network to optimize the calculation efficiency while ensuring the recognition accuracy, especially the strategy of dynamically adjusting the decoding frame rate through the spectrum quality score, and can still maintain stable recognition performance in a complex noise environment. This method has significant advantages in terms of dialect recognition accuracy, calculation efficiency, and system robustness while maintaining the general speech recognition ability. Description of the Drawings

[0015] Figure 1 It is a flowchart of a speech recognition method based on the AI large model of the present invention; Figure 2 It is a module structure diagram of a speech recognition system based on the AI large model of the present invention. Detailed Embodiments

[0016] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the present invention, rather than limiting the present invention. Additionally, it should be noted that for the convenience of description, only the parts related to the present invention are shown in the drawings, rather than all the structures. Furthermore, the embodiments in the present invention and the features in the embodiments can be combined with each other without conflict.

[0017] It should also be noted that for the convenience of description, only the parts related to the present invention are shown in the drawings, rather than all the content. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as being processed sequentially, many of the operations can be performed in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but there can also be additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0018] Embodiment 1 A speech recognition method based on an AI large model. In the specific implementation process, as Figure 1 shown, it shows a flowchart of a speech recognition method based on an AI large model. The speech recognition method includes: Step S1: Denoise and enhance the frequency band of the original speech signal through a pre-trained anti-noise suppression network to generate a Mel spectrogram matrix and a spectral quality score.

[0019] Specifically, step S1 includes the following steps: S101: Extract the multi-scale noise distribution features of the original speech signal through a pre-trained anti-noise suppression network to generate a denoised intermediate speech representation and a noise suppression intensity; S102: Perform frequency band energy analysis on the denoised intermediate speech representation, calculate the signal-to-noise ratio and the average signal-to-noise ratio of each frequency band, and perform non-linear gain compensation on the frequency bands with a signal-to-noise ratio lower than the preset signal-to-noise ratio threshold to obtain a frequency band-enhanced speech signal, and at the same time calculate the frequency band energy balance degree; S103: Map the frequency band-enhanced speech signal to the Mel scale to generate a Mel spectrogram matrix. At the same time, construct a multi-dimensional spectral quality score based on the noise suppression intensity, the frequency band energy balance degree, and the average signal-to-noise ratio of the frequency band.

[0020] In the specific implementation process, first, the input original speech signal is processed by a pre-trained anti-noise suppression network. The pre-trained anti-noise suppression network adopts an architecture that combines a complex-domain convolutional module and a temporal dependence modeling unit. The complex-domain convolutional module contains 3 layers of complex convolutional layers, with the number of filters in each layer being 32, 64, and 128 respectively. The temporal dependence modeling unit adopts a bidirectional LSTM structure with 256 hidden units. The steady-state noise in the speech signal, such as ambient background noise, and non-steady-state noise, such as sudden interference, are separated through multi-scale feature extraction technology to generate a denoised intermediate speech representation and output the noise suppression intensity. Then, band energy analysis is performed on the denoised speech signal. The instantaneous signal-to-noise ratio of each frequency band is calculated using a 128-point FFT. Through a sliding window, for example, the window duration is set to 200 ms and the step size is 10 ms, the mean signal-to-noise ratio of each frequency band is calculated. For the frequency bands with a signal-to-noise ratio lower than a preset signal-to-noise ratio threshold, for example, the typical value of the preset signal-to-noise ratio threshold is 5 dB, a non-linear gain compensation algorithm based on the companding principle is used for the frequency bands with a signal-to-noise ratio lower than 5 dB, the compression ratio is 2:1, and the startup time is 5 ms. At the same time, the entropy quantization is used to calculate the balance degree of the energy distribution of each frequency band. Finally, the processed speech signal is passed through a 40-channel Mel filter bank with a frequency range of 80 - 8000 Hz and mapped to the Mel scale, thereby generating a Mel spectrogram matrix with dimensions of T×40, where T is the number of time frames. And a comprehensive spectral quality score is constructed based on the three dimensions of noise suppression intensity, band energy balance degree, and band signal-to-noise ratio mean. Among them, the noise suppression intensity is normalized by the Sigmoid function, the band energy balance degree is normalized by Min-Max, and the band signal-to-noise ratio mean is processed by logarithmic compression. Finally, the spectral quality score is obtained through linear weighting with a preset weighting factor.

[0021] Step S2: Extract meta-feature vectors from the Mel spectrogram matrix, screen N dialects from a pre-constructed dialect heat map library based on a preset screening method to form a preselected dialect set, and calculate the matching weights of each dialect in the preselected dialect set in combination with the spectral quality score.

[0022] Specifically, in step S2, the screening steps for screening N dialects to form a preselected dialect set include: Based on a preset basic N value and the maximum signal-to-noise ratio supported by the system, multiply the preset basic N value by the ratio of the band signal-to-noise ratio mean to the maximum signal-to-noise ratio supported by the system to obtain a first corrected N value. If the spectral quality score is higher than the preset score threshold, then use the first corrected N value as the final N value. If the spectral quality score is not higher than the preset score threshold, then perform a secondary correction; The secondary correction is to multiply the first corrected N value by the spectral quality score to obtain the final N value, and screen out the dialects with the number of the final N value to form a preselected dialect set.

[0023] Specifically, in step S2, the steps of calculating the matching weights of each dialect in the preselected dialect set in combination with the spectrum quality score include: Calculating the cosine similarity between the meta-feature vector and each dialect in the preselected dialect set as the initial matching weight; Based on historical data statistics, obtain the noise sensitivity coefficient, equilibrium sensitivity coefficient, and signal-to-noise ratio sensitivity coefficient of each dialect respectively; Perform a linear combination calculation on the three sub-dimensions of the spectrum quality score, namely, noise suppression intensity, frequency band energy equilibrium, and average frequency band signal-to-noise ratio, and the noise sensitivity coefficient, equilibrium sensitivity coefficient, and signal-to-noise ratio sensitivity coefficient of each dialect in the preselected dialect set to obtain the weight adjustment coefficient of each dialect in the preselected dialect set; Adjust the initial matching weight based on the weight adjustment coefficient of each dialect in the preselected dialect set to obtain the matching weight of each dialect in the preselected dialect set.

[0024] In the specific implementation process, a deep convolutional neural network is first used to extract meta-feature vectors from the Mel-spectrogram matrix. The deep convolutional neural network adopts a four-layer convolution structure design. Each layer includes a convolution layer with a convolution kernel size of 3×3, a batch normalization layer and a ReLU activation function. The number of filters in each layer is 64, 128, 256, and 512 respectively, and the step size is set to 1×1. Each layer is followed by a 2×2 maximum pooling layer for downsampling, and finally a 256-dimensional meta-feature vector is output through a global average pooling layer; then based on the preset basic N value, an exemplary , typically set to 10, adjusted between 5-15 according to actual applications, and the maximum signal-to-noise ratio supported by the system, fixed at 40dB, the first revised N value is obtained by multiplying the preset basic N value by the ratio of the average signal-to-noise ratio of the frequency band to the maximum signal-to-noise ratio supported by the system, if the spectrum quality score is higher than the preset score threshold, exemplary, the score threshold is set to 0.7, the configurable range is 0.6-0.8, then the first revised N value is used as the final N value; if the spectrum quality score is not higher than the preset score threshold, a secondary correction is performed, and the secondary correction is performed by The final N value is obtained by multiplying the corrected N value by the spectrum quality score. Based on the determined final N value, the size of the pre-selected dialect set is obtained. The top N dialects are sorted from high to low based on cosine similarity to form the pre-selected dialect set. After the pre-selected dialect set is determined, the cosine similarity algorithm is used to calculate the cosine similarity between the meta-feature vector and the heat map of each candidate dialect to generate the initial matching weight; at the same time, the three sub-dimensions of the spectrum quality score of each dialect are obtained from the pre-constructed dialect feature database, namely, the sensitivity of noise suppression intensity, frequency band energy balance and frequency band signal-to-noise ratio average. The sensitivity coefficients are noise sensitivity coefficient, equalization sensitivity coefficient and signal-to-noise ratio sensitivity coefficient. These coefficients are obtained through multivariate linear regression analysis of historical data and at least 1000 hours of annotated speech. The noise suppression intensity, frequency band energy balance and frequency band signal-to-noise ratio mean are linearly combined with the noise sensitivity coefficient, equalization sensitivity coefficient and signal-to-noise ratio sensitivity coefficient of each dialect in the pre-selected dialect set through a preset differentiated weight correction formula to obtain the weight adjustment coefficient of each dialect in the pre-selected dialect set. The preset differentiated weight correction formula is: in, Weight adjustment factor, is the noise sensitivity coefficient, is the noise suppression strength, Equilibrium sensitivity coefficient, To balance the sensitivity, is the signal-to-noise ratio sensitivity coefficient, is the mean signal-to-noise ratio of the frequency band. The matching weight of each dialect in the pre-selected dialect set is obtained by multiplying the initial matching weight with the weight adjustment coefficient. Finally, it is normalized by the Softmax function to ensure that the sum of the matching weights is 1. Step S3: Load the pronunciation deviation heat maps of all dialects in the preselected dialect set from the pre-constructed dialect heat map library, and generate a fused heat map by combining the matching weights. At the same time, input the meta-feature vector into the hypernetwork to obtain the acoustic adaptation matrix and the language model adaptation matrix.

[0025] Specifically, in step S3, the pronunciation deviation heat map includes three dimensions: dialect type, standard phoneme set, and deviation dimension.

[0026] Specifically, in step S3, the generation process of the acoustic adaptation matrix adopts orthogonal constraint processing.

[0027] In the specific implementation process, first load the pronunciation deviation heat maps of the preselected dialect set from the pre-constructed dialect heat map library. The pronunciation deviation heat map is stored in a three-dimensional tensor structure. The three dimensions are: dialect type, 120 sub-dialects covering 7 major dialect regions in the country, standard phoneme set, 586 basic phonemes extended based on the International Phonetic Alphabet, and deviation dimension, including three sub-dimensions: initial consonant deviation, final vowel deviation, and tone deviation. Each deviation intensity value is represented by a floating point number from 0 to 1 and is processed by z-score standardization. In the stage of fusing the pronunciation deviation heat maps, a weighted average algorithm is used to weight and fuse the pronunciation deviation heat maps of the dialects according to the corresponding matching weights of the dialects to obtain a fused heat map. Bilinear interpolation is used in the fusion process to ensure smooth transition between the pronunciation deviation heat maps of different dialects. At the same time, input the meta-feature vector into the hypernetwork to generate the adaptation matrix. The hypernetwork adopts a three-layer fully connected structure. The output dimension of the acoustic adaptation matrix generation branch is d×d, where d is the dimension of the hidden layer of the acoustic model, and the typical value is 768. The orthogonality of the matrix is ensured by Gram-Schmidt orthogonalization. The output dimension of the language model adaptation matrix generation branch is V×h, where V is the size of the vocabulary, and the typical value is 5000; h is the hidden dimension of the language model, and the typical value is 1024. The low-rank decomposition technology is used to reduce the number of parameters. The first two layers of the two branches share parameters, and the third layer is independent. All hidden layers use the GELU activation function and apply Dropout to prevent overfitting. The finally generated acoustic adaptation matrix is used to adjust the attention projection direction of the acoustic model, and the language model adaptation matrix is used to adjust the weight distribution of the output layer of the language model.

[0028] Step S4: Input the fused heat map and the acoustic adaptation matrix into the pre-trained acoustic model to correct the attention weights and output the corrected phoneme probability distribution.

[0029] Specifically, step S4 includes the following steps: S401: Inject the acoustic adaptation parameter matrix into the intermediate Transformer layer of the pre-trained acoustic model to perform matrix multiplication operations with the query vector and the key vector respectively, generate the corrected query projection vector and key projection vector, and calculate the initial attention score matrix; S402: Convert the fusion heatmap into an attention bias matrix, and inject it into the intermediate Transformer layer of the pre-trained acoustic model to perform matrix addition operations with the initial attention score matrix to generate the corrected attention weights; S403: Generate the encoder feature representation based on the corrected attention weights, and output the corrected phoneme probability distribution through the phoneme classification layer.

[0030] In the specific implementation process, first inject the acoustic adaptation matrix into the Transformer intermediate layer. After performing matrix transformation on the query vector and the key vector, calculate the initial attention score matrix through dot product operation. Then, use the phoneme forced alignment tool to perform fine time alignment on each phoneme in the fusion heatmap on the input speech time axis, determine the exact start and end positions of each phoneme on the time axis through the Viterbi algorithm, and perform special marking processing on the silent segments. Then, fill the pronunciation deviation intensity value corresponding to each phoneme into the block corresponding to the time period of the phoneme in the matrix, and keep the non-phoneme area as zero to generate the initial bias matrix. Then, perform Gaussian smoothing and temperature coefficient adjusted Softmax normalization processing on each row of the initial bias matrix to generate the attention bias matrix to ensure the reasonable distribution of attention weights. Then, fuse the attention bias matrix and the initial attention score matrix according to a preset ratio, and each head independently completes the fusion calculation in the multi-attention head mechanism. Finally, obtain the corrected attention weights through Softmax normalization. After generating the value vector weighted sum and the encoded features, input them into the classification layer containing 586 phoneme categories to output the corrected phoneme probability distribution.

[0031] Step S5: Input the language model adaptation matrix into the pre-trained language model to obtain the dynamic language model probability, and combine it with the corrected phoneme probability distribution. Use a preset algorithm to generate a candidate text list, and select the candidate text with the highest confidence in the candidate text list through the pre-trained adversarial discriminant network as the final speech recognition result.

[0032] Specifically, in step S5, the preset algorithm is the beam search algorithm guided by the fusion heatmap.

[0033] Specifically, during the execution of the beam search algorithm guided by the fusion heatmap, the decoding frame rate of the acoustic model is dynamically adjusted using the spectral quality score.

[0034] Specifically, first, the language model adaptation matrix is input into the pre-trained language model, and the dynamic language model probability is obtained by adjusting the weight distribution of the output layer. In the decoding stage, a beam search algorithm guided by a fused heatmap is adopted. This algorithm first extracts the set of phonemes in the fused heatmap whose pronunciation deviation intensity exceeds the preset deviation threshold, and increases the extended priority score of the candidate paths containing these phonemes by a preset proportion. At the same time, the pruning threshold is dynamically adjusted according to the average deviation intensity of the phoneme set. In the candidate path scoring stage, the corrected phoneme probability distribution and the average deviation intensity of all phonemes in the path are weighted and summed according to the preset weighted weights to generate a comprehensive ranking score. Based on the comprehensive ranking score, the candidate texts with a preset candidate quantity threshold are retained to form a candidate text list. Finally, the generated candidate text list is verified by the pre-trained adversarial discriminative network, and a confidence score is output and the candidate text with the highest score is selected as the final recognition result. During the entire decoding process, the system dynamically adjusts the decoding frame rate according to the spectral quality score. When the comprehensive ranking score is lower than the preset first threshold, full-frame decoding is adopted. When the comprehensive ranking score is between the preset first threshold and the preset second threshold, frame skipping decoding is adopted. When the comprehensive ranking score is higher than the preset second threshold, an efficient decoding mode is enabled.

[0035] The working principle of a speech recognition method based on an AI large model provided by the present invention is as follows: First, an adversarial noise suppression network is used to denoise and enhance the frequency band of the input speech, generate a Mel spectrogram, and calculate the spectral quality score. Then, candidate dialects are screened from the dialect heatmap library based on similarity and the matching weights are calculated. Next, the pronunciation deviation heatmaps of each candidate dialect are fused to obtain a fused heatmap. At the same time, an adaptation matrix for the acoustic model and the language model is generated through a hypernetwork. In the acoustic modeling stage, the fused heatmap is converted into an attention bias matrix, which acts on the Transformer encoder together with the acoustic adaptation parameters to correct the attention weights to optimize phoneme recognition and output the corrected phoneme probability distribution. Finally, the language model adaptation matrix is input into the pre-trained language model to obtain the dynamic language model probability. Then, combined with the corrected phoneme probability distribution, a beam search algorithm guided by the fused heatmap is used to generate a candidate text list, and the candidate text with the highest confidence is selected as the speech recognition result through the adversarial discriminative network.

[0036] Embodiment 2 A speech recognition system based on an AI large model, in the specific implementation process, as Figure 2 shown, which shows a module structure diagram of a speech recognition system based on an AI large model. The speech recognition system includes: A speech enhancement and evaluation module 100, which is used to denoise and enhance the frequency band of the original speech signal through a pre-trained adversarial noise suppression network, and generate a Mel spectrogram matrix and a spectral quality score; The dialect feature matching module 200 is used to extract meta-feature vectors from the Mel spectrogram matrix, screen N dialects from the pre-constructed dialect heat map library based on a preset screening method to form a preselected dialect set, and calculate the matching weights of each dialect in the preselected dialect set in combination with the spectrum quality score; The multi-modal parameter generation module 300 is used to load the pronunciation deviation heat maps of each dialect in the preselected dialect set from the pre-constructed dialect heat map library, and generate a fused heat map in combination with the matching weights. At the same time, the meta-feature vectors are input into the hypernetwork to obtain an acoustic adaptation matrix and a language model adaptation matrix; The acoustic model correction module 400 is used to input the fused heat map and the acoustic adaptation matrix into the pre-trained acoustic model for attention weight correction, and output the corrected phoneme probability distribution; The speech recognition decision module 500 is used to input the language model adaptation matrix into the pre-trained language model to obtain the dynamic language model probability, and in combination with the corrected phoneme probability distribution, generate a candidate text list using a preset algorithm, and select the candidate text with the highest confidence in the candidate text list as the final speech recognition result through the pre-trained adversarial discriminant network.

[0037] The working principle of a speech recognition system based on an AI large model provided by the present invention is as follows: First, the complex-domain adversarial network in the speech enhancement and evaluation module 100 performs multi-scale noise suppression and frequency band compensation on the original speech, generates a high-quality Mel spectrogram and outputs a multi-dimensional spectrum quality score including noise suppression intensity, frequency band equalization degree, and signal-to-noise ratio; the dialect feature matching module 200 uses a deep convolutional network to extract the deep meta-features of the speech, dynamically adjusts the screening strategy in combination with the spectrum quality score, selects the N most matching dialects from the pre-built dialect heat map library and calculates the matching weights; the multi-modal parameter generation module 300 obtains a fused heat map through weighted fusion, and at the same time uses the hypernetwork to generate the adaptation parameter matrices of the acoustic model and the language model; the acoustic model correction module 400 converts the fused heat map into a time-aligned attention bias matrix, which acts on the middle layer of the Transformer encoder together with the acoustic adaptation matrix, and optimizes phoneme recognition by reconstructing the query vector and the key vector and correcting the attention score; the final speech recognition decision module 500 uses a beam search algorithm guided by the fused heat map, dynamically adjusts the decoding strategy and fuses the language model probability, generates a candidate text list, and then verifies and selects the candidate text with the highest confidence as the final speech recognition result through the dual-channel adversarial discriminant network.

[0038] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one flow Figure 1 one or more flows and / or blocks Figure 1 or in a plurality of blocks.

[0039] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disc memories, tape memories, or any other medium that can be used to carry or store data and is computer-readable.

[0040] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, commodity or device comprising the element.

Claims

1. A speech recognition method based on a large AI model, characterized in that, The speech recognition method includes the following steps: S1: Denoise and enhance the frequency band of the original speech signal through a pre-trained anti-noise suppression network to generate a Mel spectrogram matrix and a spectral quality score; S2: Extract meta-feature vectors from the Mel spectrogram matrix, screen N dialects from a pre-constructed dialect heat map library based on a preset screening method to form a preselected dialect set, and calculate the matching weights of each dialect in the preselected dialect set in combination with the spectral quality score; S3: Load the pronunciation deviation heat maps of each dialect in the preselected dialect set from the pre-constructed dialect heat map library, and generate a fused heat map in combination with the matching weights. At the same time, input the meta-feature vectors into a hypernetwork to obtain an acoustic adaptation matrix and a language model adaptation matrix; S4: Input the fused heat map and the acoustic adaptation matrix into a pre-trained acoustic model for attention weight correction, and output a corrected phoneme probability distribution; S5: Input the language model adaptation matrix into a pre-trained language model to obtain a dynamic language model probability, and in combination with the corrected phoneme probability distribution, generate a candidate text list using a preset algorithm, and select the candidate text with the highest confidence in the candidate text list as the final speech recognition result through a pre-trained adversarial discriminant network.

2. The speech recognition method based on an AI large model according to claim 1, characterized in that, Step S1 includes the following steps: S101: Extract the multi-scale noise distribution characteristics of the original speech signal through a pre-trained anti-noise suppression network to generate a denoised intermediate speech representation and a noise suppression intensity; S102: Perform frequency band energy analysis on the denoised intermediate speech representation, calculate the signal-to-noise ratio and the average signal-to-noise ratio of each frequency band, and perform non-linear gain compensation on the frequency bands with a signal-to-noise ratio lower than a preset signal-to-noise ratio threshold to obtain a frequency band-enhanced speech signal, and at the same time calculate the frequency band energy balance degree; S103: Map the frequency band-enhanced speech signal to the Mel scale to generate a Mel spectrogram matrix. At the same time, construct a multi-dimensional spectral quality score based on the noise suppression intensity, the frequency band energy balance degree, and the average signal-to-noise ratio of the frequency band.

3. A speech recognition method based on an AI large model according to claim 2, characterized in that, In step S2, the screening step of screening N dialects to form a preselected dialect set includes: Based on a preset base N value and the maximum signal-to-noise ratio supported by the system, multiply the preset base N value by the ratio of the average signal-to-noise ratio of the frequency band to the maximum signal-to-noise ratio supported by the system to obtain a first corrected N value. If the spectral quality score is higher than a preset score threshold, take the first corrected N value as the final N value. If the spectral quality score is not higher than the preset score threshold, perform a secondary correction; The secondary correction obtains the final N value by multiplying the first corrected N value by the spectral quality score, and screens out the dialects with the number of the final N value to form a preselected dialect set.

4. A speech recognition method based on an AI large model according to claim 3, characterized in that, In step S2, the step of calculating the matching weights of each dialect in the preselected dialect set in combination with the spectral quality score includes: Calculate the cosine similarity between the meta-feature vector and each dialect in the preselected dialect set as the initial matching weight; Based on historical data statistics, obtain the noise sensitivity coefficient, the balance sensitivity coefficient, and the signal-to-noise ratio sensitivity coefficient of each dialect respectively; Perform a linear combination calculation on the three sub - dimensions of the spectrum quality score, namely, the noise suppression intensity, the frequency band energy balance degree, and the average frequency band signal - to - noise ratio, and the noise sensitivity coefficient, the balance sensitivity coefficient, and the signal - to - noise ratio sensitivity coefficient of each dialect in the pre - selected dialect set to obtain the weight adjustment coefficient of each dialect in the pre - selected dialect set; Adjust the initial matching weights based on the weight adjustment coefficients of each dialect in the pre - selected dialect set to obtain the matching weights of each dialect in the pre - selected dialect set.

5. A speech recognition method based on an AI large model according to claim 4, characterized in that, In step S3, the pronunciation deviation heat map includes three dimensions: dialect type, standard phoneme set, and deviation dimension.

6. The speech recognition method based on the AI large model according to claim 5, wherein In step S3, the generation process of the acoustic adaptation matrix adopts orthogonal constraint processing.

7. The speech recognition method based on the AI large model according to claim 6, characterized in that Step S4 includes the following steps: S401: Inject the acoustic adaptation parameter matrix into the middle Transformer layer of the pre - trained acoustic model to perform matrix multiplication operations with the query vector and the key vector respectively, generate the corrected query projection vector and key projection vector, and calculate the initial attention score matrix; S402: Convert the fusion heat map into an attention bias matrix, and inject it into the middle Transformer layer of the pre - trained acoustic model to perform matrix addition operations with the initial attention score matrix to generate the corrected attention weights; S403: Generate the encoder feature representation based on the corrected attention weights, and output the corrected phoneme probability distribution through the phoneme classification layer.

8. A speech recognition method based on an AI large model according to claim 7, characterized in that, In step S5, the preset algorithm is the beam search algorithm guided by the fusion heat map.

9. A speech recognition method based on an AI large model according to claim 8, characterized in that, During the execution of the beam search algorithm guided by the fusion heat map, dynamically adjust the decoding frame rate of the acoustic model using the spectrum quality score.

10. A speech recognition system based on a large AI model, characterized in that, Applied to a speech recognition method based on an AI large - model as described in any one of claims 1 - 9, the speech recognition system includes: A speech enhancement and evaluation module, which is used to perform noise reduction and frequency band enhancement on the original speech signal through a pre - trained anti - noise suppression network, and generate a Mel spectrogram matrix and a spectrum quality score; A dialect feature matching module, which is used to extract meta - feature vectors from the Mel spectrogram matrix, screen N dialects from the pre - constructed dialect heat map library based on a preset screening method to form a pre - selected dialect set, and calculate the matching weights of each dialect in the pre - selected dialect set in combination with the spectrum quality score; A multi - modal parameter generation module, which is used to load the pronunciation deviation heat map of each dialect in the pre - selected dialect set from the pre - constructed dialect heat map library, generate a fusion heat map in combination with the matching weights, and at the same time, input the meta - feature vectors into the hyper - network to obtain the acoustic adaptation matrix and the language model adaptation matrix; An acoustic model correction module, which is used to input the fusion heat map and the acoustic adaptation matrix into the pre - trained acoustic model for attention weight correction, and output the corrected phoneme probability distribution; A speech recognition decision module, which is used to input the language model adaptation matrix into the pre - trained language model to obtain the dynamic language model probability, and in combination with the corrected phoneme probability distribution, use the preset algorithm to generate a candidate text list, and select the candidate text with the highest confidence in the candidate text list as the final speech recognition result through the pre - trained adversarial discriminant network.

Citation Information

Patent Citations

  • Generative adversarial network speech enhancement method based on sparse continuous constraint

    CN113066483A

  • Chongqing dialect speech recognition method of transformer composed of double encoders

    CN116416968A

  • Communication enhancement method and system fused with noise scene, and storage medium

    CN116959467A

  • Speech Recognition Using Unspoken Text and Speech Synthesis

    US20210350786A1

  • Model training method and apparatus, dialect recognition method and apparatus, and server and storage medium

    WO2022121185A1

Cited By

  • A speech recognition method, system, computer device and medium

    CN122551783A