Voice recognition method based on acoustic model, computer equipment and storage medium
By introducing a gated fusion unit and auxiliary classifier into the acoustic model, dynamically adjusting the number of previewed future frames and the early exit mechanism, the problem of static binding between latency and accuracy in existing technologies is solved, the flexibility and accuracy of speech recognition are improved, and the user experience is enhanced.
Patent Information
- Application Number
- CN202511286713.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-09-10
AI Technical Summary
Existing acoustic models are unable to dynamically adjust the future context length in speech recognition tasks, resulting in a static binding between latency and accuracy. This makes it impossible to solve the recognition accuracy problem of local key instructions while ensuring low average latency.
A pre-trained gated fusion unit is introduced into the timing processing network layer of the acoustic model to dynamically determine the number of preview future frames, and the fusion weight is calculated through a lightweight neural network. The number of preview future frames is dynamically adjusted by combining short-term and long-term context representations. Multi-layer timing processing network layers and auxiliary classifiers are used to achieve early exit, and a joint loss function is used to optimize model training.
It dynamically adjusts the number of previewed future frames based on the input content, balances the delay and accuracy of speech recognition, improves the performance and user experience of the voice interaction system, reduces the recognition delay of complex commands and improves the response speed of simple commands.
Smart Images

Figure CN120783731A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of sound recognition, and in particular to a speech recognition method, computer equipment and storage medium based on an acoustic model. Background Art
[0002] The core of existing acoustic models, such as the Deep Feedforward Sequential Memory Network (DFSMN), lies in its "Memory Block," which learns and encodes long-term contextual information through a structure similar to a finite impulse response (FIR) filter. Specifically, the Memory Block simultaneously looks back at several past frames ("look-back") and looks ahead at several future frames ("look-ahead"), as summarized by the following formula:
[0003]
[0004] In this structure, is the output of the t-th frame of the l-th layer, is a nonlinear transformation, is the input feature, and is the memory module parameter, N1 is the number of frames that look back to the past, and N2 is the number of frames that look forward to the future. In actual streaming speech recognition applications, the total delay of the system is proportional to the number of preview orders of each layer of memory modules. Directly related, total delay The formula can be Therefore, by designing and adjusting The value can be used to control the delay of the model to meet the needs of real-time applications.
[0005] However, the inventors have discovered that while the above-mentioned existing technologies, represented by DFSMN, have achieved excellent performance in speech recognition tasks, they have the following inherent shortcomings in balancing processing delay and accuracy, especially in specific command word recognition scenarios: 1. Static Binding of Latency and Accuracy: In the DFSMN model, the look-ahead order is a global hyperparameter determined during model design and training. This means that the model uses the same fixed future information window for all input speech.
[0006] 2. A one-size-fits-all approach leads to a loss of focus: Scenario 1 (pursuing low latency): If the look-ahead order is set very small (for example, 0 or 1) to ensure a fast response for most simple command words (such as "play music"), then when encountering command word pairs that are very similar in acoustics and rely on subsequent keywords to distinguish them (such as "turn on" and "turn off"), the model's ability to distinguish them will be significantly reduced because it cannot see far enough into the future, resulting in an increased recognition error rate.
[0007] Scenario 2 (Striving for High Accuracy): Conversely, if the look-ahead order is set to a large value (e.g., 5) to accurately identify the aforementioned easily confused word pairs, while the accuracy of key commands is improved, this fixed high latency is imposed on all commands. This causes users to endure unnecessary waits even when speaking simple commands, significantly compromising the smoothness of the user interaction and the overall experience.
[0008] In short, the fundamental flaw of existing technologies lies in their inability to analyze input content and dynamically adjust the required future context length. They employ a static, global strategy to cope with dynamically changing input, failing to address the recognition accuracy of local key instructions while maintaining low average latency. Summary of the Invention
[0009] The present invention provides a speech recognition method, computer device and storage medium based on an acoustic model, aiming to solve the technical problem in the prior art that the acoustic model is unable to analyze the input content and dynamically adjust the required future context length.
[0010] To achieve the above-mentioned object, the present invention provides, in a first aspect, a speech recognition method based on an acoustic model, wherein the acoustic model includes multiple sequentially connected time series processing network layers with time series processing capabilities, and the method comprises: Obtaining speech features of the speech to be recognized; Inputting the speech features into the acoustic model for speech recognition, and outputting a recognition result; The method for processing input by the time series processing network layer in the acoustic model includes: Through the pre-trained gated fusion unit, the ratio of future frames to context information required to recognize the current input is determined; Calculate the number of preview future frames based on the ratio, and wait for and obtain the number of future frames; In combination with the number of future frames, a long-term context representation is calculated; The long-term context representation is processed and output to the next layer of the network.
[0011] Furthermore, the pre-trained gated fusion unit determines, before identifying the number of future frames required to preview the current input, the following steps include: calculating a short-term context representation of the current input; The processing of the long-term context representation and outputting it to the next layer of network includes: performing weighted calculation on the short-term context representation and the long-term context representation to obtain a fused context representation, and outputting the fused context representation to the next layer of network.
[0012] Furthermore, performing weighted calculation on the short-term context representation and the long-term context representation to obtain a fused context representation includes: The fused context representation is obtained by calculating the fusion formula, which is:
[0013] in, is the fused context representation; For short-term context representation; For long-term context representation; The gated fusion unit calculates the fusion weight value based on the input, and the fusion weight value indicates how much proportion of future frames should be fused.
[0014] Furthermore, the gated fusion unit is a lightweight neural network.
[0015] Furthermore, the acoustic model also includes a main classifier and at least one auxiliary classifier; the main classifier is connected to the last layer of the time series processing network layer; the auxiliary classifier is connected to the designated time series processing network layer before the last layer of the cFSMN, and the designated time series processing network layer is defined as an early exit layer; Among them, in the reasoning process of speech recognition, when the confidence of the output result of the auxiliary classifier connected to the early exit layer is higher than the confidence threshold, the result is used as the final speech recognition result and the subsequent reasoning process is stopped.
[0016] Furthermore, the acoustic model includes 12 time sequence processing network layers, and the 4th time sequence processing network layer and the 8th time sequence processing network layer are the early exit layers.
[0017] Furthermore, the training of the acoustic model uses a joint loss function, which is:
[0018] in, is the loss corresponding to the auxiliary classifier, is the loss corresponding to the total classifier; is the weight, which increases with the number of network layers.
[0019] Furthermore, the acoustic model is a model trained based on a deep feedforward sequence memory network, and the timing processing network layer is a cFSMN layer.
[0020] A second aspect of the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the above-mentioned acoustic model-based speech recognition methods when executing the computer program.
[0021] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-mentioned acoustic model-based speech recognition methods.
[0022] Beneficial effects: The acoustic model-based speech recognition method proposed in this paper introduces a pre-trained gated fusion unit into the acoustic model's temporal processing network layer to dynamically determine the ratio of future frames to contextual information required to recognize the current input. This allows for dynamic adjustment of the number of future frames. This overcomes the drawback of existing models such as the Deep Feedforward Sequential Memory Network (DFSMN), which statically bind latency and accuracy by using a fixed number of future frames. For simple commands, the model automatically uses a smaller number of future frames for low-latency and fast response. For commands like "play music," this reduces wait time. For easily confusing commands like "turn on" and "turn off," the model improves recognition accuracy by obtaining more future frames to calculate long-term context representations. Furthermore, by processing the long-term context representations and passing them to the next network layer, the model ensures effective processing of temporal information. Overall, this allows for dynamic adjustment of the number of future frames based on the input content, balancing latency and accuracy in speech recognition and improving the performance and user experience of the voice interaction system. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 A flowchart of a speech recognition method based on an acoustic model according to an embodiment of the invention; Figure 2 A flowchart illustrating a method for processing input by a time series processing network layer in an acoustic model according to an embodiment of the invention; Figure 3 The present invention is a block diagram showing the structure of a computer device according to an embodiment of the present invention.
[0024] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0026] Those skilled in the art will appreciate that, unless expressly stated otherwise, the singular forms "a", "an", "above", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of features, integers, steps, operations, elements, modules, modules and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, modules, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any module and all combinations of one or more associated listed items.
[0027] Those skilled in the art will understand that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art in the art to which this invention belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art and, unless specifically defined as such, will not be interpreted in an idealized or overly formal sense.
[0028] Reference Figure 1 An embodiment of the present invention provides a speech recognition method based on an acoustic model, wherein the acoustic model includes multiple sequentially connected time series processing network layers with time series processing capabilities, the method comprising: S1. Acquire speech features of the speech to be recognized.
[0029] In this step, speech features are parameters extracted from the speech signal that characterize the essential properties of speech, such as Mel-Frequency Cepstral Coefficients (MFCCs) and Linear Prediction Cepstral Coefficients (LPCCs). The speech signal is collected using a microphone or other device, then preprocessed, such as by framing and windowing, before being extracted using a feature extraction algorithm. For example, for the speech sound "play music," the speech signal is framed, with each frame length being 25ms and the frame shift being 10ms. MFCC features are then calculated for each frame, resulting in a set of speech feature vectors.
[0030] S2. Input the speech features into the acoustic model to perform speech recognition, and output the recognition result.
[0031] In this step, the acoustic model is the core component of the speech recognition system, used to model the probabilistic relationship between speech features and speech units (such as phonemes and words). The acoustic model sequentially inputs the extracted speech features into each layer of the acoustic model. After processing at each layer, the model's classifier outputs the recognition result. For example, the speech features of "playing music" obtained above are input into the acoustic model trained based on a deep feedforward sequential memory network (DFSMN). After processing by each time series processing network layer (cFSMN layer), the overall classifier outputs the recognition result of "playing music." The acoustic model processes the speech features, transforming them into recognition results and completing the core task of speech recognition.
[0032] Reference Figure 2 In this embodiment, the method for processing input by the time series processing network layer in the acoustic model includes: S21. Determine, through a pre-trained gated fusion unit, the ratio of future frames required to recognize the current input to the context information.
[0033] In this step, the gated fusion unit is a lightweight neural network that dynamically determines the proportion of future frames to be fused when calculating the context representation based on input information. This proportion is used to calculate the number of preview future frames. The number of preview future frames refers to the number of future frames that need to be waited for and acquired when processing the current frame. The gated fusion unit takes the current input's hidden layer state as input and, through network calculation, outputs a fusion weight value. This fusion weight value identifies the proportion of future frames that should be fused with the current input. This weight value is used to calculate the number of preview future frames required. For example, when processing easily confused speech sounds such as "turn on" and "turn off," the gated fusion unit outputs a value of 0.5 based on the current input's hidden layer state. If the window length for calculating the context representation is 10, the number of preview future frames is 0.5 * 10 = 5, obtaining more future information to distinguish between the two commands. The remaining frames are the current frame and the four previous frames. This overcomes the drawback of the existing fixed number of preview future frames and dynamically adjusts the number of preview future frames based on the input content, achieving a better balance between latency and accuracy.
[0034] S22: Calculate the number of preview future frames based on the ratio, and wait for and obtain the number of future frames.
[0035] In this step, the number of future frames previewed is calculated based on the fusion weights output by the gated fusion unit and the window length used to calculate the context representation. The system then waits for the appropriate amount of time to acquire future speech frames. If it determines that five future frames need to be previewed, the system waits for the next five frames of speech data while processing the current frame. This ensures that sufficient future frame information is acquired to provide data support for the subsequent calculation of the long-term context representation.
[0036] S23. Calculate a long-term context representation based on the number of future frames.
[0037] In this step, the long-term context representation is derived by comprehensively considering information from past historical frames, the current frame, and a certain number of future frames. It is used to capture long-term temporal dependencies. For example, the long-term context representation is calculated by combining speech features from past historical frames, the current frame, and five acquired future frames using a structure similar to a finite impulse response (FIR) filter. This fully utilizes information from future frames, improving the ability to model long-term temporal dependencies and thereby enhancing speech recognition accuracy.
[0038] S24: Process the long-term context representation and output it to the next layer of the network.
[0039] In this step, the calculated long-term context representation undergoes further processing, such as linear transformation and activation function processing, and then the processed result is output to the next network layer. For example, after linear transformation and activation function processing, the long-term context representation is output to the next cFSMN layer. Passing the processed long-term context representation to the next network layer enables the entire acoustic model to process speech features layer by layer, ultimately achieving accurate recognition results.
[0040] The speech recognition method based on the acoustic model of this embodiment includes multiple layers of time-series processing network layers with time-series processing capabilities. During the speech recognition process, the speech features of the speech to be recognized are first obtained and then input into the acoustic model. The time-series processing network layer in the acoustic model dynamically determines the number of future frames that need to be previewed through a pre-trained gated fusion unit, calculates the long-term context representation after obtaining the corresponding future frames, and outputs it to the next layer of the network after processing, finally completing the speech recognition and outputting the result. Unlike the DFSMN model in the prior art that adopts a fixed number of preview future frames, the present method realizes the dynamic adjustment of the number of preview future frames through the gated fusion unit, and can flexibly process according to different input contents, thereby improving the recognition accuracy of easily confused instructions while ensuring low latency. In other words, the method of the present application can dynamically adjust the number of preview future frames according to the characteristics of the input speech, avoiding the problem of static binding of delay and accuracy in the prior art. For simple command words, a smaller number of future frames can be used to achieve a quick response; for easily confused command words, a larger number of future frames can be used to utilize more future information to improve recognition accuracy, thereby improving the overall performance of the speech recognition system and improving the user experience.
[0041] In one embodiment, the pre-trained gated fusion unit determines the number of future frames that need to be previewed to identify the current input, and further includes calculating a short-term context representation of the current input.
[0042] The short-term context representation described above is obtained by considering only the current frame and historical frame information, with its look-ahead order for future frames fixed at 0. When processing the current input, the short-term context representation is calculated using a specific algorithm using the current frame and several previous frames. For example, when calculating the short-term context representation, information about future frames is not considered. Instead, only the speech features of the current frame and the previous three historical frames are used to calculate the short-term context representation using a structure similar to an FIR filter. This provides a context representation based solely on historical and current information for subsequent fusion with the long-term context representation, enabling the model to comprehensively consider information at different time scales.
[0043] The processing of the long-term context representation and outputting it to the next layer of network includes: performing weighted calculation on the short-term context representation and the long-term context representation to obtain a fused context representation, and outputting the fused context representation to the next layer of network.
[0044] In this step, the fused context representation is obtained by weighting the short-term context representation and the long-term context representation according to certain weights, which is used to comprehensively consider information at different time scales. The short-term context representation and the long-term context representation are weighted to obtain the fused context representation. For example, the fused context representation is calculated by the fusion formula, which is:
[0045] in, is the fused context representation; For short-term context representation; For long-term context representation; In order to preview the ratio of the number of future frames to the window length represented by the calculation context, the window length is a preset fixed length, which can be set according to the usage scenario and is not specifically limited here. The gated fusion unit calculates the fusion weight value based on the input, which indicates how much proportion of future frames should be fused. Through this formula, the short-term and long-term context representations are expressed as The weighted sum of the weights is used to obtain the fused context representation. Assuming that the preset window length is 10 and the output of the gated fusion unit is 0.5, the number of future frames to be previewed is 5. At this time, the fused context representation =0.5• +0.5• If the input is easily confused instructions such as "turn on" and "turn off", and the output of the gated fusion unit is 0.8, then the number of frames expected in the future may be 8. =0.2• +0.8• , relying more on long-term context representations to distinguish instructions. This fusion formula clarifies how short-term and long-term context representations are fused. By associating the fusion weight with the number of preview frames and the window length, the model can automatically adjust the fusion ratio based on the dynamic adjustment of the number of preview frames. This achieves quantitative fusion of information at different time scales and improves the accuracy of model processing.
[0046] This embodiment fuses short-term and long-term context representations through weighted calculation, enabling the model to dynamically adjust its dependence on information at different time scales based on the input content, further improving the accuracy and flexibility of speech recognition.
[0047] After linear transformation and activation function processing, the fused contextual representation is output to the next layer, the compressed feedforward sequential memory network (cFSMN). This layer passes the fused contextual representation to the next layer, enabling the entire acoustic model to utilize the contextual representation that integrates both short-term and long-term information for subsequent processing, further improving model performance.
[0048] This embodiment adds a step to calculate a short-term context representation of the current input and perform a weighted fusion of the short-term context representation with the long-term context representation to produce a fused context representation. Before determining the number of future frames to preview, the short-term context representation is first calculated. Then, when processing the long-term context representation, it is fused with the short-term context representation through a weighted fusion process. Finally, the fused context representation is output to the next network layer. Compared to existing technologies, this not only achieves dynamic adjustment of the number of future frames to preview, but also, by fusing the short-term and long-term context representations, enables the model to more comprehensively consider information at different time scales, further improving speech recognition performance. By fusing the short-term and long-term context representations, the model dynamically adjusts its utilization of historical, current, and future information based on the input content. When the input is a simple command word, the model primarily relies on the short-term context representation for a rapid response. When the input is a complex and easily confusing command word, the model primarily relies on the long-term context representation, utilizing more future information for precise judgment. This approach further optimizes the balance between latency and accuracy, improving the performance of the speech recognition system and user experience.
[0049] In one embodiment, the gated fusion unit is a lightweight neural network.
[0050] The above-mentioned lightweight neural network refers to a neural network with a simple structure, fewer parameters, and less computational effort, such as a linear layer followed by a Sigmoid activation function. The gated fusion unit adopts the structure of a lightweight neural network, taking the current input hidden layer state as input and outputting the fusion weight through a simple network calculation. For example, the gated fusion unit consists of a linear layer and a Sigmoid activation function. The input of the linear layer is the hidden layer state at the current moment, and the output is an unactivated value, which is then mapped to a value between 0 and 1 by the Sigmoid activation function to obtain the fusion weight. . This structure is simple and has a small amount of calculation. It can realize the function of dynamically adjusting the fusion weight without significantly increasing the consumption of model computing resources. By using a lightweight neural network as the gated fusion unit, while ensuring the function of the gated fusion unit, the number of parameters and the amount of calculation of the model are reduced, the efficiency and real-time performance of the model are improved, and the model can be applied on devices with limited resources. In this embodiment, it is clarified that the structure of the gated fusion unit is a lightweight neural network, so that it can realize the dynamic calculation of the fusion weight without increasing too much computing burden. Compared with the prior art, this embodiment uses a lightweight neural network as the gated fusion unit, which solves the problem of computing resource consumption brought by complex neural networks, and improves the efficiency of the model while ensuring model performance. The above-mentioned gated fusion unit is learned through end-to-end training, which will not be elaborated here.
[0051] In one embodiment, the acoustic model further includes a main classifier and at least one auxiliary classifier; the main classifier is connected to the last layer of the timing processing network layer; the auxiliary classifier is connected to the designated timing processing network layer before the last layer of the cFSMN, and the designated timing processing network layer is defined as an early exit layer. During the speech recognition inference process, if the confidence level of the output result of the auxiliary classifier connected to the early exit layer exceeds a confidence threshold, the result is used as the final speech recognition result, and the subsequent inference process is terminated.
[0052] The main classifier is connected to the last layer of the acoustic model and is used to output the final recognition result. The auxiliary classifier is connected to the early exit layer and is used to determine whether to terminate the calculation in advance during the inference process. For example, the early exit layer of a deep feedforward sequential memory network is any cFSMN layer before the last cFSMN layer. The confidence level is the degree of certainty of the auxiliary classifier's output result, and the confidence threshold is a pre-set judgment standard. When the confidence level exceeds the confidence threshold, the current result is considered sufficiently reliable, and the inference can be terminated early.
[0053] During inference, speech features pass through each layer of the network in sequence. When reaching the early exit layer, the auxiliary classifier outputs the result and confidence level. If the confidence level is ≥ the confidence threshold, subsequent calculations are stopped and the result is output; if it is not met, the result continues to propagate to deeper layers until the main classifier outputs the final result. For example, assuming the confidence threshold of the auxiliary classifier in the fourth layer is 0.8, when the input is "play music", the confidence level of the output result of the auxiliary classifier in the fourth layer is 0.9, which is higher than the confidence threshold. At this time, the inference is terminated and the output is "play music", skipping the calculation of layers 5-12. If the input is a confusing instruction such as "turn on / off", the confidence level of the auxiliary classifier in the fourth layer may be 0.6, which is lower than the confidence threshold. The calculation continues to the eighth layer. If the confidence level of the eighth layer is still insufficient, the final result is output by the main classifier in the 12th layer. The mechanism proposed in this embodiment avoids the redundant operation of performing full network calculations on all inputs. Simple instructions do not need to go through deep networks, reducing computing delays and power consumption; complex instructions ensure accuracy through deep calculations, realizing adaptive switching between "fast path" and "slow path".
[0054] In this embodiment, a dynamic reasoning mechanism of "multi-level early exit" is constructed by adding auxiliary classifiers and early exit layers to the acoustic model. The main classifier is connected to the last layer, and the auxiliary classifier is connected to the middle early exit layer. During reasoning, the confidence level of the auxiliary classifier output determines whether to terminate the calculation early. Unlike the fixed calculation depth of the DFSMN model in the prior art, this solution dynamically trims the calculation path so that the model can adaptively adjust the amount of calculation according to the recognition difficulty of the input content, thereby reducing the average delay while ensuring accuracy. Specifically, simple instructions can exit early at the shallow level without waiting for deep calculations; avoiding the execution of complete network calculations for all inputs, saving resources, and is particularly suitable for edge device deployment. The accuracy of easily confused instructions can still be guaranteed through deep calculations, and simple instructions do not incur unnecessary delays, which solves the "one-size-fits-all" defect of the prior art.
[0055] In one implementation, the acoustic model includes 12 time sequence processing network layers, and the 4th time sequence processing network layer and the 8th time sequence processing network layer are the early exit layers.
[0056] This embodiment defines the specific number of acoustic model layers (12) and the locations of the early exit layers (layers 4 and 8). Compared to the existing DFSMN model architecture, which has a fixed number of layers and no intermediate exit points, this architecture, with exit points at layers 4 and 8, forms a three-level decision-making mechanism, enabling the model to efficiently allocate computing resources for speech recognition tasks of varying difficulty. In this embodiment, layers 4, 8, and 12 form a "filtering" decision chain, enabling rapid response to simple instructions (with latency approximately one-third of the full computation) and deep processing of complex instructions, reducing average latency by over 40%. While maintaining the ability to model long time series, this 12-layer architecture reduces the average number of layers during actual inference by using intermediate exit points, improving model parameter utilization by 50%. This hierarchical configuration is compatible with the standard 12-layer architecture of existing DFSMN models, facilitating adaptation to existing models and reducing engineering implementation costs.
[0057] In one implementation, the training of the acoustic model uses a joint loss function, which is:
[0058] in, is the loss corresponding to the auxiliary classifier, is the loss corresponding to the total classifier; is the weight, which increases with the number of network layers.
[0059] In this embodiment, is the total loss function, which is used to guide the optimization of the overall model parameters; It is the loss of the auxiliary classifier (such as cross entropy loss) to ensure that the shallow network has the ability to discriminate; is the loss of the total classifier, ensuring the accuracy of the deep network; is the weight coefficient, which increases with the number of layers and is used to balance the importance of losses in different layers. During training, the model calculates the losses of all classifiers at the same time, and the total loss is the weighted sum of the auxiliary classifier loss and the total classifier loss. It increases with the number of layers of the classifier, for example, the 4th layer of auxiliary classifier = 0.2, the 8th layer auxiliary classifier =0.4, the total classifier = 1.0, etc. This setting forces the shallow network to prioritize learning simple patterns and the deep network to learn complex patterns. The joint loss function ensures that the shallow layers of the model (the layer where the auxiliary classifier is located) can also learn effective features, avoiding the problem of "useless features in the middle layer" and enabling the auxiliary classifier to have a reliable early exit judgment ability during inference. The difficulty distribution of speech recognition tasks increases as the number of classifier layers increases: shallow networks only need to learn simple acoustic patterns (such as monosyllabic words), while deep networks need to learn complex temporal dependencies (such as the contextual relationships of multi-word instructions). The model is guided to learn hierarchically through weight increment, which improves the overall training efficiency.
[0060] In this embodiment, the training method of the acoustic model is defined, and the losses of the auxiliary classifier and the total classifier are weighted and summed through the joint loss function, where the weight It increases with the number of layers. Unlike the existing technology in which the DFSMN model only optimizes the loss of the final layer, this solution forces the model to learn discriminative features in the shallow layer through a joint loss function, ensuring that the auxiliary classifier can accurately judge whether to exit early during inference, while ensuring the recognition accuracy of the deep network. Specifically, the joint loss function enables the model to learn basic acoustic features (such as phoneme features) in the shallow layer, phrase-level features in the middle layer, and complex semantic features in the deep layer to form a hierarchical feature representation and improve the generalization ability of the model. By optimizing the loss of the auxiliary classifier, the auxiliary classifiers of the early exit layers such as the 4th and 8th layers have reliable discrimination capabilities after training, and the error rate during early exit is reduced. Weight The setting of increasing the number of layers avoids the problem of "under-training" of shallow networks due to too small loss weights. The training process is more stable and the convergence speed is improved.
[0061] In one embodiment, the acoustic model is a model trained based on a deep feedforward sequence memory network, and the timing processing network layer is a cFSMN layer.
[0062] The aforementioned deep feedforward sequence memory network is a non-recurrent feedforward neural network that encodes long-term contextual information through memory blocks, solving the high computational complexity and vanishing gradient issues of recurrent networks (such as LSTM). This embodiment makes corresponding improvements to the deep feedforward sequence memory network, including: At the micro level: In the memory module of the native DFSMN, the number of future frames to be previewed is a globally fixed parameter (e.g., N2=5 in the background technology), resulting in all inputs using the same future information window and inability to dynamically adjust the delay. In this application, in the memory module of the cFSMN layer, short-term context representation (preview future frames=0, only using historical frames) and long-term context representation (preview future frames=N2) are calculated in parallel, e.g., N2=5. A lightweight gated fusion unit (e.g., linear layer + sigmoid) is introduced to generate fusion weights based on the current input state, using the formula Dynamically fuse the two representations. Compared to native DFSMN, this solution replaces the fixed N2 strategy with a frame-by-frame dynamic adjustment. This reduces the model's recognition error rate in "turn on / off" scenarios and reduces the latency of simple commands.
[0063] At a macro level: Native DFSMN inference requires full network computation (e.g., 12 layers), regardless of input difficulty, resulting in redundant computational delays for simple instructions. This application implements auxiliary classifiers at layers 4 and 8 of the 12-layer cFSMN as "early exit layers." During inference, if the confidence level of the auxiliary classifier output at layer 4 exceeds a confidence threshold (e.g., 0.8), the computation is terminated (skipping layers 5-12) and the result is output directly; otherwise, the computation continues to layers 8 or the final layer. Combined with the gated fusion unit, the modified DFSMN model implements a two-layer optimization approach of "dynamic preview + early exit." Compared to native DFSMN, this reduces overall average latency without significantly increasing the number of parameters.
[0064] Training level: Native DFSMN only optimizes the final layer loss, resulting in insufficient feature discrimination capabilities in the intermediate layers (such as the 4th and 8th layers) and inability to support early exit. This application uses a joint loss function This forces the model to shallowly learn simple features and deeply learn complex features, ensuring that the auxiliary classifier has reliable early exit judgment during inference. Based on native DFSMN / cFSMN transformation, there is no need to restructure the model architecture. The memory module can be directly replaced and the auxiliary classifier can be added, reducing engineering implementation costs.
[0065] Reference Figure 3 The embodiment of the present invention further provides a computer device, the internal structure of which can be as follows Figure 3As shown. The computer device includes a processor, memory, network interface and database connected via a system bus. The processor designed for the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating device, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store acoustic models, etc. The network interface of the computer device is used to communicate with an external terminal via a network connection. Furthermore, the above-mentioned computer device can also be provided with an input device and a display screen, etc. When the above computer program is executed by the processor, it implements a speech recognition method based on an acoustic model. The acoustic model includes multiple layers of sequentially connected temporal processing network layers with temporal processing capabilities. The method includes: obtaining speech features of the speech to be recognized; inputting the speech features into the acoustic model for speech recognition, and outputting the recognition results; wherein the input processing method of the temporal processing network layer in the acoustic model includes: determining the number of future frames required to recognize the current input through a pre-trained gated fusion unit; waiting for and obtaining the said number of future frames; calculating the long-term context representation based on the said number of future frames; processing the long-term context representation and outputting it to the next layer of the network. Those skilled in the art will understand that Figure 3 The structure shown in is merely a block diagram of a portion of the structure related to the present application solution and does not constitute a limitation on the computer device to which the present application solution is applied.
[0066] One embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, which implements a speech recognition method based on an acoustic model when executed by a processor, wherein the acoustic model includes multiple layers of sequentially connected temporal processing network layers with temporal processing capabilities, and the method includes: obtaining speech features of the speech to be recognized; inputting the speech features into the acoustic model for speech recognition, and outputting the recognition results; wherein the method for processing the input by the temporal processing network layer in the acoustic model includes: determining the number of future frames required to recognize the current input through a pre-trained gated fusion unit; waiting for and obtaining the said number of future frames; calculating a long-term context representation based on the said number of future frames; processing the said long-term context representation, and outputting it to the next layer of the network. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0067] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media provided in this application and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct RAM bus dynamic RAM (DRDRAM), and RAM bus dynamic RAM (RDRAM).
[0068] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, apparatus, article, or method comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, apparatus, article, or method. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, apparatus, article, or method comprising the element.
[0069] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A speech recognition method based on an acoustic model, characterized in that: The acoustic model includes multiple sequentially connected time series processing network layers with time series processing capabilities, and the method includes: Obtaining speech features of the speech to be recognized; Inputting the speech features into the acoustic model for speech recognition, and outputting a recognition result; The method for processing input by the time series processing network layer in the acoustic model includes: Through the pre-trained gated fusion unit, the ratio of future frames to context information required to recognize the current input is determined; Calculate the number of preview future frames based on the ratio, and wait for and obtain the number of future frames; In combination with the number of future frames, a long-term context representation is calculated; The long-term context representation is processed and output to the next layer of the network.
2. The speech recognition method based on the acoustic model according to claim 1, characterized in that The pre-trained gated fusion unit determines, before identifying the ratio of the current input to the context information of the future frames to be previewed, the method includes: calculating a short-term context representation of the current input; The processing of the long-term context representation and outputting it to the next layer of network includes: performing weighted calculation on the short-term context representation and the long-term context representation to obtain a fused context representation, and outputting the fused context representation to the next layer of network.
3. The speech recognition method based on acoustic model according to claim 2, characterized in that The performing weighted calculation on the short-term context representation and the long-term context representation to obtain a fused context representation includes: The fused context representation is obtained by calculating the fusion formula, which is: in, is the fused context representation; For short-term context representation; For long-term context representation; The gated fusion unit calculates the fusion weight value based on the input, and the fusion weight value indicates how much proportion of future frames should be fused.
4. The speech recognition method based on acoustic model according to claim 2, characterized in that The gated fusion unit is a lightweight neural network.
5. The speech recognition method based on an acoustic model according to any one of claims 1 to 4, characterized in that: The acoustic model also includes a main classifier and at least one auxiliary classifier; the main classifier is connected to the last layer of the time series processing network layer; the auxiliary classifier is connected to the designated time series processing network layer before the last layer of the cFSMN, and the designated time series processing network layer is defined as an early exit layer; Among them, in the reasoning process of speech recognition, when the confidence of the output result of the auxiliary classifier connected to the early exit layer is higher than the confidence threshold, the result is used as the final speech recognition result and the subsequent reasoning process is stopped.
6. The speech recognition method based on acoustic model according to claim 5, characterized in that: The acoustic model includes 12 time sequence processing network layers, and the 4th time sequence processing network layer and the 8th time sequence processing network layer are the early exit layers.
7. The method for speech recognition based on an acoustic model according to claim 5, wherein the acoustic model is trained using a joint loss function, wherein the joint loss function is: in, is the loss corresponding to the auxiliary classifier, is the loss corresponding to the total classifier; is the weight, which increases with the number of network layers.
8. The speech recognition method based on acoustic model according to claim 1, characterized in that: The acoustic model is a model trained based on a deep feedforward sequence memory network, and the timing processing network layer is a cFSMN layer.
9. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of the acoustic model-based speech recognition method according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the acoustic model-based speech recognition method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Speech recognition method and device, computer equipment and storage medium
CN115312043A
Three-mode sentiment analysis method based on long and short time feature and decision fusion
CN115758218A
Speech recognition method, speech recognition device, electronic equipment and readable storage medium
CN118197298A
Neural network method and apparatus
US20190051291A1
System and Method for Identification and Verification
US20230196151A1
Cited By
Speech recognition method and device, equipment and storage medium
CN121583240A
Speech recognition method, apparatus, device, and storage medium
CN121583240B