Speech Encoding Pretraining for Causal Streaming Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current pre-training methods for speech models result in poor performance in speech processing tasks, particularly in scenarios requiring causality, such as streaming speech processing.
Innovation Solution
A method and apparatus for speech processing that involves acquiring a speech feature sequence, generating subsequent speech tokens based on preceding speech features using a speech encoding model, and training the model on these tokens to achieve causality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If current pre-training methods are used for speech models, then the model can be trained efficiently, but the performance in speech processing tasks (especially streaming speech processing) deteriorates
Solution Approach 1:
The patent inverts the traditional training approach by training the speech encoding model to predict future speech tokens rather than reconstructing current or past tokens. This causal prediction direction (from past to future) aligns with the requirements of streaming speech processing, resolving the contradiction between training efficiency and task performance
Solution Approach 2:
The patent changes the training objective parameter from reconstruction error to prediction error of future tokens. By modifying the loss function to measure the difference between predicted future tokens and actual future tokens, the model learns causal relationships that improve streaming speech processing performance while maintaining training feasibility
2Adaptability or versatility
If traditional speech model training is applied, then the training process is straightforward, but the model fails to capture causal relationships in speech sequences
Solution Approach 1:
The patent applies preliminary action by pre-defining the causal structure in the training objective, where the model is explicitly trained to predict future tokens based on past and present features. This preliminary structuring of causal relationships during training enables the model to naturally capture causality without adding complex architectural components
Solution Approach 2:
The patent implements feedback mechanisms through the loss function that compares predicted future tokens with actual future tokens. This feedback loop guides the model to continuously improve its causal predictions, enhancing adaptability while keeping the training approach manageable through standard optimization techniques
Data Source
AI summary
According to an embodiment of the disclosure, a method, apparatus, device and computer-readable storage medium for speech processing are provided. The method includes: acquiring a speech feature sequence corresponding to a speech sample, the speech feature in the speech feature sequence corresponding to a speech frame in the speech sample. For the target speech feature in the speech feature sequence, one or more subsequent speech tokens respectively corresponding to the one or more subsequent speech features are generated based on the target speech feature and the one or more preceding speech features and according to the speech encoding model. A speech encoding model is trained based on the one or more subsequent speech tokens. This training approach enables the speech encoding model to learn high-quality speech representations.


