Efficient simultaneous interpretation method based on expert routing threshold
By constructing a hybrid expert routing strategy model based on expert routing thresholds, the limitations of computational efficiency and multilingual support in streaming speech translation are resolved, high efficiency and low latency of multilingual streaming translation are achieved, and the application value of the model is enhanced.
Patent Information
- Application Number
- CN202510940032.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-10
AI Technical Summary
Existing technologies in streaming speech translation have limitations in computational efficiency, multilingual support, and training costs, and lack a simple, efficient, and explainable reading and writing strategy learning solution.
A hybrid expert routing strategy model is constructed using an expert routing threshold-based method, including a streaming speech encoder, a text decoder, a routing threshold module, and a hybrid expert post-processing module. Multilingual streaming translation is achieved through training, and the expert routing strategy is used to decide whether to output or wait for the reading of audio clips.
It achieves high efficiency and quality of multi-language streaming translation, maintains low latency and multi-task training capabilities, and improves the application value and computing efficiency of the model.
Smart Images

Figure CN120766657A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of speech processing, and in particular to an efficient simultaneous interpretation method based on expert routing threshold. BACKGROUND
[0002] With the continuous progress of deep learning technology, especially the wide application of sequence generation models based on Transformer architecture in the field of speech processing, end-to-end speech translation systems have made significant breakthroughs. Such systems can directly convert source language speech signals into target language text without the complex multi-stage processing flow in traditional systems. Current technology has made conventional offline translation tasks practical, and academic research is shifting to the more challenging field of streaming speech translation (SST), i.e., real-time translation systems that can achieve the effect of simultaneous interpretation at international conferences.
[0003] In the streaming speech translation architecture, the system must simultaneously process two core tasks: first, the understanding and encoding of speech content, and second, the real-time output decision of the translation result. This real-time requirement requires the system to dynamically determine whether the currently acquired speech data is sufficient to generate a translation result with a confidence threshold while continuously receiving speech input. This "when to speak" decision mechanism directly affects the overall performance of the system, constituting a key balance point between latency and translation quality. Therefore, the design and optimization of the decision strategy module is a core technical link in the development of streaming systems.
[0004] The decision strategy system includes two interrelated technical dimensions: reading policy for input speech and output policy for output results. The core task of the reading policy is to perform semantic segmentation of continuous speech signals, and the output policy determines whether to generate a translation result and the scope of the generated content based on the current processing state.
[0005] In the speech input processing link, due to the lack of clear physical boundary markers in natural speech, the segmentation process faces significant challenges. Existing mainstream segmentation strategies mainly include three technical paths:
[0006] Fixed-length segmentation strategy: equal-length segmentation is performed by presetting a time window length (e.g., 280 milliseconds) or a fixed number of frames (e.g., 400 frames). This method is computationally efficient but may disrupt the natural semantic integrity of speech.
[0007] ●Word-boundary-based Segmentation: This strategy uses phonemes or word boundaries detected by the speech recognition system to perform aligned segmentation, which can better maintain the continuity of acoustic features compared to fixed-length methods.
[0008] ●Semantic-aware Segmentation: Segmentation points are determined by analyzing the semantic features of speech. Some advanced methods integrate segmentation decisions into the model training process to achieve end-to-end optimization.
[0009] There are more diverse options for obtaining output strategies, including but not limited to:
[0010] 1) Wait-k Policy: Wait for a certain period of time before starting output. The output quantity each time is a fixed ratio of the input quantity. This is the most basic delay control method.
[0011] 2) Monotonic Attention Method: A method that monotonically models the probability of whether a model needs to be generated and optimizes its posterior probability. The disadvantage is that it requires calculating the expectation during training, which is computationally intensive.
[0012] 3) Connectionist Temporal Classification (CTC) alignment-assisted method: This method uses the CTCL algorithm to establish an input-output alignment relationship to establish an output strategy. However, its disadvantage is that one model can only process one language pair.
[0013] 4) Attention alignment auxiliary method: Directly using the attention weights in the transformer as the basis for alignment is based on empirical conclusions and is unreliable.
[0014] 5) Neural Network Transducer Architecture: This alignment method has stronger sequence modeling capabilities, but it also faces the problem of large training computational load and its performance ceiling is not as good as the transformer architecture.
[0015] Therefore, those skilled in the art are committed to developing an efficient simultaneous interpretation method based on expert routing threshold. Summary of the Invention
[0016] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is that the existing technical solutions have limitations in terms of computational efficiency, multilingual support and training costs, and lack a simple, efficient and explainable reading and writing strategy learning solution.
[0017] To achieve the above objectives, the present invention provides an efficient simultaneous interpretation method based on expert routing thresholds. The method is based on the classic Transformer architecture model and constructs an expert routing strategy model based on expert routing thresholds. By training the expert routing strategy model, multilingual streaming translation is achieved. The expert routing strategy model includes:
[0018] A streaming speech encoder, using a Transformer architecture, is compatible with traditional speech-to-text translation systems. The streaming speech encoder adopts a hybrid design consisting of block-by-block autoregressive blocks and non-autoregressive blocks.
[0019] The text decoder follows the standard autoregressive Transformer architecture and processes both the complete offline speech and the randomly truncated speech prefix to generate the offline hidden state and the prefix hidden state.
[0020] a routing threshold module implemented by a two-layer feed-forward network with a sigmoid head, wherein the routing threshold module projects the final hidden state of the text decoder into a scalar value and uses the scalar value to determine the expert weight;
[0021] A hybrid expert post-processing module adopts a Transformer-like architecture and shares a language model head with the text decoder. The hybrid expert post-processing module combines prefix information and global information to predict the target translation sequence and utilizes the output result of the routing threshold module.
[0022] Furthermore, the streaming speech encoder uses 20 of the block-wise autoregressive blocks Chunk-AR and 4 of the non-autoregressive blocks NAR, wherein:
[0023] The Chunk-AR only calculates new chunks and optimizes inference efficiency through a cache key-value mechanism;
[0024] The NAR recomputes the entire sequence, maintains translation quality by capturing global context, and adds a learnable end-of-stream flag in front of the NAR input, which is configured to indicate whether the current block terminates the audio stream.
[0025] Furthermore, the hybrid expert post-processing module comprises a plurality of blocks, each of which comprises a dual expert module, a pre-order output attention module and a feed-forward network module, wherein:
[0026] The dual-expert module includes a prefix expert and a global expert. The global expert is configured as a two-layer MLP module. The prefix expert adopts a standard cross-attention mechanism. The combined weight of the prefix expert and the global expert is used to determine whether the current input prefix contains sufficient information required to generate the target tag.
[0027] The pre-order output attention module is configured to strictly isolate global information, no longer pay attention to the hidden state in the same layer, and only pay attention to the pre-order output of the decoder to prevent global information leakage.
[0028] Furthermore, the final output of the hybrid expert post-processing module is combined with the contribution of the dual expert module through a gated residual connection, and the specific result is:
[0029]
[0030] in, is the output of the i-th position, is the input at position i, is the global expert input for position i, is the prefix expert input for position i, is the global expert weight of the i-th position, is the prefix expert weight of the i-th position, h offline For offline hidden state, H global is a global embedding, is the projection weight, [,] is the vector concatenation, and MHA is the multi-head attention.
[0031] Furthermore, the expert routing strategy model is based on a pre-trained offline speech-to-text translation S2TT model, and the training of the expert routing strategy model is achieved through an offline pre-training stage and a synchronous training stage.
[0032] Furthermore, in the offline pre-training stage, the standard offline S2TT target offline loss is used Train the model until convergence.
[0033] Furthermore, in the synchronous training phase, combined with the offline loss Prefix loss and post-mixing losses The total loss of the model during the synchronous training phase is:
[0034]
[0035] in, is the total loss, Offline loss, is the prefix loss, is the post-mixing loss, ω r ,ω p is the weight value, p ref is the output distribution of the hybrid expert refiner, p dec is the output distribution of the text decoder, λ is the hyperparameter, h prefixHide the prefix status.
[0036] Furthermore, during the streaming inference process, the hybrid expert post-processing module uses an expert routing strategy to decide whether to continue outputting or wait and read subsequent audio segments. The expert routing strategy is configured as follows:
[0037] Determine whether the routing threshold score output by the autoregressive decoding is lower than the preset threshold. If it is lower than the threshold, continue to output; otherwise, wait and read the subsequent audio clip:
[0038]
[0039] Among them, p t,i is the routing threshold score at input time t and target location i, and ε is the threshold.
[0040] Furthermore, the hybrid expert post-processing module uses a normalization method to align the average score of the routing gate with the relative information content between the prefix and the global context. The normalization method calculates a normalized loss based on the prefix length and the full sequence length. The specific calculation method of the normalized loss is:
[0041]
[0042] in, is the normalized loss, l p is the prefix length, l g is the complete sequence length, l b is the buffer hyperparameter, is the average routed gate output of the sequence.
[0043] Furthermore, during the training of the expert routing strategy model, the normalization method is combined to control the mean and variance of the routing gate score, and the complete training loss value target is:
[0044]
[0045] in, is the training loss value, is the total loss during the synchronous training phase, is the normalized loss, ω n is the normalized weight.
[0046] In a preferred embodiment of the present invention, compared with the prior art, the present invention has the following beneficial effects:
[0047] 1. This invention adopts a hybrid expert MoE threshold solution to learn strategies, giving full play to the self-learning ability of neural networks. End-to-end training automatically mines the alignment dependency information contained in the data to form a reliable streaming translation strategy.
[0048] 2. The present invention trains offline translation in a multi-task manner while training the streaming read-write strategy. Since the training of the hybrid expert threshold does not conflict with the offline training tasks, end-to-end multi-task training can be performed to ensure that its offline translation capability is not reduced.
[0049] 3. This invention applies the mixed expert MoE threshold solution to streaming TTS, achieving low-latency streaming TTS at the word level. It cooperates with the translation model to open up the link from speech to speech translation, maintains low latency, expands the application value of the model, and has better sound quality than other existing solutions.
[0050] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the accompanying drawings to fully understand the purpose, characteristics and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 1 is a schematic diagram of an expert routing strategy model architecture based on an expert routing threshold according to a preferred embodiment of the present invention;
[0052] Figure 2 This is a schematic diagram of the expert routing strategy flow reasoning process of a preferred embodiment of the present invention;
[0053] Figure 3 It is a schematic diagram of the architecture of a hybrid expert post-processing module of a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0054] The following describes several preferred embodiments of the present invention with reference to the accompanying drawings to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.
[0055] In the drawings, components with identical structures are denoted by the same reference numerals, and components with similar structures or functions are denoted by similar reference numerals. The size and thickness of each component shown in the drawings are arbitrary and are not limited by the present invention. For clarity, the thickness of components in some places in the drawings is appropriately exaggerated.
[0056] like Figures 1-3As shown, the present invention addresses the limitations of existing solutions in terms of computational efficiency, multilingual support, and training costs. By adopting an innovative hybrid expert-based output strategy training framework, the present invention offers a novel solution with low training costs, integrated multilingual translation capabilities, and strong performance, achieving currently optimal comprehensive performance indicators. Furthermore, the present invention is not limited to streaming speech translation but can also be applied to streaming speech-to-text (streaming TTS), making it even more widely applicable.
[0057] like Figure 1 As shown, an embodiment of the present invention provides an efficient simultaneous interpretation method based on expert routing threshold. Based on the classic Transformer architecture model, an expert routing strategy model based on expert routing threshold is constructed. By training the expert routing strategy model, multi-language streaming translation is achieved.
[0058] The expert routing strategy model provided by the embodiment of the present invention includes:
[0059] 1) Streaming speech encoder
[0060] It adopts the Transformer architecture to maintain compatibility with traditional speech-to-text translation systems; the streaming speech encoder adopts a hybrid design consisting of block-by-block autoregressive blocks and non-autoregressive blocks.
[0061] In this embodiment, the streaming speech encoder uses 20 block-by-block autoregressive blocks Chunk-AR and 4 non-autoregressive blocks NAR, where:
[0062] Chunk-AR only calculates new chunks and optimizes inference efficiency through a cache key-value mechanism;
[0063] NAR recomputes the entire sequence to maintain translation quality by capturing global context and adds a learnable end-of-stream flag before the NAR input to indicate whether the current block terminates the audio stream.
[0064] 2) Text decoder
[0065] Following the standard autoregressive Transformer architecture, the complete offline speech and the randomly truncated speech prefix are processed simultaneously to generate the offline hidden state and the prefix hidden state.
[0066] 3) Routing threshold module
[0067] It is implemented by a two-layer feed-forward network with a Sigmoid head, which projects the final hidden state of the text decoder into a scalar value and uses the scalar value to determine the expert weight.
[0068] 4) Hybrid expert post-processing module
[0069] It adopts a Transformer-like architecture and shares the language model head with the text decoder. The hybrid expert post-processing module combines prefix information and global information to predict the target translation sequence and utilizes the output of the routing gate module.
[0070] In this embodiment, the hybrid expert post-processing module includes multiple blocks, each of which includes a dual expert module, a pre-order output attention module, and a feedforward network module, wherein:
[0071] Dual-expert module: This module includes a prefix expert and a global expert. The global expert is configured as a two-layer MLP module. The prefix expert uses a standard cross-attention mechanism. The combined weight of the prefix expert and the global expert is used to determine whether the current input prefix contains sufficient information required to generate the target token.
[0072] Pre-order output attention module: It is configured to strictly isolate global information, no longer pay attention to the hidden state within the same layer, and only focus on the pre-order output of the decoder to prevent global information leakage.
[0073] The final output of the hybrid expert post-processing module is combined with the contribution of the dual expert module through a gated residual connection. The specific results are:
[0074]
[0075] in, is the output of the i-th position, is the input at position i, is the global expert input for position i, is the prefix expert input for position i, is the global expert weight of the i-th position, is the prefix expert weight of the i-th position, h offline For offline hidden state, H global is a global embedding, is the projection weight, [,] is the vector concatenation, and MHA is the multi-head attention.
[0076] In this embodiment, the expert routing strategy model is based on a pre-trained offline speech-to-text translation S2TT model, and the training of the expert routing strategy model is achieved through an offline pre-training stage and a synchronous training stage.
[0077] In the offline pre-training stage, the standard offline S2TT target offline loss is used. Train the model until convergence.
[0078] In the synchronous training phase, combined with offline loss Prefix loss and post-mixing losses The model is trained, and the total loss in the synchronous training stage is:
[0079]
[0080]
[0081] wherein, is the total loss, is the offline loss, is the prefix loss, is the post-processing loss, ω r ,ω p is a weight value, p ref is the output distribution of the hybrid expert refiner, p dec is the output distribution of the text decoder, λ is a hyperparameter, h prefix is the prefix hidden state.
[0082] In the streaming inference process, the hybrid expert post-processing module adopts an expert routing strategy to determine whether to continue outputting or waiting and reading the subsequent audio segment, and the expert routing strategy is configured as:
[0083] The routing threshold score of the autoregressive decoding output is determined, and if the routing threshold score is lower than a preset threshold, the output is continued; otherwise, the subsequent audio segment is waited and read:
[0084]
[0085] wherein, p t,i is the routing threshold score at input time t and target position i, and ε is a threshold value.
[0086] In this embodiment, the hybrid expert post-processing module aligns the average score of the routing gate with the relative information amount between the prefix and the global context by using a normalization method, which calculates a normalization loss based on the prefix length and the complete sequence length. The specific calculation method of the normalization loss is:
[0087]
[0088] wherein, is the normalization loss, l p is the prefix length, l g is the complete sequence length, l b is a buffer hyperparameter, is the average routing gate output of the sequence.
[0089] In the training process of the expert routing strategy model, the mean and variance of the routing gate score are controlled in combination with the normalization method, and the complete training loss value target is:
[0090]
[0091] in, is the training loss value, is the total loss during the synchronous training phase, is the normalized loss, ω n is the normalized weight.
[0092] The efficient simultaneous interpretation method based on expert routing threshold provided by the embodiment of the present invention has the following technical features compared with the prior art:
[0093] 1. In view of the lack of a simple, efficient, and explainable reading and writing strategy learning solution in existing streaming speech translation solutions, the present invention adopts a mixed expert MoE threshold solution to learn strategies. By setting up two experts, one translates based on prefix input (streaming input) and the other translates based on fuzzified global input (offline input) information. According to the principle of mixed expert threshold, the model automatically assigns different weights to the two input information, and this weight can be used as the basis for strategic decision-making. The present invention fully utilizes the self-learning ability of neural networks, and end-to-end training automatically mines the alignment dependency information contained in the data to form a reliable streaming translation strategy.
[0094] 2. Existing streaming speech translation solutions based on the Transformer architecture often have a large performance gap compared to offline translation systems. The present invention trains offline translation in a multi-task manner while training streaming read and write strategies. Since the training of the hybrid expert threshold does not conflict with the offline training tasks, end-to-end multi-task training can be performed to ensure that its offline translation capabilities do not decrease. In the experimental results, compared to the offline model, the streaming model based on the hybrid expert threshold has a performance drop of less than 7% at 1.5 seconds, a performance drop of less than 5% at 2 seconds, and a performance drop of less than 3% at 3 seconds. In comparison, the performance drop of the competitor Seamless at 3 seconds is more than 7%.
[0095] 3. Existing speech-to-speech simultaneous interpretation solutions struggle to balance sound quality and translation quality. This invention applies the hybrid expert MoE threshold approach to streaming TTS, employing the same streaming principles and strategies for autoregressive streaming TTS as for streaming translation. This invention achieves word-level low-latency streaming TTS, and in conjunction with the translation model, it bridges the gap between speech-to-speech translation, maintaining low latency and expanding the model's application value. Compared to existing solutions, this approach offers superior sound quality.
[0096] The present invention will be described in detail below in conjunction with preferred embodiments of the present invention.
[0097] The term "training" or "learning" used in the present invention refers to the process of updating the configuration parameters, optimizing the system performance by utilizing experience or training data. For example, the translation system can gradually optimize the translation performance, such as improving the translation accuracy, through the training or learning process. The training or learning process can be ended based on certain convergence conditions, and the terms "training" or "learning" can be used interchangeably. The term "derivation" or "inference" used refers to the process of utilizing a model or system trained or having learned capabilities to perform a specific task for real-world data. The term "loss function" refers to a mathematical formula used to calculate the gap between the "model output" and the "true answer", which is referred to as the loss value. The process of model training is to make the score of the loss function lower and lower. In model training, the model inputs data, the model gives an answer, and then the loss value is calculated by the loss function, and then the model adjusts its parameters according to the loss value, and this process is repeated continuously. The cross-entropy (Cross Entropy) loss function is used in the present invention.
[0098] In traditional translation models, different language corpora (texts, speech) are mapped to high-dimensional vectors, and after a series of calculations, the vectors are converted to text. Taking a speech-to-text translation system as an example, given the audio frames of the source language X = {x1, x2,..., xM}, and the sentence tokens of the target language Y = {y1, y2,..., yN}, where M and N represent the sequence lengths. Usually, a speech-to-text translation model constructs a conversion from X to Y through a converter. The converter is jointly trained so that the conditional probability of Y' given X' is maximized, and the conditional probability P(Y'|X') can be determined, for example, according to the following formula:
[0099]
[0100] where θ represents the trainable parameters of the model.
[0101] The present invention is mainly based on the classic Transformer architecture model or similar architecture model. For other architecture models, the method of the present invention may not be applicable. The Transformer architecture generally includes an encoder and a decoder (or only a decoder). The encoder inputs the source corpus sequence X and outputs the representation for the decoder to use, and the decoder inputs the historical text token y <i and predicts the next text token y i , and finally generates the target corpus sequence step by step. The term "autoregressive" refers to this way of generating a complete sequence based on the history of the sequence, for example, the generative model of the classic Transformer architecture is an autoregressive model.
[0102] The core principle of synchronous translation training is to enable the model to independently judge whether the input in the currently received speech window is sufficient to generate more output content, and how much content should be output. Previous methods often have limitations, such as relying on manually labeled signals, which will restrict the self-learning ability of the model; or modifying the model architecture, but this will reduce the translation quality. In contrast, the framework of the present invention uses the expert mixed threshold method to achieve unsupervised policy learning while maintaining a structure consistent with the offline model, thereby ensuring that the translation performance is not affected. The prefix-based training strategy adopted by the present invention highly restores the streaming reasoning scenario in the real world.
[0103] The method provided by the embodiment of the present invention achieves high-performance multi-language streaming translation through three key innovations:
[0104] A hybrid-expert post-processing module, used only during training, enables unsupervised policy learning without increasing inference time overhead;
[0105] Minimal modifications to the standard Transformer architecture for broad applicability;
[0106] ●A unified framework that supports speech-to-text and text-to-speech streaming tasks.
[0107] The overall architecture of the expert routing threshold policy framework provided in this embodiment consists of four key components:
[0108] 1. Streaming Voice Encoder
[0109] 2. Text Decoder
[0110] 3. Routing threshold module
[0111] 4. Hybrid Expert Post-Processing Module
[0112] The streaming speech encoder and text decoder adopt the Transformer architecture, which is compatible with traditional speech-to-text translation (S2TT) systems. The encoder adopts a hybrid design, which consists of 20 chunk-by-chunk autoregressive (Chunk-AR) blocks and 4 non-autoregressive (NAR) blocks. After each reading, the Chunk-AR block only calculates the new chunk, while the NAR block recalculates the entire sequence. The Chunk-AR block optimizes inference efficiency by caching the key-value mechanism, while the NAR block maintains translation quality by capturing the global context. In order to handle streaming inputs that lack an explicit end-of-sequence (EOS) marker, this embodiment adds a learnable end-of-stream (EoSt) flag (a binary embedding) before the input of the NAR block to indicate whether the current block terminates the audio stream. The text decoder follows the standard autoregressive Transformer architecture.
[0113] The core innovations of the method provided by the embodiments of the present invention lie in the routing threshold module and the hybrid-of-experts post-processing module, which enable unsupervised learning of synchronous translation strategies without architectural modifications. The MoE post-processing module adopts a Transformer-like architecture and shares the language model head with the text decoder. It combines prefix information and global information to predict the target translation sequence. This refiner is only activated during training and does not introduce additional computational overhead during inference.
[0114] Hybrid expert post-processing module, its architecture is as follows Figure 3 As shown in Figure 2, the routing gate determines the mixed ratio of prefix experts and global experts. This ratio reflects the model's confidence in the prefix sequence, resulting in a natural reading and writing strategy. The self-attention module is replaced by the Prefix Output Attention (POA) module to prevent global information leakage. This module contains N refiner Each block contains two specialized experts: prefix expert (E p ) and global experts (E g ). The combined weights of these experts implicitly define a strategy for deciding whether the current input prefix contains enough information to generate the target token. In addition to the dual-expert module, each block also contains a previous-order output attention module (similar to self-attention) and a feed-forward network module, similar to the standard Transformer decoder block.
[0115] Routing gate module: To ensure the consistency of decision making, each layer of the hybrid expert post-processing module shares the output of a global routing gate module, which is implemented by a two-layer feed-forward network with a Sigmoid head. The routing gate converts the final hidden state (O dec ) is projected into a scalar value p∈[0,1] to determine the expert weight:
[0116]
[0117] Dual expert architecture: Figure 3 As shown, during the training process, the encoder processes the complete offline speech and the randomly truncated prefix at the same time to generate the corresponding hidden state h offline and h prefix The routing gate and the dual expert architecture will then decide which hidden states to use. While the routing gate generally prefers the more informative h offline , but the embodiment of the present invention introduces an information bottleneck to balance this preference. Specifically, for h offline Apply time average pooling, which is E g Generate a compressed global embedding H global .
[0118] Global Expert E g It is a two-layer MLP module, and its calculation process is as follows:
[0119]
[0120] in, represents the input of the i-th position, [·;·] represents vector splicing, are the learnable projection weights (omitting the bias).
[0121] Prefix Expert E p Using the standard cross-attention mechanism:
[0122]
[0123] Where MHA stands for multi-head attention. The final output combines the contributions of the two experts through a gated residual connection:
[0124]
[0125] Preventing global information leakage: In the standard Transformer architecture, the self-attention mechanism inherently leaks global information in the sequence. This poses a key problem for the design: even if the routing gate at a certain position If it is set to 0, the hidden state of this position may still access the global context through self-attention, thus destroying the expected expert specialization. In order to strictly isolate global information, the embodiment of the present invention replaces self-attention with the pre-order output attention mechanism. Each position no longer pays attention to the hidden state in the same layer, but only pays attention to the pre-order output of the decoder. Unnecessary information flow is effectively prevented while preserving sequential dependencies.
[0126] Training framework: The expert routing strategy model is based on the pre-trained offline S2TT model. The training process is divided into two stages, such as Figure 1 As shown, in the first stage, the model is based on (offline loss value) for pre-training. In the second stage, the model is combined with (prefix loss value) and (mixed post-processing loss value) for training:
[0127] 1. Offline pre-training: First use the standard offline S2TT target (Add streaming block masks to the encoder) Train the model until convergence.
[0128] 2. Synchronous training: Then, two additional loss functions are introduced to achieve synchronization capabilities:
[0129]
[0130] Among them, p ref and p dec are the output distributions of the MoE refiner and text decoder, respectively. λ is a predefined hyperparameter used to limit the loss to locations with high confidence. For learning reading and writing strategies, Used to enhance the ability of prefix-based translation. At this stage, the total loss is the weighted sum of the above three losses:
[0131]
[0132] In this embodiment, w is set r =w p =0.2 to prioritize offline training.
[0133] In the streaming reasoning process, Figure 2 As shown, the hybrid expert post-processing module is not required, so its parameter count does not burden inference. At each autoregressive decoding step, a routing threshold score is output. Based on this score, the expert routing strategy adopts a simple threshold-based strategy: when the score is below the threshold, the output continues; when the score is above the threshold, the expert routing strategy waits and reads the next audio segment.
[0134]
[0135] Among them, p t,i is the score of the routing threshold at input time t and target location i.
[0136] The raw output scores of the mixture of experts often lack inherent interpretability because they are simply the result of neural network optimization via gradient descent. These scores exhibit inconsistent statistical properties, both in terms of mean and variance, across different tasks, language pairs, and training datasets. This discrepancy poses significant challenges during inference, requiring thresholds to be adjusted for specific tasks or language pairs, which severely limits the generalizability of the method.
[0137] To solve this problem, the present invention introduces a heuristic method to align the average score of the routing gate with the relative information content between the prefix and the global context. Since the sequence length can be used as a practical proxy for the amount of information, the prefix length (l p ) and the complete sequence length (l g ) formulates a normalized loss:
[0138]
[0139] in, represents the average routing gate output of the sequence, l b is a buffer hyperparameter (set to 1.5 seconds in this example) to prevent abrupt changes in the score near the prefix boundary. The weighting factor of 0.5 takes into account the partial information available at the prefix edge.
[0140] Based on previous work, the present embodiment adds Gaussian noise before Sigmoid activation to promote the discretization of routing gate output. Although this discretization does not directly improve model performance, it significantly enhances the robustness of the model by reducing the sensitivity of different tasks to threshold selection. In this embodiment, zero-mean unit variance Gaussian noise (σ R = 1) is applied to the log-odds before sigmoid activation.
[0141] Combined with normalization techniques, this approach can accurately control the mean and variance of the routing gate scores. The complete training loss target is:
[0142]
[0143] In this embodiment, a smaller normalized weight w is set. n =0.01.
[0144] In speech-to-speech translation (S2ST) systems, it's common to integrate a speech-to-text translation (S2TT) module with a text-to-speech (TTS) component. However, in simultaneous translation scenarios, a streaming TTS system is crucial to maintaining low latency across the entire pipeline. While specialized streaming TTS systems exist, this example demonstrates that the expert routing threshold strategy provides a general and straightforward approach to converting a standard Transformer-based autoregressive TTS system into a streaming variant.
[0145] In this embodiment, the open source CozyVoice 2 is used as the base model, which consists of a decoder-only language model and a flow-based acoustic model. During the adaptation process, the flow model is retained and the language model is fine-tuned using the above-mentioned techniques and loss functions. During the training process, a routing threshold module and a hybrid expert post-processing module with the same architecture are added. Since the backbone network does not have an encoder, this embodiment uses the last hidden state output of the text markup as h prefix and use the hidden state of the end-of-text (EOS) marker as H global .
[0146] In speech-to-text translation (S2TT), when the input reaches the end of the stream, the read-write strategy is usually discontinued, and generation continues until the EOS marker appears. However, in TTS tasks, the judgment of the end of speech is relatively vague. Therefore, if the same strategy as S2TT is adopted, the problem of hallucination may occur at the end of generation. Therefore, this embodiment adopts a hybrid strategy to determine the end of generation:
[0147]
[0148] Among them, E t / s Indicates the end of text / speech (EOS) mark, T text / speech Indicates the text or speech mark in the current sequence, p ·,i is the routing output of the current position after reading all input streams, λ end =0.9 indicates the threshold control condition.
[0149] Through this invention, the present inventors not only developed a more complete S2ST system that can achieve real-time speech-to-speech translation, but also verified the compatibility of the expert routing gated flow strategy in different tasks and model architectures.
[0150] Compared with the prior art, the efficient simultaneous interpretation method based on expert routing threshold provided by the embodiment of the present invention has the following beneficial effects:
[0151] 1. Model performance: In experiments with multiple test sets, the performance index of the present invention decreased by less than 5% compared with the offline model at an average delay of 2 seconds, reflecting its robustness.
[0152] 2. Multi-language support: This invention uses a single model to train and test 30 language pairs in 6 languages, achieving good results and stable performance.
[0153] 3. Computational overhead: The computational overhead of the proposed method is the same as that of the classic Transformer during inference. According to experimental tests, when the model is deployed on an Nvidia-H100 graphics card using native PyTorch implementation, the additional latency due to the model computational overhead is only 50ms.
[0154] 4. Speech-to-speech translation: While maintaining industry-leading translation and sound quality, it provides a low-latency streaming speech-to-speech translation link, which is more practical in life.
[0155] 5. Potential for More Applications: This invention has been shown to achieve good results in both streaming translation and streaming TTS tasks, demonstrating the versatility of the method and its potential for application to more streaming sequence generation tasks.
[0156] The above describes in detail the preferred embodiments of the present invention. It should be understood that those skilled in the art can make numerous modifications and variations based on the concepts of the present invention without inventive effort. Therefore, any technical solutions that can be derived by those skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.
Claims
1. An efficient simultaneous interpretation method based on expert routing threshold, characterized in that: The method is based on the classic Transformer architecture model and constructs an expert routing strategy model based on expert routing thresholds. By training the expert routing strategy model, multi-language streaming translation is achieved. The expert routing strategy model includes: A streaming speech encoder, using a Transformer architecture, is compatible with traditional speech-to-text translation systems. The streaming speech encoder adopts a hybrid design consisting of block-by-block autoregressive blocks and non-autoregressive blocks. The text decoder follows the standard autoregressive Transformer architecture and processes both the complete offline speech and the randomly truncated speech prefix to generate the offline hidden state and the prefix hidden state. a routing threshold module implemented by a two-layer feed-forward network with a sigmoid head, wherein the routing threshold module projects the final hidden state of the text decoder into a scalar value and uses the scalar value to determine the expert weight; A hybrid expert post-processing module adopts a Transformer-like architecture and shares a language model head with the text decoder. The hybrid expert post-processing module combines prefix information and global information to predict the target translation sequence and utilizes the output result of the routing threshold module.
2. The method according to claim 1, wherein The streaming speech encoder uses 20 block-by-block autoregressive blocks Chunk-AR and 4 non-autoregressive blocks NAR, wherein, The Chunk-AR only calculates new chunks and optimizes inference efficiency through a cache key-value mechanism; The NAR recomputes the entire sequence, maintains translation quality by capturing global context, and adds a learnable end-of-stream flag in front of the NAR input, which is configured to indicate whether the current block terminates the audio stream.
3. The method according to claim 2, wherein The hybrid expert post-processing module includes multiple blocks, each of which includes a dual expert module, a pre-order output attention module and a feedforward network module, wherein: The dual-expert module includes a prefix expert and a global expert. The global expert is configured as a two-layer MLP module. The prefix expert adopts a standard cross-attention mechanism. The combined weight of the prefix expert and the global expert is used to determine whether the current input prefix contains sufficient information required to generate the target tag. The pre-order output attention module is configured to strictly isolate global information, no longer pay attention to the hidden state in the same layer, and only pay attention to the pre-order output of the decoder to prevent global information leakage.
4. The method according to claim 3, wherein The final output of the hybrid expert post-processing module is combined with the contribution of the dual expert module through a gated residual connection. The specific results are: in, is the output of the i-th position, is the input at position i, is the global expert input for position i, is the prefix expert input for position i, is the global expert weight of the ith position, is the prefix expert weight of the i-th position, h offline For offline hidden state, H global is a global embedding, is the projection weight, [,] is the vector concatenation, and MHA is the multi-head attention.
5. The method according to claim 4, wherein The expert routing strategy model is based on a pre-trained offline speech-to-text translation S2TT model, and the training of the expert routing strategy model is achieved through an offline pre-training stage and a synchronous training stage.
6. The method according to claim 5, wherein In the offline pre-training stage, the standard offline S2TT target offline loss is used. Train the model until convergence.
7. The method according to claim 6, wherein In the synchronous training phase, combined with the offline loss Prefix loss and post-mixing losses The total loss of the model during the synchronous training phase is: in, is the total loss, Offline loss, is the prefix loss, is the post-mixing processing loss, ω r ,ω p is the weight value, p ref is the output distribution of the hybrid expert refiner, p dec is the output distribution of the text decoder, λ is the hyperparameter, h prefix Hide the prefix status.
8. The method according to claim 7, wherein During the streaming inference process, the hybrid expert post-processing module uses an expert routing strategy to decide whether to continue outputting or wait and read subsequent audio segments. The expert routing strategy is configured as follows: Determine whether the routing threshold score output by the autoregressive decoding is lower than the preset threshold. If it is lower than the threshold, continue to output; otherwise, wait and read the subsequent audio clip: Among them, p t,i is the routing threshold score at input time t and target location i, and ε is the threshold.
9. The method according to claim 8, wherein The hybrid expert post-processing module uses a normalization method to align the average score of the routing gate with the relative information content between the prefix and the global context. The normalization method calculates the normalization loss based on the prefix length and the full sequence length. The specific calculation method of the normalization loss is: in, is the normalized loss, l p is the prefix length, l g is the complete sequence length, l b is the buffer hyperparameter, is the average routed gate output of the sequence.
10. The method according to claim 9, wherein During the training process of the expert routing strategy model, the mean and variance of the routing gate scores are controlled in combination with the normalization method. The complete training loss value target is: in, is the training loss value, is the total loss during the synchronous training phase, is the normalized loss, ω n is the normalized weight.
Citation Information
Cited By
Multilingual speech synthesis method and related device
CN121708899A