Chinese-English Speech Recognition Method and System Based on Zip-MoE Model Grouped Mixture-of-Experts Layer
Through the grouping hybrid expert layer structure of the Zip-MoE model, the language confusion problem of Chinese and English speech recognition system in the Chinese and English speech mix scenario is solved, and efficient and flexible speech recognition and recognition accuracy is achieved, which is suitable for speech recognition tasks in multiple scenarios.
Patent Information
- Application Number
- CN202510607710.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-13
AI Technical Summary
When facing the Chinese-English speech recognition system, it is difficult to effectively solve the problem of language confusion when facing the Chinese-English mixed speech scenario, resulting in low recognition accuracy and complex training and reasoning processes.
A grouping hybrid expert layer structure based on Zip-MoE model is adopted, including a Chinese expert group, an English expert group and a language router. The language router is trained through the CTC loss function to perform language classification, and the unsupervised router is used for characterization and decoupling, combining the Bypass module and the Softmax function to optimize the output of the expert group.
It realizes efficient and flexible adaptation of Chinese and English speech recognition, supports frame-level language recognition in streaming scenarios, improves recognition accuracy and inference efficiency, and supports the deployment of the model when resource constraints are encountered.
Smart Images

Figure CN120126451B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and particularly to a Chinese-English speech recognition method and system based on a grouped mixture-of-experts layer of a Zip-MoE model. Background Art
[0002] Speech interaction is the most natural and widespread interaction method in human-computer interaction. Among them, ASR (Automatic Speech Recognition) technology is used to convert human speech into corresponding text, and the experience of speech interaction is closely related to the accuracy of automatic speech recognition. In many scenarios of daily life, such as in a meeting scenario, some English professional terms will be used, or there are often some English words in the process of dialogue interaction with a voice assistant. The recognition effect of the English part affects the experience of speech interaction. Therefore, a speech recognition system should not only support Chinese, but also have excellent recognition effects on English and Chinese-English mixed speech.
[0003] At present, most Chinese-English speech recognition systems do not optimize the model for the Chinese-English mixed speech scenario, but directly adopt the standard model framework for end-to-end training. This solution is difficult to handle the problem of language confusion, which in turn affects the speech recognition accuracy in Chinese-English monolingual scenarios and Chinese-English mixed speech scenarios.
[0004] In the prior art, patent CN 111816169 A is equipped with an encoder for Chinese and an encoder for English respectively. Each encoder is pre-trained with the corresponding monolingual data, and then fine-tuned with Chinese-English mixed data. Finally, the representations of the two encoders are fused for decoding. However, it is necessary to pre-train the Chinese and English encoders respectively, and the overall process is cumbersome, the model calculation amount is large, and it is difficult to achieve efficient training and inference.
[0005] The dual-encoder architecture is a typical architecture for dealing with the problem of language confusion. Such as the LSCA architecture, by designing a Chinese encoder and an English encoder to decouple the Chinese and English representations. Although it alleviates the problem of language confusion to a certain extent, this method has a large model calculation overhead and a complex training process;
[0006] On the other hand, the LAE architecture proposes to partially share the dual-encoder to reduce the calculation overhead. This architecture still has a complex training process and it is difficult to achieve efficient model inference;
[0007] In order to enable the speech recognition system to have the advantages of both language decoupling and efficient inference at the same time, LSR-MoE proposes to use a mixture-of-experts model to model the Chinese-English speech recognition task, and only one expert is activated during the inference process to achieve efficient inference. However, LSR-MoE does not introduce language information, and the model performance is poor.
[0008] To solve the above problems, this solution proposes a Chinese-English speech recognition method and system based on a grouped mixture-of-experts layer of a Zip-MoE model. Summary of the Invention
[0009] In view of one or more technical deficiencies in the above-mentioned prior art, the present application proposes the following technical solutions.
[0010] Based on the first aspect of the present application, a Chinese-English speech recognition system based on a grouped mixture-of-experts layer of a Zip-MoE model is proposed, including:
[0011] The Zip-MoE model includes 6 encoder blocks, and a Bypass module is included between every two encoder blocks. The Bypass module is used to learn the weights of the weighted output of the previous encoder block and the current encoder block;
[0012] Among them, the first 3 encoder blocks are of the standard Zipformer structure;
[0013] The last 3 encoder blocks adopt the Zipformer-MoE structure, and the Zipformer-MoE structure includes a grouped mixture-of-experts layer, and the grouped mixture-of-experts layer is used to replace the last feed-forward neural network (FNN) of the standard Zipformer structure;
[0014] The grouped mixture-of-experts layer includes a Chinese expert group, an English expert group, and a language router. Both the Chinese expert group and the English expert group are composed of a number of expert networks, and each is configured with an independent unsupervised router.
[0015] Furthermore, the language router is arranged after the downsampling operation of the Zipformer-MoE structure encoder block, and the language of each frame of speech is obtained by training the language router through the CTC loss function. The training formula of the language router is expressed as:
[0016] ;
[0017] Among them, represents the loss function of the language router, represents the word-level label of language recognition, represents the speech feature after downsampling, represents the weight of the linear classification layer of the language recognition task, D represents the embedding dimension, 3 represents the Chinese language, the English language, and a CTC blank label.
[0018] Furthermore, the language router outputs the language classification result of each frame of speech based on the speech feature of the current Zipformer-MoE structure encoding block , and based on the language classification result Decouple the mixed-language representation to obtain a Chinese representation and an English representation. The decoupling formula is:
[0019] ;
[0020] Among them, represents the Chinese representation, , represents the English representation, , represents the mixed-language representation, represents the language classification result, , represents the number of voice frames routed to the Chinese expert group, represents the number of voice frames routed to the English expert group, T represents the total number of voice frames, represents the Chinese mask matrix, represents the English mask matrix, B represents the batch size, D represents the embedding dimension, represents element-wise multiplication, is an indicator function used to obtain the frame-level language mask matrix.
[0021] During the process of performing language classification, the language router does not depend on subsequent frames and can support streaming frame-synchronous language recognition.
[0022] Furthermore, the unsupervised router has a fully connected structure. The unsupervised router uses the Softmax function to process the Chinese expert group and the English expert group to obtain the unsupervised routing weights, which are expressed as:
[0023] ;
[0024] Among them, represents the weight of the unsupervised router, s represents the attribute of the language group, , zh represents the Chinese expert group, en represents the English expert group, represents the mixed-language representation, including the Chinese representation and the English representation, represents the unsupervised router.
[0025] Based on the second aspect of this application, a method for Chinese-English speech recognition according to the system described in any one of the above is also proposed, including:
[0026] S1: Input the speech into the Zip-MoE model, perform downsampling on adjacent frames of the speech, learn the weighted weights of the output of the previous encoder block and the output of the current encoder block through the Bypass module, and perform weighted fusion on the speech according to the weights learned by the Bypass module to reduce the frame rate of the speech;
[0027] S2: Train the language router using the CTC loss function. The encoder block learns the mixed-language representation of the speech and restores the frame rate of the speech through an upsampling operation.
[0028] S3: Input the mixed-language representation into the grouped mixture-of-experts layer. The language router outputs the language classification result for each frame of speech and decouples the mixed-language representation into a Chinese representation and an English representation based on the language classification result.
[0029] S4: The unsupervised router routes the Chinese representation and the English representation to the Chinese expert group and the English expert group respectively. After processing the Chinese representation and the English representation through the Softmax function, the unsupervised routing weights of the language groups are obtained, and the k expert indices of the language groups are fused according to the unsupervised routing weights to obtain the output of the language groups.
[0030] S5: Sum up the outputs of the language groups to obtain the final output of the grouped mixture-of-experts layer, and perform Chinese-English speech recognition based on the final output.
[0031] The grouped mixture-of-experts layer inputs Chinese representations and English representations with language consistency, enabling the unsupervised router to focus on the fine-grained differences within the language and achieve more refined expert modeling.
[0032] Furthermore, the k expert indices are selected through the Top-k strategy, and the specific formula is as follows:
[0033] ;
[0034] where represents the mixed-language representation the k expert indices i corresponding to the speech frame t in , , is the Top-K expert set corresponding to the language group s at the speech frame t.
[0035] Furthermore, the language groups include a Chinese expert group and an English expert group, and the output of the language groups is expressed as:
[0036] ;
[0037] where represents the output of the language group, represents the i-th expert network within the language group s, represents the unsupervised routing weight of the i-th expert within the language group s, , Representing a mixed language representation The k expert indices i corresponding to the speech frame t in
[0038] The final output of the grouped mixture-of-experts layer is expressed as:
[0039] ;
[0040] Wherein, represents the final output of the grouped mixture-of-experts layer, , represents the output of the language group.
[0041] Furthermore, the unsupervised router adopts an implicit learning mechanism, autonomously learns the feature distribution of the input data, and dynamically allocates speech frames to the most matching experts.
[0042] Based on the third aspect of the present application, a computer program product is also proposed, which has one or more computer programs, and when the computer programs are executed by a computer processor, the method described in any one of the above is implemented.
[0043] The technical effect of the present invention is as follows: Based on the Zipformer structure encoder block, the Zip-MoE model is constructed in the present application. The idea of expert grouping design is adopted, and the flexible expansion of the model is realized through expert groups. A novel hierarchical routing strategy is proposed to decouple Chinese and English features. The speech routing and the unsupervised routing work together. The speech routing realizes the recognition of frame-level languages, and the unsupervised routing further finely models other attributes of the speech. The grouped mixture-of-experts layer selects the number of expert activations according to needs during training, can adapt to streaming scenarios with different latency requirements, supports a flexible Top-k inference mechanism, and can balance model performance and inference efficiency. Moreover, the present application also supports parameter pruning of the trained grouped mixture-of-experts layer, so that a smaller model can be deployed when resources are limited, greatly improving the accuracy and flexibility of speech recognition. Description of the Drawings
[0044] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present application will become more obvious.
[0045] Figure 1 is the overall framework diagram of the Zip-MoE model of a Chinese-English speech recognition system based on the grouped mixture-of-experts layer of the Zip-MoE model provided by the embodiment of the present application.
[0046] Figure 2 is the structural diagram of the grouped mixture-of-experts layer of a Chinese-English speech recognition system based on the grouped mixture-of-experts layer of the Zip-MoE model provided by the embodiment of the present application.
[0047] Figure 3 It is a flowchart of a Chinese-English speech recognition method based on the grouped mixture-of-experts layer of the Zip-MoE model according to an embodiment of the present application.
[0048] Figure 4 It is a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners
[0049] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings.
[0050] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.
[0051] Figure 1 It shows an overall framework diagram of the Zip-MoE model of a Chinese-English speech recognition system based on the grouped mixture-of-experts layer of the Zip-MoE model, including:
[0052] The Zip-MoE model includes 6 encoder blocks. Except that there is no Bypass module between the first two encoder blocks, there is a Bypass module between every two of the remaining encoder blocks. The Bypass module is used to learn the weights of the weighted outputs of the previous encoder block and the current encoder block.
[0053] Among them, the first 3 encoder blocks are of the standard Zipformer structure;
[0054] The last 3 encoder blocks adopt the Zipformer-MoE structure. The Zipformer-MoE structure includes a grouped mixture-of-experts layer, and the grouped mixture-of-experts layer is used to replace the last feed-forward neural network (FNN) of the standard Zipformer structure.
[0055] As Figure 2 shown, the grouped mixture-of-experts layer includes a Chinese expert group, an English expert group, and a language router. Both the Chinese expert group and the English expert group are composed of several expert networks, and each is configured with an independent unsupervised router.
[0056] It should be noted that the language router is set after the downsampling operation of the Zipformer-MoE structure encoder block, and the language of each frame of speech is obtained by training the language router with the CTC loss function. The training formula of the language router is expressed as:
[0057] ;
[0058] where, represents the loss function of the language router, represents the word-level label for language identification, represents the speech feature after downsampling, represents the weight of the linear classification layer for the language identification task, D represents the embedding dimension, 3 represents the Chinese language, the English language, and a CTC blank label.
[0059] It should be noted that the language router outputs the language classification result of each frame of speech based on the speech feature of the current Zipformer-MoE structure encoding block , and based on the language classification result decouples the mixed language representation to obtain the Chinese representation and the English representation. The decoupling formula is:
[0060] ;
[0061] where, represents the Chinese representation, , represents the English representation, , represents the mixed language representation, represents the language classification result, , represents the number of speech frames routed to the Chinese expert group, represents the number of speech frames routed to the English expert group, T represents the total number of speech frames, represents the Chinese mask matrix, represents the English mask matrix, B represents the batch size, D represents the embedding dimension, represents element-wise multiplication, is an indicator function used to obtain the frame-level language mask matrix.
[0062] It should be noted that the unsupervised router has a fully connected structure. The unsupervised router uses the Softmax function to process the Chinese expert group and the English expert group to obtain the unsupervised routing weight. The unsupervised routing weight is expressed as:
[0063] ;
[0064] Among them, represents the weight of the unsupervised router, s represents the attributes of the language group, , zh represents the Chinese expert group, en represents the English expert group, represents the mixed language representation, including Chinese representation and English representation, represents the unsupervised router.
[0065] It should be noted that the Zip-MoE model has different speech frame rates between different encoder blocks. The Zip-MoE model learns time-domain information with different granularities according to different speech frame rates. The speech frame rate is reduced by downsampling operations between encoder blocks. The encoder blocks learn the features of speech under the condition of low frame rate, and then the speech frame rate is restored by upsampling operations.
[0066] In a specific embodiment, all the encoder blocks are 16 layers in total. Among them, the numbers of layers of the first three Zipformer-structured encoder blocks are 2 layers, 2 layers, and 3 layers respectively, and the numbers of layers of the last three Zipformer-MoE-structured encoder blocks are 4 layers, 3 layers, and 2 layers.
[0067] It should be noted that each Zipformer-MoE-structured encoder block shares the weight of the same language router. There are N experts and one unsupervised router in each of the Chinese expert group and the English expert group, and the number of experts in the language group is flexibly variable.
[0068] It should be noted that the language router solves the problem of language confusion, enabling the speech representations of different languages to be respectively routed to different language expert groups to extract specific language representations, selecting the corresponding language expert for each frame of the input speech, and making the calculation of the grouped mixed expert layer efficient. The language router adopts frame-synchronous language recognition that supports streaming and does not depend on subsequent frames during the process of performing language classification. This is also one of the important invention points of this application.
[0069] It should be noted that the grouped mixed expert layer inputs Chinese representations and English representations with language consistency, enabling the unsupervised router to focus on the fine-grained differences within the language, such as different dialects and accents and other attributes in the speech of the same language, and realizing more refined expert modeling.
[0070] It should be noted that this application constructs a Zip-MoE model based on the Zipformer structure encoder. Adopting the idea of expert grouping design, it realizes the flexible expansion of the model through expert groups, and proposes a novel hierarchical routing strategy to decouple Chinese and English features. The voice routing and unsupervised routing work together. The voice routing realizes the frame-level language identification, and the unsupervised routing further finely models other attributes of the voice. The grouped mixture-of-experts layer selects the number of activated experts according to needs during training, can adapt to streaming scenarios with different latency requirements, supports a flexible Top-k inference mechanism, and can balance model performance and inference efficiency. Moreover, this application also supports parameter pruning of the trained grouped mixture-of-experts layer, so that a smaller model can be deployed when resources are limited, greatly improving the accuracy and flexibility of speech recognition.
[0071] Figure 3 The flowchart shows a Chinese-English speech recognition method based on the grouped mixture-of-experts layer of the Zip-MoE model, including:
[0072] S1: Input the speech into the Zip-MoE model, perform downsampling operations on adjacent frames of the speech, learn the weighted weights of the output of the previous encoder block and the current encoder block through the Bypass module, and perform weighted fusion on the speech according to the weights learned by the Bypass module to reduce the frame rate of the speech;
[0073] S2: Train the language router through the CTC loss function. The encoder block learns the mixed-language representation of the speech and restores the frame rate of the speech through upsampling operations;
[0074] S3: Input the mixed-language representation into the grouped mixture-of-experts layer. The language router outputs the language classification result of each frame of speech, and decouples the mixed-language representation into Chinese representation and English representation based on the language classification result;
[0075] S4: The unsupervised router routes the Chinese representation and the English representation to the Chinese expert group and the English expert group respectively, processes the Chinese representation and the English representation through the Softmax function to obtain the unsupervised routing weights of the language groups, and fuses the k expert indices of the language groups according to the unsupervised routing weights to obtain the output of the language groups;
[0076] S5: Sum up the outputs of the language groups to obtain the final output of the grouped mixture-of-experts layer, and perform Chinese-English recognition of the speech according to the final output.
[0077] It should be noted that the k expert indices are selected through the Top-k strategy, and the specific formula is as follows:
[0078] ;
[0079] Among them, represents the k expert indices i corresponding to the speech frame t in the mixed language representation , and , , is the Top-K expert set corresponding to the language group s at the speech frame t.
[0080] It should be noted that the language group includes a Chinese expert group and an English expert group, and the output of the language group is expressed as:
[0081] ;
[0082] Among them, represents the output of the language group, represents the i-th expert network within the language group s, represents the unsupervised routing weight of the i-th expert within the language group s, , represents the k expert indices i corresponding to the speech frame t in the mixed language representation ;
[0083] The final output of the grouped mixed expert layer is expressed as:
[0084] ;
[0085] Among them, represents the final output of the grouped mixed expert layer, , represents the output of the language group.
[0086] It should be noted that the unsupervised router adopts an implicit learning mechanism, autonomously learns the feature distribution of the input data, and dynamically assigns speech frames to the most matching experts.
[0087] It should be noted that the dynamic Top-k strategy of this application makes the combination of experts within the language expert group more diverse and possible, enhancing the expressive power of the Zip-MoE model.
[0088] In a specific embodiment, the number of experts k follows a discrete uniform distribution, , and at least one expert within the expert group is activated each time, and at most half of the experts within the expert group are activated. During the model inference process, the user can specify different k values to balance the model performance and inference efficiency. For example, specifying a larger k value can obtain better model performance, and specifying a smaller k value can obtain a faster model inference speed.
[0089] It should be noted that the grouped mixture-of-experts layer proposed in this application has high computational efficiency. Through the design of the Chinese expert group and the English expert group, the flexible expansion of the Zip-MoE model can be realized, the number of experts can be flexibly expanded to increase the model capacity, and a novel hierarchical routing strategy is proposed to achieve efficient decoupling of Chinese and English features.
[0090] It should be noted that this application constructs a Zip-MoE model based on the Zipformer encoder block, introduces a grouped mixture-of-experts layer, and proposes a novel hierarchical routing strategy to decouple Chinese and English features. The voice routing and unsupervised routing work together. The voice routing realizes real-time frame-level language identification and routes the speech features to the matching language expert group. The unsupervised routing further finely models attributes such as accent and dialect. Only a small number of matching language group experts are activated during the inference process, enabling efficient computing.
[0091] It should be noted that the entire training process of the Zip-MoE model is simple and efficient. Without pre-training, the training data in Chinese and English can be directly mixed for training, and multi-scenario adaptation can be achieved in one training, significantly improving the recognition accuracy and scalability of Chinese and English speech recognition tasks.
[0092] The following refers to Figure 4 , which shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of this application. Figure 4 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.
[0093] As Figure 4 shown, the computer system includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage section 408 into the random access memory (RAM) 403. In the RAM 403, various programs and data required for system operation are also stored. The CPU 401, ROM 402, and RAM 403 are connected to each other through a bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.
[0094] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 410 as needed so that a computer program read therefrom is installed into the storage section 408 as needed.
[0095] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable storage medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 409 and / or installed from the removable medium 411. When the computer program is executed by the central processing unit (CPU) 401, the above functions defined in the method of the present application are performed. It should be noted that the computer-readable storage medium of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable storage medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0096] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).
[0097] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0098] The modules described in the embodiments of this application can be implemented in software or in hardware.
[0099] As another aspect, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist separately without being assembled into the electronic device. The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to: input voice into the Zip-MoE model, perform downsampling operations on adjacent frames of the voice, learn the weighted weights of the output of the previous encoder block and the output of the current encoder block through the Bypass module, and perform weighted fusion on the voice according to the weights learned by the Bypass module to reduce the frame rate of the voice; train the language router through the CTC loss function, the encoder block learns the mixed-language representation of the voice, and restores the frame rate of the voice through upsampling operations; input the mixed-language representation into the grouped mixture-of-experts layer, the language router outputs the language classification result of each frame of voice, and decouples the mixed-language representation into a Chinese representation and an English representation based on the language classification result; the unsupervised router routes the Chinese representation and the English representation to the Chinese expert group and the English expert group respectively, processes the Chinese representation and the English representation through the Softmax function to obtain the unsupervised routing weights of the language groups, and fuses the k expert indexes of the language groups according to the unsupervised routing weights to obtain the output of the language groups; sum up the outputs of the language groups to obtain the final output of the grouped mixture-of-experts layer, and perform Chinese-English recognition of the voice according to the final output.
[0100] Finally, it should be noted that the above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present application.
Claims
1. A Chinese-English speech recognition system based on the grouped mixture-of-experts layer of the Zip-MoE model, characterized in that, It includes: The Zip-MoE model includes 6 encoder blocks, and there is a Bypass module between every two encoder blocks. The Bypass module is used to learn the weights for weighting the output of the previous encoder block and the output of the current encoder block. Among them, the first 3 encoder blocks are of the standard Zipformer structure. The last 3 encoder blocks adopt the Zipformer-MoE structure, and the Zipformer-MoE structure includes a grouped mixture-of-experts layer, which is used to replace the last feed-forward neural network (FNN) of the standard Zipformer structure. The grouped mixture-of-experts layer includes a Chinese expert group, an English expert group, and a language router. Both the Chinese expert group and the English expert group are composed of several expert networks, and each is configured with an independent unsupervised router. The language router is set after the downsampling operation of the encoder block in the Zipformer-MoE structure, and the language router is trained through the CTC loss function to obtain the language of each frame of speech. The training formula of the language router is expressed as: ; Among them, represents the loss function of the language router, represents the word-level label of language recognition, represents the speech feature after downsampling, represents the weight of the linear classification layer for the language recognition task, D represents the embedding dimension, 3 represents the Chinese language, the English language, and a blank label for CTC; The language router outputs the language classification result of each frame of speech based on the speech features of the current Zipformer-MoE structure encoding block , and based on the language classification result decouples the mixed language representation to obtain a Chinese representation and an English representation. The decoupling formula is: ; Among them, represents Chinese representation, , represents English representation, represents mixed language representation, represents the language classification result, , represents the number of voice frames routed to the Chinese expert group, represents the number of voice frames routed to the English expert group, and T represents the total number of voice frames, represents the Chinese mask matrix, represents the English mask matrix, B represents the batch size, and D represents the embedding dimension, represents element-wise multiplication, is an indicator function used to obtain the frame-level language mask matrix.
2. The system according to claim 1, wherein, The unsupervised router is of a fully connected structure. The unsupervised router uses the Softmax function to process the Chinese expert group and the English expert group to obtain the unsupervised routing weights, and the unsupervised routing weights are expressed as: ; Among them, represents the weight of the unsupervised router, s represents the attribute of the language group, , zh represents the Chinese expert group, and en represents the English expert group, represents the mixed language representation, including Chinese representation and English representation, represents the unsupervised router.
3. A method for Chinese-English speech recognition according to the system described in any one of claims 1-2, characterized in that, It includes: S1: Input speech into the Zip-MoE model, perform downsampling operations on adjacent frames of the speech, learn the weights for weighting the output of the previous encoder block and the output of the current encoder block through the Bypass module, and perform weighted fusion on the speech according to the weights learned by the Bypass module to reduce the frame rate of the speech. S2: Train the language router through the CTC loss function. The encoder block learns the mixed-language representation of the speech and restores the frame rate of the speech through an upsampling operation. S3: Input the mixed-language representation into the grouped mixture-of-experts layer. The language router outputs the language classification result of each frame of speech, and decouples the mixed-language representation into a Chinese representation and an English representation based on the language classification result. S4: The unsupervised router routes the Chinese representation and the English representation to the Chinese expert group and the English expert group respectively, processes the Chinese representation and the English representation through the Softmax function to obtain the unsupervised routing weights of the language groups, and fuses the k expert indices of the language groups according to the unsupervised routing weights to obtain the output of the language groups. S5: Sum up the outputs of the language groups to obtain the final output of the grouped mixture-of-experts layer, and perform Chinese-English speech recognition based on the final output.
4. The method according to claim 3, characterized in that, The k expert indices are selected through the Top-k strategy, and the specific formula is as follows: ; Among them, represents the k expert indices i corresponding to the speech frame t in the mixed language representation, , , and is the Top-K expert set corresponding to the language group s at the speech frame t.
5. The method according to claim 4, characterized in that The language groups include a Chinese expert group and an English expert group, and the output of the language groups is expressed as: ; Among them, represents the output of the language group, represents the i-th expert network within the language group s, represents the unsupervised routing weight of the i-th expert within the language group s, , represents the mixed language representation for the k expert indices i corresponding to the speech frame t; The final output of the grouped mixture-of-experts layer is expressed as: ; Among them, represents the final output of the grouped mixture-of-experts layer, , represents the output of the language group.
6. The method according to claim 3, wherein The unsupervised router adopts an implicit learning mechanism, autonomously learns the feature distribution of the input data, and dynamically assigns speech frames to the most matching experts.
7. A computer program product having one or more computer programs thereon, characterized in that, When the computer program is executed by a computer processor, it implements the method according to any one of claims 3 to 6.
Citation Information
Patent Citations
Chinese and English hybrid speech recognition model training method and device
CN111816169A
Speech recognition method, speech recognition model self-supervision pre-training method and electronic equipment
CN118248130A