Zip-MoE model grouping mixed expert layer-based Chinese and English speech recognition method and system
By adopting a grouping hybrid expert layer based on the Zip-MoE model in the Chinese-English voice recognition system, using language routers and unsupervised routers to work together, the language confusion problem in the Chinese-English mixed speech scenario is solved, efficient Chinese-English voice recognition is achieved, and recognition accuracy and flexibility are improved.
Patent Information
- Application Number
- CN202510607710.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-13
AI Technical Summary
The existing Chinese-English voice recognition system is difficult to deal with language confusion in the Chinese-English mixed speech scenario, resulting in a decrease in the accuracy of speech recognition. The existing technology models are large in computing and complex in training processes, making it difficult to achieve efficient training and reasoning.
The packet hybrid expert layer based on the Zip-MoE model is adopted, and the language router and unsupervised router work together to achieve decoupling and efficient inference of Chinese and English voice features. Specific measures include: using a Zip-MoE model with 6 encoder blocks, adding a Bypass module between each two encoder blocks, using a grouped hybrid expert layer to replace the last feedforward network of the standard Zipformer structure, including a Chinese expert group, an English expert group and a language router, and training the language router through the CTC loss function.
It realizes efficient recognition of the Chinese and English voice recognition system in the mixed speech scenarios of Chinese and English, reduces the overhead of model calculation, simplifies the training process, improves the accuracy and flexibility of speech recognition, and supports streaming frame synchronization language recognition.
Smart Images

Figure CN120126451A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and particularly to a Chinese-English speech recognition method and system based on a grouped mixture-of-experts layer of a Zip-MoE model. Background Art
[0002] Speech interaction is the most natural and widespread interaction method in human-computer interaction. Among them, the ASR automatic speech recognition technology is used to convert human speech into corresponding text, and the experience of speech interaction is closely related to the accuracy of automatic speech recognition. In many scenarios of daily life, such as in a meeting scenario, some English professional terms will be used, or there are often some English words in the process of dialogue interaction with a voice assistant. The recognition effect of the English part affects the experience of speech interaction. Therefore, the speech recognition system should not only support Chinese, but also have excellent recognition effects for English and Chinese-English mixed speech.
[0003] Currently, most Chinese-English speech recognition systems do not optimize the model for the Chinese-English mixed speech scenario, but directly adopt the standard model framework for end-to-end training. This solution is difficult to handle the problem of language confusion, which in turn affects the speech recognition accuracy in Chinese-English monolingual scenarios and Chinese-English mixed speech scenarios.
[0004] In the prior art, patent CN 111816169 A is equipped with an encoder for Chinese and an encoder for English respectively. Each encoder adopts corresponding monolingual pre-training, and then fine-tunes with Chinese-English mixed data. Finally, the representations of the two encoders are fused for decoding. However, it is necessary to pre-train the Chinese and English encoders respectively, and the overall process is cumbersome, and the model calculation amount is large, making it difficult to achieve efficient training and inference.
[0005] The dual-encoder architecture is a typical architecture for dealing with the problem of language confusion. For example, the LSCA architecture decouples Chinese and English representations by designing a Chinese encoder and an English encoder. Although it alleviates the problem of language confusion to a certain extent, this method has a large model calculation overhead and a complex training process; On the other hand, the LAE architecture proposes to partially share the dual-encoder to reduce the calculation overhead. This architecture still has a complex training process and is difficult to achieve efficient model inference; In order to enable the speech recognition system to have the advantages of both language decoupling and efficient inference at the same time, LSR-MoE proposes to use a mixture-of-experts model to model the Chinese-English speech recognition task, and only activate one expert during the inference process to achieve efficient inference. However, LSR-MoE does not introduce language information, and the model performance is poor.
[0006] To solve the above problems, this solution proposes a Chinese-English speech recognition method and system based on a grouped mixture-of-experts layer of a Zip-MoE model. Summary of the Invention
[0007] In view of one or more technical deficiencies in the above-mentioned prior art, the present application proposes the following technical solutions.
[0008] Based on the first aspect of the present application, a Chinese-English speech recognition system based on a grouped mixture-of-experts layer of a Zip-MoE model is proposed, including: The Zip-MoE model includes 6 encoder blocks, and a Bypass module is included between every two encoder blocks. The Bypass module is used to learn the weights of the weighted output of the previous encoder block and the current encoder block. Among them, the first 3 encoder blocks are of the standard Zipformer structure; The last 3 encoder blocks adopt the Zipformer-MoE structure, and the Zipformer-MoE structure includes a grouped mixture-of-experts layer, and the grouped mixture-of-experts layer is used to replace the last feed-forward network FNN of the standard Zipformer structure. The grouped mixture-of-experts layer includes a Chinese expert group, an English expert group, and a language router. Both the Chinese expert group and the English expert group are composed of a number of expert networks, and an independent unsupervised router is configured respectively.
[0009] Furthermore, the language router is arranged after the downsampling operation of the Zipformer-MoE structure encoder block, and the language of each frame of speech is obtained by training the language router through the CTC loss function. The training formula of the language router is expressed as: ; Among them, represents the loss function of the language router, represents the word-level label of language recognition, represents the speech feature after downsampling, represents the weight of the linear classification layer of the language recognition task, D represents the embedding dimension, 3 represents the Chinese language, the English language, and a CTC blank label.
[0010] Furthermore, the language router outputs the language classification result of each frame of speech based on the speech feature of the current Zipformer-MoE structure encoding block , and based on the language classification result the mixed language representation is decoupled to obtain a Chinese representation and an English representation. The decoupling formula is: ; Among them, represents the Chinese representation, , Represents the English representation, , represents the mixed-language representation, represents the language classification result, , represents the number of speech frames routed to the Chinese expert group, represents the number of speech frames routed to the English expert group, and T represents the total number of speech frames, represents the Chinese mask matrix, represents the English mask matrix, B represents the batch size, and D represents the embedding dimension, represents the element-wise multiplication, is an indicator function used to obtain the frame-level language mask matrix.
[0011] The language router does not depend on subsequent frames during the process of language classification and can support streaming frame-synchronous language recognition.
[0012] Furthermore, the unsupervised router has a fully connected structure. The unsupervised router uses the Softmax function to process the Chinese expert group and the English expert group to obtain the unsupervised routing weights, and the unsupervised routing weights are expressed as: ; where, represents the weight of the unsupervised router, s represents the attribute of the language group, , zh represents the Chinese expert group, and en represents the English expert group, represents the mixed-language representation, including the Chinese representation and the English representation, represents the unsupervised router.
[0013] Based on the second aspect of the present application, a method for Chinese-English speech recognition according to the system described in any one of the above is also proposed, including: S1: Input the speech into the Zip-MoE model, perform downsampling on adjacent frames of the speech, learn the weighted weights of the output of the previous encoder block and the output of the current encoder block through the Bypass module, and perform weighted fusion on the speech according to the weights learned by the Bypass module to reduce the frame rate of the speech; S2: Train the language router through the CTC loss function. The encoder block learns the mixed-language representation of the speech and restores the frame rate of the speech through upsampling operation; S3: Input the mixed-language representation into the grouped mixture-of-experts layer. The language router outputs the language classification result of each frame of speech, and decouples the mixed-language representation into the Chinese representation and the English representation based on the language classification result; S4: The unsupervised router routes the Chinese representation and the English representation to the Chinese expert group and the English expert group respectively, processes the Chinese representation and the English representation through the Softmax function to obtain the unsupervised routing weights of the language group, and fuses the k expert indexes of the language group according to the unsupervised routing weights to obtain the output of the language group; S5: Sum up the outputs of the language groups to obtain the final output of the grouped mixture-of-experts layer, and perform Chinese-English speech recognition based on the final output.
[0014] The grouped mixture-of-experts layer inputs Chinese representations and English representations with language consistency, enabling the unsupervised router to focus on the fine-grained differences within the language, and achieving more refined expert modeling.
[0015] Further, the k expert indexes are selected through the Top-k strategy, and the specific formula is as follows: ; Where, represents the k expert indexes i corresponding to the speech frame t in the mixed-language representation , , , is the Top-K expert set corresponding to the language group s at the speech frame t.
[0016] Further, the language group includes a Chinese expert group and an English expert group, and the output of the language group is expressed as: ; Where, represents the output of the language group, represents the i-th expert network within the language group s, represents the unsupervised routing weight of the i-th expert within the language group s, , represents the k expert indexes i corresponding to the speech frame t in the mixed-language representation ; The final output of the grouped mixture-of-experts layer is expressed as: ; Where, represents the final output of the grouped mixture-of-experts layer, , represents the output of the language group.
[0017] Further, the unsupervised router adopts an implicit learning mechanism, autonomously learns the feature distribution of the input data, and dynamically assigns speech frames to the most matching experts.
[0018] Based on the third aspect of the present application, a computer program product is further proposed, which has one or more computer programs that, when executed by a computer processor, implement the method described in any one of the above.
[0019] The technical effects of the present invention are as follows: The present application constructs a Zip-MoE model based on the Zipformer structural encoder block, adopts the idea of expert grouping design, realizes the flexible expansion of the model through expert groups, and proposes a novel hierarchical routing strategy to decouple Chinese and English features. The voice routing and unsupervised routing work together. The voice routing realizes the frame-level language identification, and the unsupervised routing further finely models other attributes of the voice. The grouped mixture-of-experts layer selects the number of expert activations according to needs during training, can adapt to streaming scenarios with different latency requirements, supports a flexible Top-k inference mechanism, and can balance model performance and inference efficiency. Moreover, the present application also supports parameter pruning of the trained grouped mixture-of-experts layer, so that a smaller model can be deployed when resources are limited, greatly improving the accuracy and flexibility of speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Other features, objects, and advantages of the present application will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings.
[0021] Figure 1 FIG. is an overall framework diagram of the Zip-MoE model of a Chinese-English speech recognition system based on the grouped mixture-of-experts layer of the Zip-MoE model according to an embodiment of the present application.
[0022] Figure 2 FIG. is a structural diagram of the grouped mixture-of-experts layer of a Chinese-English speech recognition system based on the grouped mixture-of-experts layer of the Zip-MoE model according to an embodiment of the present application.
[0023] Figure 3 FIG. is a flowchart of a Chinese-English speech recognition method based on the grouped mixture-of-experts layer of the Zip-MoE model according to an embodiment of the present application.
[0024] Figure 4 FIG. is a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The present application will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the related invention and not for limiting the invention. Additionally, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings.
[0026] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will detail this application with reference to the accompanying drawings and in combination with the embodiments.
[0027] Figure 1 Fig. shows the overall framework diagram of the Zip-MoE model of a Chinese-English speech recognition system based on the grouped mixture-of-experts layer of the Zip-MoE model, including: The Zip-MoE model includes 6 encoder blocks. Except that there is no Bypass module between the first two encoder blocks, there is a Bypass module between every two of the remaining encoder blocks. The Bypass module is used to learn the weights of the weighted outputs of the previous encoder block and the current encoder block; Among them, the first 3 encoder blocks are of the standard Zipformer structure; The last 3 encoder blocks adopt the Zipformer-MoE structure, and the Zipformer-MoE structure includes a grouped mixture-of-experts layer, which is used to replace the last feed-forward network FNN of the standard Zipformer structure; As Figure 2 shown, the grouped mixture-of-experts layer includes a Chinese expert group, an English expert group, and a language router. Both the Chinese expert group and the English expert group are composed of a number of expert networks, and each is configured with an independent unsupervised router.
[0028] It should be noted that the language router is set after the downsampling operation of the Zipformer-MoE structure encoder block, and the language router is trained through the CTC loss function to obtain the language of each frame of speech. The training formula of the language router is expressed as: ; Among them, represents the loss function of the language router, represents the word-level label of language recognition, represents the speech features after downsampling, represents the weight of the linear classification layer of the language recognition task, D represents the embedding dimension, and 3 represents the Chinese language, the English language, and a CTC blank label.
[0029] It should be noted that the language router outputs the language classification result of each frame of speech based on the speech features of the current Zipformer-MoE structure encoding block and decouples the mixed language representation based on the language classification result to obtain the Chinese representation and the English representation. The decoupling formula is: ; Among them, represents the Chinese representation, , represents the English representation, , represents the mixed language representation, represents the language classification result, , represents the number of speech frames routed to the Chinese expert group, represents the number of speech frames routed to the English expert group, and T represents the total number of speech frames, represents the Chinese mask matrix, represents the English mask matrix, B represents the batch size, and D represents the embedding dimension, represents element-wise multiplication, is an indicator function used to obtain the frame-level language mask matrix.
[0030] It should be noted that the unsupervised router is a fully connected structure. The unsupervised router uses the Softmax function to process the Chinese expert group and the English expert group to obtain the unsupervised routing weights, and the unsupervised routing weights are expressed as: ; Among them, represents the weight of the unsupervised router, s represents the attribute of the language group, , zh represents the Chinese expert group, and en represents the English expert group, represents the mixed language representation, including the Chinese representation and the English representation, represents the unsupervised router.
[0031] It should be noted that the Zip-MoE model has different speech frame rates between different encoder blocks. The Zip-MoE model learns different granularities of time-domain information according to different speech frame rates, reduces the speech frame rate through downsampling operations between encoder blocks, and the encoder blocks learn the features of the speech under the condition of low frame rate, and then restore the speech frame rate through upsampling operations.
[0032] In a specific embodiment, there are a total of 16 layers in all encoder blocks. Among them, the numbers of layers of the first three Zipformer structure encoder blocks are 2 layers, 2 layers, and 3 layers respectively, and the numbers of layers of the last three Zipformer-MoE structure encoder blocks are 4 layers, 3 layers, and 2 layers.
[0033] It should be noted that each Zipformer-MoE structure encoder block shares the weight of the same language router. There are N experts and one unsupervised router in each of the Chinese expert group and the English expert group, and the number of experts in the language group is flexibly variable.
[0034] It should be noted that the language router solves the problem of language confusion, enabling the speech representations in different languages to be respectively routed to different language expert groups to extract specific language representations. For each frame of input speech, the corresponding language expert is selected, making the calculation of the grouped mixture of experts layer efficient. The language router adopts frame-synchronous language recognition that supports streaming and does not depend on subsequent frames during the process of language classification, which is also one of the important invention points of this application.
[0035] It should be noted that the grouped mixture of experts layer inputs Chinese representations and English representations with language consistency, enabling the unsupervised router to focus on the fine-grained differences within the language, such as different dialects and accents in the speech of the same language, to achieve more refined expert modeling.
[0036] It should be noted that this application constructs a Zip-MoE model based on the Zipformer structure encoder, adopts the idea of expert grouping design, flexibly expands the model through expert groups, and proposes a novel hierarchical routing strategy to decouple Chinese and English features. The speech routing and the unsupervised routing work together. The speech routing realizes frame-level language recognition, and the unsupervised routing further performs fine-grained modeling on other attributes of the speech. The grouped mixture of experts layer selects the number of activated experts according to needs during training, can adapt to streaming scenarios with different latency requirements, supports a flexible Top-k inference mechanism, and can balance model performance and inference efficiency. Moreover, this application also supports parameter pruning of the trained grouped mixture of experts layer to enable the deployment of a smaller model when resources are limited, greatly improving the accuracy and flexibility of speech recognition.
[0037] Figure 3 The flowchart of a Chinese-English speech recognition method based on the grouped mixture of experts layer of the Zip-MoE model is shown, including: S1: Input speech into the Zip-MoE model, perform downsampling operations on adjacent frames of the speech, learn the weighted weights of the output of the previous encoder block and the current encoder block through the Bypass module, and perform weighted fusion on the speech according to the weights learned by the Bypass module to reduce the frame rate of the speech; S2: Train the language router through the CTC loss function. The encoder block learns the mixed language representation of the speech and restores the frame rate of the speech through upsampling operations; S3: Input the mixed language representation into the grouped mixture of experts layer. The language router outputs the language classification result of each frame of speech, and decouples the mixed language representation into Chinese representation and English representation based on the language classification result; S4: The unsupervised router routes the Chinese representation and the English representation to the Chinese expert group and the English expert group respectively, processes the Chinese representation and the English representation through the Softmax function to obtain the unsupervised routing weights of the language group, and fuses the k expert indices of the language group according to the unsupervised routing weights to obtain the output of the language group; S5: Sum up the outputs of the language groups to obtain the final output of the grouped mixture-of-experts layer, and perform Chinese-English speech recognition based on the final output.
[0038] It should be noted that the k expert indices are selected through the Top-k strategy, and the specific formula is as follows: ; where represents the k expert indices i corresponding to the speech frame t in the mixed language representation , , , is the Top-K expert set corresponding to the language group s at the speech frame t.
[0039] It should be noted that the language group includes a Chinese expert group and an English expert group, and the output of the language group is expressed as: ; where represents the output of the language group, represents the i-th expert network within the language group s, represents the unsupervised routing weight of the i-th expert within the language group s, , represents the k expert indices i corresponding to the speech frame t in the mixed language representation ; The final output of the grouped mixture-of-experts layer is expressed as: ; where represents the final output of the grouped mixture-of-experts layer, , represents the output of the language group.
[0040] It should be noted that the unsupervised router adopts an implicit learning mechanism, autonomously learns the feature distribution of the input data, and dynamically assigns speech frames to the most matching experts.
[0041] It should be noted that the dynamic Top-k strategy of this application makes the expert combinations within the language expert group more diverse and possible, enhancing the expressive ability of the Zip-MoE model.
[0042] In a specific embodiment, the number of experts k follows a discrete uniform distribution. Each time, at least one expert in the expert group is activated, and at most half of the experts in the expert group are activated. During the model inference process, the user can specify different k values to balance the model performance and inference efficiency. For example, specifying a larger k value can obtain better model performance, and specifying a smaller k value can obtain a faster model inference speed.
[0043] It should be noted that the grouped mixture-of-experts layer proposed in this application has high computational efficiency. Through the design of the Chinese expert group and the English expert group, the flexible expansion of the Zip-MoE model can be realized. The number of experts can be flexibly expanded to increase the model capacity, and a novel hierarchical routing strategy is proposed to achieve efficient decoupling of Chinese and English features.
[0044] It should be noted that this application constructs a Zip-MoE model based on the Zipformer encoder block, introduces a grouped mixture-of-experts layer, and proposes a novel hierarchical routing strategy to decouple Chinese and English features. The voice routing and unsupervised routing work together. The voice routing realizes real-time frame-level language recognition and routes the voice features to the matching language expert group. The unsupervised routing further finely models attributes such as accents and dialects. During the inference process, only a small number of matching language group experts are activated, enabling efficient computation.
[0045] It should be noted that the entire training process of the Zip-MoE model is simple and efficient. Without pre-training, the Chinese and English training data can be directly mixed for training, and multi-scenario adaptation can be performed in one training, significantly improving the recognition accuracy and scalability of Chinese and English speech recognition tasks.
[0046] Next, refer to Figure 4 , which shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of this application. Figure 4 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.
[0047] As Figure 4 shown, the computer system includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage section 408 into the random access memory (RAM) 403. In the RAM 403, various programs and data required for system operation are also stored. The CPU 401, ROM 402, and RAM 403 are connected to each other through a bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.
[0048] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including a liquid crystal display (LCD) and the like, as well as speakers; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 410 as needed so that a computer program read therefrom is installed in the storage section 408 as needed.
[0049] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable storage medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 409 and / or installed from the removable medium 411. When the computer program is executed by the central processing unit (CPU) 401, the above functions defined in the methods of the present application are performed. It should be noted that the computer-readable storage medium of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable storage medium other than the computer-readable storage medium, and the computer-readable storage medium can send, propagate, or transmit a program for use by or combined with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0050] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0051] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0052] The modules described in the embodiments of this application can be implemented in software or in hardware.
[0053] As another aspect, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or may exist alone without being assembled into the electronic device. The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to: input speech into the Zip-MoE model, perform downsampling operations on adjacent frames of the speech, learn the weighted weights of the output of the previous encoder block and the output of the current encoder block through the Bypass module, and perform weighted fusion on the speech according to the weights learned by the Bypass module to reduce the frame rate of the speech; train the language router through the CTC loss function, the encoder block learns the mixed-language representation of the speech, and restores the frame rate of the speech through upsampling operations; input the mixed-language representation into the grouped mixture-of-experts layer, the language router outputs the language classification result of each frame of speech, and decouples the mixed-language representation into a Chinese representation and an English representation based on the language classification result; the unsupervised router routes the Chinese representation and the English representation to the Chinese expert group and the English expert group respectively, processes the Chinese representation and the English representation through the Softmax function to obtain the unsupervised routing weights of the language groups, and fuses the k expert indices of the language groups according to the unsupervised routing weights to obtain the output of the language groups; sum up the outputs of the language groups to obtain the final output of the grouped mixture-of-experts layer, and perform Chinese-English recognition of the speech according to the final output.
[0054] Finally, it should be noted that the above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the present application.
Claims
1. A Chinese and English speech recognition system based on Zip-MoE model grouping mixed expert layer, characterized in that: include: The Zip-MoE model includes 6 encoder blocks, and a Bypass module is included between every two encoder blocks, and the Bypass module is used to learn the weight of the previous encoder block output and the current encoder block output; Among them, the first three encoder blocks are standard Zipformer structures; The last three encoder blocks adopt the Zipformer-MoE structure, which includes a grouped mixed expert layer, which is used to replace the last feedforward network FNN of the standard Zipformer structure; The grouped mixed expert layer includes a Chinese expert group, an English expert group and a language router. The Chinese expert group and the English expert group are both composed of a number of expert networks and are respectively configured with an independent unsupervised router.
2. The system according to claim 1, characterized in that The language router is set after the downsampling operation of the Zipformer-MoE structure encoder block, and the language router is trained by the CTC loss function to obtain the language of each frame of speech. The training formula of the language router is expressed as: ; in, represents the loss function of the language router, Word-level labels indicating language identification, represents the speech features after downsampling, represents the weight of the linear classification layer for the language identification task, D represents the embedding dimension, 3 represents Chinese, English, and a blank tag for CTC.
3. The system according to claim 1, characterized in that The language router outputs the language classification result of each frame of speech based on the speech features of the current Zipformer-MoE structure coding block , and based on the language classification results The mixed language representation is decoupled to obtain Chinese representation and English representation. The decoupling formula is: ; in, Represents Chinese representation, , Indicates English representation, , Represents mixed language representation, Indicates the language classification result. , Indicates the number of voice frames routed to the Chinese expert group. represents the number of speech frames routed to the English expert group, T represents the total number of speech frames, represents the Chinese mask matrix, represents the English mask matrix, B represents the batch size, D represents the embedding dimension, represents element-wise multiplication, is the indicator function used to obtain the language mask matrix at the frame level.
4. The system according to claim 1, characterized in that The unsupervised router is a fully connected structure. The unsupervised router processes the Chinese expert group and the English expert group using the Softmax function to obtain an unsupervised routing weight. The unsupervised routing weight is expressed as: ; in, represents the weight of the unsupervised router, s represents the attribute of the language group, ,zh means Chinese expert group,en means English expert group, Represents mixed language representations, including Chinese representations and English representations, Represents an unsupervised router.
5. A method for Chinese and English speech recognition according to the system as claimed in any one of claims 1 to 4, characterized in that: include: S1: Input speech into the Zip-MoE model, perform downsampling operation on adjacent frames of speech, learn the weighted weight of the previous encoder block output and the current encoder block output through the Bypass module, and perform weighted fusion on the speech according to the weight learned by the Bypass module to reduce the frame rate of the speech; S2: The language router is trained by a CTC loss function, the encoder block learns the mixed language representation of the speech, and restores the frame rate of the speech by upsampling operation; S3: inputting the mixed language representation into the grouped mixed expert layer, the language router outputting the language classification result of each frame of speech, and decoupling the mixed language representation into Chinese representation and English representation based on the language classification result; S4: the unsupervised router routes the Chinese representation and the English representation to the Chinese expert group and the English expert group respectively, processes the Chinese representation and the English representation by a Softmax function to obtain an unsupervised routing weight of the language group, and fuses k expert indexes of the language group according to the unsupervised routing weight to obtain an output of the language group; S5: The outputs of the language groups are summed up to obtain the final output of the grouped mixed expert layer, and Chinese and English speech recognition is performed based on the final output.
6. The method according to claim 5, characterized in that The k expert indexes are selected through the Top-k strategy, and the specific formula is as follows: ; in, Representing mixed language representations The k expert indexes i corresponding to the speech frame t in , , is the Top-K expert set corresponding to language group s at speech frame t.
7. The method according to claim 6, characterized in that The language group includes a Chinese expert group and an English expert group, and the output of the language group is represented as: ; in, The output of the language group, represents the i-th expert network in language group s, represents the unsupervised routing weight of the i-th expert in language group s, , Representing mixed language representations k expert indexes i corresponding to speech frame t in; The final output of the grouped hybrid expert layer is expressed as: ; in, represents the final output of the grouped hybrid expert layer, , Output representing the language group.
8. The method according to claim 5, characterized in that The unsupervised router adopts an implicit learning mechanism to autonomously learn the feature distribution of input data and dynamically assign speech frames to the best matching expert.
9. A computer program product having one or more computer programs thereon, characterized in that: When the computer program is executed by a computer processor, the method according to any one of claims 5 to 8 is implemented.
Citation Information
Patent Citations
Chinese and English hybrid speech recognition model training method and device
CN111816169A
Speech recognition method, speech recognition model self-supervision pre-training method and electronic equipment
CN118248130A
Multilingual speech recognition method and device and storage medium
CN118351830A
Speech recognition method and device, electronic equipment and storage medium
CN119943050A
Mixture-of-expert conformer for streaming multilingual asr
WO2024186965A1
Cited By
Speech recognition method and device
CN120319222A
Optimization method and device of hybrid expert system, computer equipment and readable storage medium
CN120449952A
Large language model training method and reasoning method
CN120875045A
Speech recognition method and device, electronic equipment and storage medium
CN121528218A
Dialect emotion speech synthesis method
CN121922102A