Speech recognition method and apparatus, computer-readable medium, and electronic device
By pruning the parameters of the multilingual pre-trained model and constructing a sparse sub-network for cross-lingual adaptive training, the language interference problem in cross-lingual representation learning is solved, and the performance and efficiency of multilingual speech recognition are improved.
Patent Information
- Application Number
- CN202210204891.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-03
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2042-03-03
AI Technical Summary
Existing cross-linguistic representation learning methods have failed to effectively address the interference problem between different languages, leading to a decline in multilingual speech recognition performance, especially in the recognition of major languages and minor languages.
By pruning the parameters of the multilingual pre-trained model, a sparse subnetwork is constructed, forming a sparse subnetwork that shares some parameters. Cross-language adaptive training is then performed to achieve specific modeling for different languages.
In the process of cross-linguistic representation learning, it significantly improves the speech recognition performance of major and minor languages, reduces computational costs, and improves the deployment efficiency of the model on terminal devices.
Smart Images

Figure CN114582329B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the computer technology field, and in particular, to a speech recognition method, device, computer readable medium and electronic equipment. BACKGROUND
[0002] The current cross-language representation learning method does not consider the diversity of pronunciation between different languages, and still uses the model structure of monolingual representation learning. The model does not have a module specially used for modeling the characteristics of a specific language, so it often faces the problem of mutual interference between languages. This problem will be more serious when the number of languages increases and the unsupervised data increases. When this multi-language pre-training model is used for a downstream multi-language speech recognition task, it will cause the recognition performance of large language (such as Chinese and English) to decrease significantly, and there is a large gap compared with a monolingual pre-training model.
[0003] Therefore, there is an urgent need for a speech recognition model that can solve the language interference problem of cross-language representation learning from the perspective of self-adaptation. SUMMARY
[0004] This summary is provided to introduce a selection of concepts, which are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in limiting the scope of the claimed subject matter.
[0005] In a first aspect, the present disclosure provides a speech recognition method, comprising: obtaining a target speech signal to be recognized containing multiple languages; recognizing the semantics of the target speech signal by a speech recognition model fusing sparse sub-networks of various languages; the sparse sub-networks are obtained by parameter pruning processing on a multi-language pre-training model, and the multi-language pre-training model is trained according to speech signals containing the multiple languages.
[0006] In a second aspect, the present disclosure provides a speech recognition device, comprising: an obtaining module configured to obtain a target speech signal to be recognized containing multiple languages; and a recognition module configured to recognize the semantics of the target speech signal by a speech recognition model fusing sparse sub-networks of various languages; the sparse sub-networks are obtained by parameter pruning processing on a multi-language pre-training model, and the multi-language pre-training model is trained according to speech signals containing the multiple languages.
[0007] In a third aspect, the present disclosure provides a computer readable medium having stored thereon a computer program, which, when executed by a processing device, implements the steps of the speech recognition method described above.
[0008] In a fourth aspect, the present disclosure provides a computer device, comprising: a storage device having a computer program stored thereon; and a processing device configured to execute the computer program in the storage device to implement the steps of the voice recognition method described above.
[0009] By the above technical solution, the target voice signal containing multiple languages is obtained, and the semantic of the target voice signal is recognized by fusing a sparse sub-network voice recognition model of various languages. The sparse sub-network is obtained by parameter pruning processing on a multi-language pre-training model, and the multi-language pre-training model is trained according to voice signals containing multiple languages. The present disclosure solves the language interference problem of cross-language representation learning from the adaptive perspective, respectively performs parameter pruning processing on the entire multi-language pre-training model for different languages, constructs a group of sparse sub-networks sharing part of the parameters for training, thereby giving the voice recognition model the ability to model specifically for different languages, and greatly improving the large and small languages in the cross-language representation learning process.
[0010] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0011] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings in which:
[0012] Figure 1 FIG. 1 is a structural schematic diagram of a computer system according to an exemplary embodiment of the present disclosure.
[0013] Figure 2 FIG. 2 is a flowchart of a voice recognition method according to an exemplary embodiment of the present disclosure.
[0014] Figure 3 FIG. 3 is a flowchart of a training method of a voice recognition model according to an exemplary embodiment of the present disclosure.
[0015] Figure 4 FIG. 4 is a flowchart of a sub-step of step S202 according to an exemplary embodiment of the present disclosure.
[0016] Figure 5 FIG. 5 is a block diagram of a voice recognition device according to an exemplary embodiment of the present disclosure.
[0017] Figure 6 FIG. 6 is a structural schematic diagram of an electronic device according to an exemplary embodiment of the present disclosure.
[0018] BRIEF DESCRIPTION OF DRAWINGS
[0019] 120 - terminal; 140 - server; 20 - speech recognition apparatus; 201 - acquisition module; 203 - recognition module; 205 - processing module; 600 - computer device; 601 - processing apparatus; 602 - ROM; 603 - RAM; 604 - bus; 605 - I / O interface; 606 - input apparatus; 607 - output apparatus; 608 - storage apparatus; 609 - communication apparatus. DETAILED DESCRIPTION
[0020] Embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein, but rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It will be understood that the drawings of the present disclosure are for illustrative purposes only and are not intended to limit the scope of the present disclosure.
[0021] It should be understood that each of the steps recited in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit performing the steps shown. The scope of the present disclosure is not limited in this respect.
[0022] The term "comprising" and variations thereof as used in the present disclosure are open-ended, that is, "including but not limited to". The term "based on" is "based, at least in part, on". The term "one embodiment" means "at least one embodiment".
[0023] It should be noted that the modification of "one" or "multiple" mentioned in the present disclosure is illustrative rather than limiting, and those skilled in the art should understand that, unless otherwise explicitly indicated in the context, it should be understood as "one or more".
[0024] The names of the messages or information exchanged between the plurality of apparatuses in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0025] Figure 1 A structural schematic diagram of a computer system provided by one exemplary embodiment of the present disclosure is shown, which includes a terminal 120 and a server 140.
[0026] The terminal 120 and the server 140 are connected to each other through a wired or wireless network.
[0027] The terminal 120 can include at least one of a smartphone, a notebook computer, a desktop computer, a tablet computer, a smart speaker, and a smart robot.
[0028] The terminal 120 includes a display; the display can be used to display the speech recognition result.
[0029] The terminal 120 includes a first memory and a first processor. The first memory stores a first program; the first program is invoked by the first processor to implement the speech recognition method provided by the present disclosure. The first memory can include but is not limited to the following: RAM, ROM, PROM, EPROM, and EEPROM.
[0030] The first processor can be composed of one or more integrated circuit chips. Alternatively, the first processor can be a general-purpose processor, such as a CPU or an NP. Illustratively, the speech recognition model in the terminal can be trained by the terminal; or trained by a server, and obtained by the terminal from the server.
[0031] The server 140 includes a second memory and a second processor. The second memory stores a second program; the second program is invoked by the second processor to implement the speech recognition method provided by the present disclosure. Illustratively, the second memory stores a speech recognition model; the speech recognition model is invoked by the second processor to implement the speech recognition method. Alternatively, the second memory can include but is not limited to the following: RAM, ROM, PROM, EPROM, and EEPROM. Alternatively, the second processor can be a general-purpose processor, such as a CPU or an NP.
[0032] The server can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal can be a smartphone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, and the like, but is not limited thereto. The terminal and the server can be connected directly or indirectly through wired or wireless communication, and the present disclosure is not limited in this regard.
[0033] In recent years, pre-trained language models have developed rapidly, and the number of parameters of pre-trained language models has also increased, resulting in higher and higher computing costs. In order to improve the efficiency of pre-trained language models, various model compression methods have been proposed, including model pruning.
[0034] Based on this, the voice recognition method provided by an example embodiment of the present disclosure includes: obtaining a target voice signal containing multiple languages to be recognized, and identifying the semantics of the target voice signal by fusing a sparse sub-network voice recognition model of various languages; the sparse sub-network is obtained by performing parameter pruning processing on a multi-language pre-trained model, and the multi-language pre-trained model is trained according to voice signals containing multiple languages, which include some large languages such as Chinese and English, and some small languages such as French and Spanish. The present disclosure solves the language interference problem of cross-language representation learning from the perspective of adaptation, performs parameter pruning processing on the entire multi-language pre-trained model for different languages, constructs a group of sparse sub-networks sharing part of the parameters for training, thereby giving the voice recognition model the ability to model specifically for different languages, and greatly improving the performance of large and small languages in the cross-language representation learning process.
[0035] It should be noted that the voice recognition method provided by the present embodiment will be described in detail below, and the description not mentioned herein can be referred to the description below Figure 2 , which will not be repeated here.
[0036] Please refer to Figure 2 , Figure 2 the flowchart of the voice recognition method provided by an example embodiment of the present disclosure. The method is executed by a computer device, for example, by a terminal or a server in the computer system shown in Figure 1 . Figure 2 The voice recognition method shown in includes the following steps:
[0037] In step S101, a target voice signal containing multiple languages to be recognized is obtained.
[0038] It should be noted that the plurality of languages includes some large languages such as Chinese and English, and some small languages such as French and Spanish.
[0039] In step S102, the semantic of the target speech signal is recognized by fusing the speech recognition models of the sparse sub-networks of various languages.
[0040] The sparse sub-network is obtained by parameter pruning processing on the multi-language pre-training model, and the multi-language pre-training model is trained according to the speech signal containing a plurality of languages.
[0041] Referring to Figure 3 , Figure 3 The flowchart of the training method of the speech recognition model provided for one exemplary embodiment of the present disclosure. The training method of the speech recognition model includes the following steps:
[0042] In step S201, the speech signal containing a plurality of languages is obtained as a training sample.
[0043] It should be noted that the plurality of languages includes some large languages such as Chinese and English, and some small languages such as French and Spanish. The sample quantity of the large language is much larger than that of the small language, for example, the sample quantity of the large language can be 1 million, and the sample quantity of the small language can be only a few ten thousand. And the training sample is unsupervised speech data, that is, the speech data without manual annotation.
[0044] In step S202, a multi-language pre-training model is trained according to the training sample.
[0045] It should be noted that the multi-language pre-training model is used for speech recognition, and the multi-language pre-training model can recognize different languages respectively.
[0046] In the training stage, the Wav2vec 2.0 framework is used for cross-language speech representation learning, which mainly includes a feature extractor, a context network and a quantization module. The feature extractor is composed of a multi-layer convolutional neural network; the context network is composed of multiple layers of Transformer layers, which is used for semantic learning of the output speech signal of the feature extractor, and outputs a representation vector with context information; the quantization module is used to provide original unsupervised speech data for contrastive learning, and to quantize the output speech signal of the feature extractor. In order to stabilize the speech representation learning, a diversity loss is additionally added to promote the use of the quantization module by the multi-language pre-training model, and to avoid the collapse phenomenon of the quantization module.
[0047] It should be noted that step S202 includes sub-step S2021, sub-step S2022, sub-step S2023, sub-step S2024 and sub-step S2025, and the specific way of training the multilingual pre-training model will be described in detail in the sub-steps of step S202. Please refer to Figure 4 , Figure 4 is a flowchart of the sub-steps of step S202 according to an example embodiment of the present disclosure.
[0048] In sub-step S2021, the speech signal is converted into a plurality of low-dimensional signal frames.
[0049] It should be noted that before converting the speech signal into a plurality of low-dimensional signal frames, the training samples of the language whose sample quantity is below the first threshold value need to be up-sampled to expand the training sample quantity of the language below the first threshold value in the sampling data; for example, the sample quantity of a small language may only have tens of thousands of samples, so the training samples of the small language need to be up-sampled. The language whose sample quantity is above the second threshold value is uniformly sampled, for example, the sample quantity of a large language can be 1 million, at this time, only uniform sampling is needed to obtain sufficient sampling data.
[0050] The input speech signal is converted into a plurality of low-dimensional signal frames by the aforementioned feature extractor, wherein each signal frame is a fixed-length speech representation signal. For example, each signal frame can be about 25 ms long with a step of 20 ms.
[0051] In sub-step S2022, any one of the plurality of signal frames is masked to obtain a masked speech signal.
[0052] For example, a certain speech signal is divided into 10 frames in total, so any one or two of the 10 signal frames can be masked. If a certain speech signal is divided into 100 frames in total, then any 10 or 20 of the 100 signal frames can be masked to obtain a masked speech signal.
[0053] In sub-step S2023, the masked speech signal is input into the initial multilingual pre-training model for semantic learning to predict the masked signal frame.
[0054] The masked speech signal is subjected to semantic learning by the aforementioned context network, and the masked signal frame is reconstructed according to the result of semantic learning to predict the masked signal frame. At the same time, the quantization module receives the speech signal output by the feature extractor, which is not masked, so the predicted masked signal frame can be compared with the unsupervised speech data provided by the quantization module.
[0055] In sub-step S2024, when the predicted masked signal frame is consistent with the actual masked signal frame, it is determined that the prediction is correct and the parameters of the initial multilingual pre-training model are updated.
[0056] When the predicted masked signal frame is consistent with the actual masked signal frame output by the quantization module, it is determined that the signal frame predicted by the context network is correct, and the parameters of the initial multilingual pre-training model are updated.
[0057] In sub-step S2025, the step of updating the parameters of the initial multilingual pre-training model is repeatedly performed to obtain the multilingual pre-training model.
[0058] The sub-steps S2022-S2024 are repeatedly performed, and each time the masked signal frame is not repeated, then the semantic learning of the masked speech signal is performed by the context network to predict the masked signal frame, the predicted masked signal frame is compared with the actual masked signal frame output by the quantization module, when the comparison is consistent, it is determined that the signal frame predicted by the context network is correct, and the parameters of the initial multilingual pre-training model are updated to obtain the multilingual pre-training model.
[0059] For example, the sub-steps S2022-S2024 can be repeatedly performed a predetermined number of times, which can be a predetermined proportion of the number of signal frames of the speech signal, such as 50% of 10 signal frames, that is, 5 times. The predetermined number of times can be obtained based on human experience or according to other feasible methods, which will not be described here.
[0060] In step S203, the multilingual pre-training model is subjected to parameter pruning processing corresponding to a plurality of languages respectively to obtain a sparse sub-network corresponding to each language.
[0061] It should be noted that the parameter pruning processing mainly includes the following two ways, a lottery ticket hypothesis way and a Taylor expansion way. In the present disclosure, any one of the pruning ways can be adopted, and the two pruning ways will be described in detail below.
[0062] The step of performing parameter pruning on the multilingual pre-training model corresponding to the plurality of languages based on the lottery ticket hypothesis includes: taking the voice signals of each language as training samples to train the multilingual pre-training model, wherein the multilingual pre-training model refers to the multilingual pre-training model trained in step S202, and the training convergence condition is the same as that in step S202; obtaining all parameters of the multilingual pre-training model corresponding to each language; constructing a parameter matrix according to the parameters; then constructing a masking matrix corresponding to the parameter matrix, the length and width of the masking matrix being consistent with the parameter matrix; then obtaining the absolute value of each parameter in the parameter matrix; pruning a predetermined proportion of parameters according to the size of the absolute value, for example, a predetermined proportion of parameters can be pruned from small to large according to the size of the absolute value, and the predetermined proportion can be 20%, 30%, or 50%, etc., which is not limited herein; finally, setting the masking state of the pruned parameters in the corresponding position of the masking matrix to a first value, and setting the masking state of the remaining positions to a second value, wherein the first value is 0 and the second value is 1 in an embodiment.
[0063] The step of performing parameter pruning on the multilingual pre-training model corresponding to the plurality of languages based on the Taylor expansion includes: taking the voice signals of each language as training samples to train the multilingual pre-training model, wherein the multilingual pre-training model refers to the multilingual pre-training model trained in step S202, and the training convergence condition is the same as that in step S202; obtaining all parameters of the multilingual pre-training model corresponding to each language; predicting the loss value caused by pruning each parameter on the multilingual pre-training model through first-order Taylor expansion of the parameter, wherein the formula for predicting the loss value caused by pruning each parameter on the multilingual pre-training model includes:
[0064] |g 2 w 2 |
[0065] wherein g is the gradient of the parameter, w is the weight of the parameter, and | is an absolute value operator; then pruning a predetermined proportion of parameters according to the size of the loss value, for example, a predetermined proportion of parameters can be pruned from small to large according to the size of the loss value, and the predetermined proportion can be 20%, 30%, or 50%, etc., which is not limited herein.
[0066] The multilingual pre-training model obtained after parameter pruning can accelerate the inference calculation speed, meet the minimum delay limit, reduce the consumed memory, and be more easily deployed on the terminal side, such as a mobile phone, and more convenient for model training and fine-tuning.
[0067] By any one of the above two pruning methods, the voice signals of each language are taken as training samples to train the multilingual pre-training model, and a sparse subnetwork corresponding to each language is obtained.
[0068] In step S204, the parameters of each sparse sub-network are updated by multi-language self-adaptive pre-training of each sparse sub-network corresponding to a language to obtain shared parameters and private parameters between each sparse sub-network.
[0069] In the training process, each training batch is composed of training samples of only one language, where the batch refers to a parameter update of model weights by one backward propagation using a small part of samples in the training samples, and the small part of samples is referred to as "a batch of data".
[0070] For input sample data of each language, only the sparse sub-network corresponding to the language is used for forward propagation, and sparse sub-network loss is calculated, and only the parameters corresponding to the sparse sub-network are updated during backward propagation. In this way, the final sparse sub-network automatically allocates shared parameters and private parameters between different languages in the network, thereby achieving the effect of self-adaptive training.
[0071] In step S205, a speech recognition model of a sparse sub-network of various languages is obtained according to the shared parameters and the private parameters.
[0072] The speech recognition model based on sparse shared sub-network cross-language speech representation proposed in the disclosure can exceed the baseline cross-language speech representation learning method on both large and small languages. On the disclosed Common Voice data set, the proposed speech recognition method has a relative 9.8% average phoneme error rate reduction compared to the baseline system on a 100M model, and a 7.4% phoneme error rate reduction on a 300M model. Moreover, this method can greatly alleviate the language interference problem suffered by large languages, with a relative 17.8% and 16.7% phoneme error rate reduction on the 100M model and the 300M model.
[0073] In summary, the speech recognition method provided by the disclosure includes obtaining a target speech signal containing multiple languages to be recognized, and recognizing the semantics of the target speech signal by a speech recognition model of a sparse sub-network of various languages, wherein the sparse sub-network is obtained by parameter pruning processing on a multi-language pre-training model, and the multi-language pre-training model is trained according to speech signals containing multiple languages. The disclosure solves the language interference problem of cross-language representation learning from the perspective of self-adaptation, performs parameter pruning processing on the entire multi-language pre-training model for different languages, constructs a group of sparse sub-networks sharing part of the parameters for training, thereby giving the speech recognition model the ability to model specifically for different languages, and greatly improving the large and small languages in the cross-language representation learning process.
[0074] Figure 5is a speech recognition device block diagram shown in one exemplary embodiment of the present disclosure. Referring to Figure 5 The device 20 comprises an acquisition module 201 and an identification module 203.
[0075] The acquisition module 201 is configured to acquire a target speech signal to be identified containing multiple languages;
[0076] The identification module 203 is configured to identify the semantics of the target speech signal by fusing speech recognition models of sparse sub-networks of various languages; the sparse sub-networks are obtained by parameter pruning processing on a multi-language pre-training model, and the multi-language pre-training model is trained according to speech signals containing the multiple languages.
[0077] The device 20 further comprises a processing module 205.
[0078] Optionally, the processing module 205 is configured to acquire speech signals containing the multiple languages as training samples;
[0079] The multi-language pre-training model is trained according to the training samples; and the multi-language pre-training model is used for speech recognition.
[0080] The multi-language pre-training model is subjected to parameter pruning processing corresponding to the multiple languages respectively to obtain a sparse sub-network corresponding to each language;
[0081] Parameters of each sparse sub-network are updated by multi-language adaptive pre-training of the corresponding language to obtain shared parameters and exclusive parameters between the sparse sub-networks;
[0082] A speech recognition model of sparse sub-networks of various languages is obtained according to the shared parameters and the exclusive parameters.
[0083] Optionally, the processing module 205 is further configured to convert the speech signal into a plurality of low-dimensional signal frames; the signal frames are speech representation signals of a fixed time length;
[0084] Any one of the plurality of signal frames is masked to obtain a masked speech signal;
[0085] The masked speech signal is input into an initial multi-language pre-training model for semantic learning to predict a masked signal frame;
[0086] When the predicted masked signal frame is consistent with the actual masked signal frame, it is determined that the prediction is correct, and parameters of the initial multi-language pre-training model are updated;
[0087] The step of updating the parameters of the initial multi-language pre-training model is repeatedly performed to obtain the multi-language pre-training model.
[0088] Optionally, the processing module 205 is further configured to up-sample training samples of a language whose sample quantity is lower than a first threshold, so as to expand the training sample quantity of the language whose sample quantity is lower than the first threshold in the sampling data.
[0089] Uniform sampling is performed on a language whose sample quantity is higher than a second threshold.
[0090] Optionally, the processing module 205 is further configured to perform parameter pruning processing on the multilingual pre-training model corresponding to each of the plurality of languages based on a lottery assumption manner, to obtain a sparse sub-network corresponding to each language.
[0091] Or perform parameter pruning processing on the multilingual pre-training model corresponding to each of the plurality of languages based on a Taylor expansion manner, to obtain a sparse sub-network corresponding to each language.
[0092] Optionally, the processing module 205 is further configured to train the multilingual pre-training model by taking the voice signal of each language as a training sample.
[0093] Obtain all parameters of the multilingual pre-training model corresponding to each language.
[0094] Construct a parameter matrix according to the parameters.
[0095] Construct a masking matrix corresponding to the parameter matrix.
[0096] Obtain the absolute value of each parameter in the parameter matrix.
[0097] According to the size of the absolute value, the parameters of a predetermined proportion are pruned.
[0098] The masking state of the pruned parameters at the corresponding position in the masking matrix is set to a first value, and the masking state of the remaining positions is set to a second value.
[0099] Optionally, the processing module 205 is further configured to train the multilingual pre-training model by taking the voice signal of each language as a training sample.
[0100] Obtain all parameters of the multilingual pre-training model corresponding to each language.
[0101] After first-order Taylor expansion of the parameters, predict the loss value caused by pruning of each parameter to the multilingual pre-training model.
[0102] According to the size of the loss value, the parameters of a predetermined proportion are pruned.
[0103] Reference will be made to the following description Figure 6which shows a structural diagram of an electronic device (e.g., a terminal device or a server) 600 suitable for use in implementing embodiments of the present disclosure. The terminal device in embodiments of the present disclosure can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a car terminal (e.g., a car navigation terminal), and the like, as well as a stationary terminal such as a digital TV, a desktop computer, and the like. Figure 1 The electronic device shown is merely one example, and should not be taken as limiting the functionality or use of embodiments of the present disclosure. Figure 6 The electronic device shown is merely one example, and should not be taken as limiting the functionality or use of embodiments of the present disclosure.
[0104] As shown in FIG. 6, Figure 6 The electronic device 600 can include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 602 or programs loaded from a storage device 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0105] In general, the following devices can be connected to the I / O interface 605: input devices 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 608 including, for example, a magnetic tape, a hard disk, and the like; and communication devices 609. The communication devices 609 can allow the electronic device 600 to communicate wirelessly or via a wire with other devices to exchange data. Although Figure 6 The electronic device 600 is shown with various devices, but it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices can alternatively be implemented or present.
[0106] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication devices 609, or installed from the storage devices 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-described functions defined in the methods of the present disclosure are performed.
[0107] Note that the computer readable medium described above in the present disclosure can be a computer readable signal medium or a computer readable storage medium or any combination thereof. The computer readable storage medium can be, for example and without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination thereof. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the present disclosure, the computer readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer readable program code is contained. Such a propagated data signal can take any of a variety of forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer readable signal medium can also be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device. The program code contained on the computer readable medium can be transmitted by any suitable medium, including but not limited to wire, cable, RF, etc., or any suitable combination thereof.
[0108] In some embodiments, the terminals, servers can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communications of any form or medium (e.g., a communications network). Examples of communications networks include local area networks ("LANs"), wide area networks ("WANs"), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future developed networks.
[0109] The computer readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device and be not assembled in the electronic device.
[0110] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: acquire a target speech signal to be recognized containing multiple languages; and recognize semantics of the target speech signal by fusing speech recognition models of sparse sub-networks of various languages; the sparse sub-networks are obtained by performing parameter pruning processing on a multi-language pre-training model, and the multi-language pre-training model is trained according to speech signals containing the multiple languages.
[0111] Computer program code for carrying out operations of the present disclosure can be written in any one or more of a variety of programming languages or combinations of languages including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0112] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in a different order than that noted in the figures. For example, two blocks noted in succession can in fact be executed substantially concurrently or in the opposite order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flow diagrams, and combinations of blocks in the block diagrams and / or flow diagrams, can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or can be implemented by a combination of dedicated hardware-based systems and computer instructions.
[0113] The modules involved in the embodiments of the present disclosure can be implemented in a software manner or in a hardware manner. In some cases, the name of the module does not constitute a limitation on the module itself.
[0114] The functionality described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0115] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0116] According to one or more embodiments of the present disclosure, example 1 provides a speech recognition method, comprising: obtaining a target speech signal containing multiple languages to be recognized;
[0117] recognizing the semantic of the target speech signal by fusing a speech recognition model of sparse sub-networks of various languages;
[0118] The sparse sub-networks are obtained by parameter pruning processing on a multi-language pre-training model, and the multi-language pre-training model is trained according to speech signals containing the multiple languages.
[0119] According to one or more embodiments of the present disclosure, example 2 provides the method of example 1, and the training method of the speech recognition model comprises the following steps:
[0120] obtaining speech signals containing the multiple languages as training samples;
[0121] training the multi-language pre-training model according to the training samples; the multi-language pre-training model is used for speech recognition;
[0122] performing parameter pruning processing on the multi-language pre-training model corresponding to the multiple languages respectively to obtain a sparse sub-network corresponding to each language;
[0123] updating parameters of each of the sparse sub-networks by multi-language self-adaptive pre-training of the sparse sub-networks corresponding to the languages to obtain shared parameters and private parameters between the sparse sub-networks;
[0124] obtaining a speech recognition model of the sparse sub-networks of various languages according to the shared parameters and the private parameters.
[0125] According to one or more embodiments of the present disclosure, example 3 provides the method of example 2, wherein the step of training the multi-language pre-training model according to the training samples comprises:
[0126] converting the speech signal into a plurality of low-dimensional signal frames; the signal frames are fixed-length speech representation signals;
[0127] masking any one of the plurality of signal frames to obtain a masked speech signal;
[0128] inputting the masked speech signal into an initial multi-language pre-training model for semantic learning to predict the masked signal frame;
[0129] when the predicted masked signal frame is consistent with the actual masked signal frame, determining that the prediction is correct and updating parameters of the initial multi-language pre-training model;
[0130] repeating the step of updating the parameters of the initial multi-language pre-training model to obtain the multi-language pre-training model.
[0131] According to one or more embodiments of the present disclosure, example 4 provides the method of example 3, wherein the step of training the multi-language pre-training model according to the training samples further comprises:
[0132] upsampling the training samples of a language whose sample number is lower than a first threshold to expand the number of training samples of the language whose sample number is lower than the first threshold in the sampling data;
[0133] uniformly sampling a language whose sample number is higher than a second threshold.
[0134] According to one or more embodiments of the present disclosure, example 5 provides the method of example 2, wherein the step of performing parameter pruning processing on the multi-language pre-training model corresponding to the plurality of languages respectively to obtain a sparse sub-network corresponding to each language comprises:
[0135] performing parameter pruning processing on the multi-language pre-training model corresponding to the plurality of languages respectively based on a lottery ticket hypothesis to obtain a sparse sub-network corresponding to each language;
[0136] or the multi-language pre-training model corresponding to the plurality of languages is respectively pruned based on a Taylor expansion manner to obtain a sparse sub-network corresponding to each language.
[0137] According to one or more embodiments of the present disclosure, example 6 provides the method of example 5, wherein the step of pruning the parameters of the multi-language pre-training model corresponding to the plurality of languages based on the lottery ticket hypothesis manner comprises:
[0138] training the multi-language pre-training model respectively based on the voice signals of each language as training samples;
[0139] obtaining all parameters of the multi-language pre-training model corresponding to each language;
[0140] constructing a parameter matrix according to the parameters;
[0141] constructing a masking matrix corresponding to the parameter matrix;
[0142] obtaining the absolute values of each parameter in the parameter matrix;
[0143] cutting a predetermined proportion of the parameters according to the absolute values;
[0144] setting the masking state of the cut parameters in the corresponding position of the masking matrix to a first value, and setting the masking state of the remaining positions to a second value.
[0145] According to one or more embodiments of the present disclosure, example 7 provides the method of example 5, wherein the step of pruning the parameters of the multi-language pre-training model corresponding to the plurality of languages based on the Taylor expansion manner comprises:
[0146] training the multi-language pre-training model respectively based on the voice signals of each language as training samples;
[0147] obtaining all parameters of the multi-language pre-training model corresponding to each language;
[0148] predicting the loss value caused by pruning each parameter to the multi-language pre-training model after expanding the parameters by the first-order Taylor expansion;
[0149] cutting a predetermined proportion of the parameters according to the loss values.
[0150] According to one or more embodiments of the present disclosure, example 8 provides the method of example 7, wherein the formula for predicting the loss value caused by pruning each parameter to the multi-language pre-training model comprises:
[0151] |g 2 w 2 |
[0152] Where g is the gradient of the parameter, and w is the weight of the parameter.
[0153] According to one or more embodiments of the present disclosure, Example 9 provides a speech recognition device, including: an acquisition module, configured to acquire a target speech signal to be recognized containing multiple languages;
[0154] The recognition module is used to recognize the semantics of the target speech signal through a speech recognition model that integrates sparse sub-networks of various languages;
[0155] The sparse subnetwork is obtained by pruning the parameters of a multilingual pre-trained model, which is trained based on speech signals containing the multiple languages.
[0156] According to one or more embodiments of the present disclosure, Example 10 provides a computer-readable medium having a computer program stored thereon that, when executed by a processing device, implements the steps of the aforementioned speech recognition method.
[0157] According to one or more embodiments of this disclosure, Example 11 provides an electronic device, including: a storage device having a computer program stored thereon; and a processing device for executing the computer program in the storage device to implement the steps of the aforementioned speech recognition method.
[0158] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0159] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0160] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims. With respect to the devices in the above-described embodiments, in which various modules perform operations, the specific manner in which the various modules perform the operations has been described in detail in the embodiments relating to the method, and will not be described in detail here.
Claims
1. A voice recognition method, characterized by, The method comprises the steps of: acquiring a target speech signal containing multiple languages to be recognized; recognizing the semantic of the target speech signal by a speech recognition model fusing sparse sub-networks of various languages; the sparse sub-networks are obtained by parameter pruning processing of a multi-language pre-training model, the multi-language pre-training model is trained according to speech signals containing the multiple languages, and the parameter pruning processing mainly includes any one of the following pruning methods: a lottery ticket hypothesis method and a Taylor expansion method.
2. The method of claim 1, wherein, The training method of the speech recognition model comprises the following steps: acquiring speech signals containing the multiple languages as training samples; training the multi-language pre-training model according to the training samples; the multi-language pre-training model is used for speech recognition; performing parameter pruning processing on the multi-language pre-training model corresponding to the multiple languages respectively to obtain sparse sub-networks corresponding to each language; updating the parameters of each sparse sub-network by multi-language adaptive pre-training of the corresponding language to obtain shared parameters and exclusive parameters between the sparse sub-networks; obtaining a speech recognition model fusing sparse sub-networks of various languages according to the shared parameters and the exclusive parameters.
3. The method of claim 2, wherein, The step of training the multi-language pre-training model according to the training samples comprises the following steps: convert the speech signals into multiple low-dimensional signal frames; the signal frames are fixed-length speech representation signals; mask any one of the multiple signal frames to obtain a masked speech signal; input the masked speech signal into an initial multi-language pre-training model for semantic learning to predict the masked signal frame; when the predicted masked signal frame is consistent with the actual masked signal frame, it is determined that the prediction is correct and the parameters of the initial multi-language pre-training model are updated; repeat the step of updating the parameters of the initial multi-language pre-training model to obtain the multi-language pre-training model.
4. The method of claim 3, wherein, The step of training the multi-language pre-training model according to the training samples further comprises the following steps: perform up-sampling on the training samples of a language whose sample number is lower than a first threshold to expand the sample number of the training samples of the language lower than the first threshold in the sampling data; perform uniform sampling on a language whose sample number is higher than a second threshold.
5. The method of claim 2, wherein, The step of performing parameter pruning processing on the multi-language pre-training model corresponding to the multiple languages respectively to obtain sparse sub-networks corresponding to each language comprises the following steps: perform parameter pruning processing on the multi-language pre-training model corresponding to the multiple languages respectively based on the lottery ticket hypothesis method to obtain sparse sub-networks corresponding to each language; or perform parameter pruning processing on the multi-language pre-training model corresponding to the multiple languages respectively based on the Taylor expansion method to obtain sparse sub-networks corresponding to each language.
6. The method of claim 5, wherein, The step of performing parameter pruning processing on the multi-language pre-training model corresponding to the multiple languages respectively based on the lottery ticket hypothesis method comprises the following steps: train the multi-language pre-training model respectively by taking the speech signals of each language as training samples; acquire all parameters of the multi-language pre-training model corresponding to each language; construct a parameter matrix according to the parameters; constructing a masking matrix corresponding to the parameter matrix; obtaining absolute values of each parameter in the parameter matrix; clipping a predetermined proportion of the parameters according to the size of the absolute values; setting the masking state of the clipped parameters at the corresponding positions in the masking matrix to a first value, and setting the masking state of the remaining positions to a second value.
7. The method of claim 5, wherein, The step of performing parameter pruning on the multilingual pre-training model corresponding to the plurality of languages based on the Taylor expansion method includes: training the multilingual pre-training model by taking the voice signals of each language as training samples; obtaining all parameters of the multilingual pre-training model corresponding to each language; predicting the loss value caused by pruning each parameter of the multilingual pre-training model by first-order Taylor expansion of the parameter; clipping a predetermined proportion of the parameters according to the size of the loss value.
8. A speech recognition apparatus characterized by comprising: It includes: an acquisition module configured to acquire a target voice signal containing a plurality of languages to be recognized; an identification module configured to identify the semantic of the target voice signal by fusing a sparse subnetwork voice recognition model of various languages. The sparse subnetwork is obtained by performing parameter pruning on a multilingual pre-training model, the multilingual pre-training model is trained according to voice signals containing the plurality of languages, and the parameter pruning mainly includes any one of the following pruning methods: a lottery ticket hypothesis method and a Taylor expansion method.
9. A computer readable medium having stored thereon a computer program, characterized in that The computer program is executed by the processing device to realize the steps of the method of any one of claims 1-7.
10. An electronic device, comprising: It includes: a storage device having a computer program stored thereon; a processing device configured to execute the computer program in the storage device to realize the steps of the method of any one of claims 1-7.
Citation Information
Patent Citations
Putonghua and Cantonese hybrid speech recognition model training method and system
CN111816160A