Speech recognition model processing method, device, equipment and storage medium
By combining federated learning and evolutionary algorithms, the speech recognition model parameters of multiple target scenarios are integrated, which solves the problem of poor speech recognition performance in multiple scenarios and achieves more efficient speech recognition results.
Patent Information
- Application Number
- CN202111629366.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-28
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2041-12-28
AI Technical Summary
Existing speech recognition models have difficulty effectively recognizing audio data in multiple different scenarios, resulting in poor recognition results.
Adopting the concept of federated learning, the speech recognition models of multiple target scenarios are obtained through the RPA platform. The model parameters are iterated multiple times using the evolutionary algorithm, and the model parameters of multiple target scenarios are integrated to form a federated speech recognition model.
The recognition effect of the speech recognition model in multiple scenarios has been improved, the breadth and depth of recognition have been increased, and the accuracy of recognition has been improved.
Smart Images

Figure CN114333803B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device and storage medium for processing a speech recognition model. Background Art
[0002] Automatic speech recognition (ASR) studies speech, enabling machines to automatically recognize and understand spoken language through speech signal processing and pattern recognition. Speech recognition technology is a high-tech process that allows machines to convert speech signals into corresponding text or commands through the recognition and understanding process. Speech recognition is a broad interdisciplinary field, closely related to acoustics, phonetics, linguistics, information theory, pattern recognition theory, and neurobiology.
[0003] With the continuous development of deep learning technology, ASR is being applied to an increasing number of scenarios, such as smart home control, in-vehicle control, intelligent customer service, healthcare, education, and finance. The audio data corresponding to each scenario varies significantly, and universal speech recognition models cannot effectively recognize audio data from all scenarios. As AI applications continue to expand, this problem will become increasingly prominent. Summary of the Invention
[0004] The present invention provides a method, device, equipment and storage medium for processing a speech recognition model to improve the recognition effect of a multi-scene speech recognition model.
[0005] In a first aspect, the present invention provides a method for processing a speech recognition model, which is applied to a Robotic Process Automation (RPA) platform and includes:
[0006] In response to a model selection operation on a graphical user interface of the RPA platform, multiple speech recognition models for target scenarios are obtained from a database of the RPA platform, where the speech recognition model for each target scenario is obtained by training a basic speech recognition model through transfer learning based on an audio sample of the target scenario and a text annotation corresponding to the audio sample;
[0007] In response to a triggering operation of a model training control on the graphical user interface, model parameters of the speech recognition models for the multiple target scenarios are obtained, and model parameters of a federated speech recognition model are obtained based on the model parameters of the speech recognition models for the multiple target scenarios; the federated speech recognition model is used to recognize audio data of the multiple target scenarios, and the model parameters of the federated speech recognition model are determined by performing N iterations on the model parameters of the speech recognition models for the multiple target scenarios, where N is a positive integer;
[0008] The model parameters of the federated speech recognition model are stored in a preset path.
[0009] In an optional embodiment of the first aspect, obtaining the model parameters of the federated speech recognition model based on the model parameters of the speech recognition models of the multiple target scenarios includes:
[0010] Obtaining model parameters of the speech recognition models of the multiple target scenarios to obtain multiple groups of model parameters;
[0011] Using an evolutionary algorithm to adjust the multiple sets of model parameters;
[0012] The model parameters of the federated speech recognition model are determined according to the adjusted multiple groups of model parameters.
[0013] In an optional embodiment of the first aspect, the adjusting the multiple sets of model parameters using an evolutionary algorithm includes:
[0014] randomly selecting at least two groups of model parameters from the multiple groups of model parameters;
[0015] The at least two groups of model parameters are cross-exchanged to obtain multiple groups of adjusted model parameters.
[0016] In an optional embodiment of the first aspect, the adjusting the multiple sets of model parameters using an evolutionary algorithm includes:
[0017] randomly selecting at least two groups of model parameters from the multiple groups of model parameters;
[0018] The at least two groups of model parameters are cross-exchanged, and the at least two groups of model parameters after the cross-exchange are randomly adjusted according to the coefficient of variation to obtain multiple groups of adjusted model parameters.
[0019] In an optional embodiment of the first aspect, determining the model parameters of the federated speech recognition model based on the adjusted multiple sets of model parameters includes:
[0020] Determining candidate model parameters of the federated speech recognition model based on the adjusted multiple sets of model parameters;
[0021] Obtaining evaluation parameters of the federated speech recognition model corresponding to the candidate model parameters;
[0022] If the evaluation parameter is within a preset evaluation parameter range, the candidate model parameter is used as the model parameter of the federated speech recognition model.
[0023] In an optional embodiment of the first aspect, determining candidate model parameters of the federated speech recognition model based on the adjusted multiple groups of model parameters includes:
[0024] The corresponding parameters in the adjusted multiple sets of model parameters are summed and averaged, or the corresponding parameters in the adjusted multiple sets of model parameters are weighted and summed according to preset weight values of different target scene model parameters to obtain candidate model parameters of the federated speech recognition model.
[0025] In an optional embodiment of the first aspect, the method further includes:
[0026] If the evaluation parameter is outside the preset evaluation parameter range, the step of adjusting the multiple sets of model parameters using the evolutionary algorithm is performed again until the evaluation parameter corresponding to the candidate model parameter is within the preset evaluation parameter range.
[0027] In an optional embodiment of the first aspect, the speech recognition models of the multiple target scenes include a speech recognition model of a first target scene, and the first target scene is any one of the multiple target scenes;
[0028] The training process of the speech recognition model for the first target scenario includes:
[0029] Obtaining an audio sample of the first target scene and a text annotation corresponding to the audio sample;
[0030] Acquire the basic speech recognition model, where the basic speech recognition model includes a basic acoustic model and a basic language model, and the output of the basic acoustic model corresponds to the input of the basic language model;
[0031] Using the audio sample of the first target scene as the input of the basic acoustic model and the text annotation corresponding to the audio sample as the output of the basic language model, to train the basic speech recognition model;
[0032] When the evaluation parameters of the speech recognition model of the first target scene are within a preset evaluation parameter range, the speech recognition model of the first target scene is obtained.
[0033] In an optional embodiment of the first aspect, obtaining a text annotation corresponding to the audio sample includes:
[0034] Acquire an audio sample of the first target scene input by a user, and generate a reference text of the audio sample using automatic speech recognition (ASR) technology;
[0035] In response to a user's modification operation on the reference text, a text annotation corresponding to the audio sample is obtained.
[0036] In a second aspect, the present invention provides a device for processing a speech recognition model, comprising:
[0037] an acquisition module, configured to respond to a model selection operation on a graphical user interface of the RPA platform and acquire speech recognition models for multiple target scenarios from a database of the RPA platform, wherein the speech recognition model for each target scenario is obtained by training a basic speech recognition model through transfer learning based on audio samples of the target scenario and text annotations corresponding to the audio samples;
[0038] a processing module, configured to, in response to a triggering operation of a model training control on the graphical user interface, obtain model parameters of the speech recognition models for the multiple target scenarios, and obtain model parameters of a federated speech recognition model based on the model parameters of the speech recognition models for the multiple target scenarios; the federated speech recognition model is configured to recognize audio data for the multiple target scenarios, and the model parameters of the federated speech recognition model are determined by performing N iterations on the model parameters of the speech recognition models for the multiple target scenarios, where N is a positive integer;
[0039] The storage module is used to store the model parameters of the federated speech recognition model in a preset path.
[0040] In an optional embodiment of the second aspect, the acquisition module is configured to acquire model parameters of the speech recognition models of the multiple target scenes to obtain multiple groups of model parameters;
[0041] A processing module, configured to adjust the plurality of model parameters using an evolutionary algorithm;
[0042] The model parameters of the federated speech recognition model are determined according to the adjusted multiple groups of model parameters.
[0043] In an optional embodiment of the second aspect, the processing module is configured to:
[0044] randomly selecting at least two groups of model parameters from the multiple groups of model parameters;
[0045] The at least two groups of model parameters are cross-exchanged to obtain multiple groups of adjusted model parameters.
[0046] In an optional embodiment of the second aspect, the processing module is configured to:
[0047] randomly selecting at least two groups of model parameters from the multiple groups of model parameters;
[0048] The at least two groups of model parameters are cross-exchanged, and the at least two groups of model parameters after the cross-exchange are randomly adjusted according to the coefficient of variation to obtain multiple groups of adjusted model parameters.
[0049] In an optional embodiment of the second aspect, the processing module is configured to:
[0050] Determining candidate model parameters of the federated speech recognition model based on the adjusted multiple sets of model parameters;
[0051] Obtaining evaluation parameters of the federated speech recognition model corresponding to the candidate model parameters;
[0052] If the evaluation parameter is within a preset evaluation parameter range, the candidate model parameter is used as the model parameter of the federated speech recognition model.
[0053] In an optional embodiment of the second aspect, the processing module is configured to:
[0054] The corresponding parameters in the adjusted multiple sets of model parameters are summed and averaged, or the corresponding parameters in the adjusted multiple sets of model parameters are weighted and summed according to preset weight values of different target scene model parameters to obtain candidate model parameters of the federated speech recognition model.
[0055] In an optional embodiment of the second aspect, the processing module is configured to:
[0056] If the evaluation parameter is outside the preset evaluation parameter range, the step of adjusting the multiple sets of model parameters using the evolutionary algorithm is performed again until the evaluation parameter corresponding to the candidate model parameter is within the preset evaluation parameter range.
[0057] In an optional embodiment of the second aspect, the speech recognition models of the multiple target scenes include a speech recognition model of a first target scene, and the first target scene is any one of the multiple target scenes;
[0058] an acquisition module, configured to acquire an audio sample of the first target scene and a text annotation corresponding to the audio sample; and acquire the basic speech recognition model, wherein the basic speech recognition model includes a basic acoustic model and a basic language model, and the output of the basic acoustic model corresponds to the input of the basic language model;
[0059] a processing module, configured to use the audio sample of the first target scene as the input of the basic acoustic model and the text annotation corresponding to the audio sample as the output of the basic language model, and train the basic speech recognition model;
[0060] When the evaluation parameters of the speech recognition model of the first target scene are within a preset evaluation parameter range, the speech recognition model of the first target scene is obtained.
[0061] In an optional embodiment of the second aspect, the acquisition module is configured to acquire an audio sample of the first target scene input by a user, and the processing module is configured to generate a reference text of the audio sample using automatic speech recognition (ASR) technology;
[0062] The acquisition module is further configured to acquire the text annotation corresponding to the audio sample in response to a user's modification operation on the reference text.
[0063] In a third aspect, the present invention provides an electronic device, comprising:
[0064] A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the method according to any one of the first aspects of the present invention is implemented.
[0065] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method as described in any one of the first aspects of the present invention is implemented.
[0066] In a fifth aspect, the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method as described in any one of the first aspects of the present invention.
[0067] An embodiment of the present invention provides a method, apparatus, device and storage medium for processing a speech recognition model. The method includes: responding to a model selection operation on a graphical user interface of an RPA platform, obtaining speech recognition models of multiple target scenarios from a database of the RPA platform, iterating N times on the model parameters of the speech recognition models of the multiple target scenarios to determine the model parameters of the federated speech recognition model, obtaining a final federated speech recognition model, and storing its model parameters in a preset path. The speech recognition model of each target scenario is obtained by training a basic speech recognition model through transfer learning based on the audio sample of the target scenario and the text annotation corresponding to the audio sample, and N is a positive integer. The federated speech recognition model obtained by the above scheme integrates the model parameters of multiple target scenarios, improves the breadth and depth of recognition of the speech recognition model, and has a better recognition effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 A schematic diagram of a scenario of a method for processing a speech recognition model provided by an embodiment of the present invention;
[0069] Figure 2 Schematic diagram of the process of processing the speech recognition model provided by the embodiment of the present invention Figure 1 ;
[0070] Figure 3 Schematic diagram of a graphical user interface provided by a server according to an embodiment of the present invention Figure 1 ;
[0071] Figure 4Schematic diagram of a graphical user interface provided by a server according to an embodiment of the present invention Figure 2 ;
[0072] Figure 5 Schematic diagram of the process of processing the speech recognition model provided by the embodiment of the present invention Figure 2 ;
[0073] Figure 6 Schematic diagram of a graphical user interface provided by a server according to an embodiment of the present invention Figure 3 ;
[0074] Figure 7 Schematic diagram of a graphical user interface provided by a server according to an embodiment of the present invention Figure 4 ;
[0075] Figure 8 Schematic diagram of a graphical user interface provided by a server according to an embodiment of the present invention Figure 5 ;
[0076] Figure 9 Schematic diagram of a graphical user interface provided by a server according to an embodiment of the present invention Figure 6 ;
[0077] Figure 10 Schematic diagram of a graphical user interface provided by a server according to an embodiment of the present invention Figure 7 ;
[0078] Figure 11 A structural block diagram of a processing device for a speech recognition model provided by an embodiment of the present invention;
[0079] Figure 12 This is a structural block diagram of an electronic device provided by an embodiment of the present invention.
[0080] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0081] Exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0082] The terms "first," "second," and the like in the description, claims, and accompanying drawings of the embodiments of the present invention are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, such that the embodiments of the present invention described herein can be practiced in an order other than that illustrated or described herein.
[0083] It should be understood that the terms "include" and "have" and any variations thereof as used herein are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus that includes a series of steps or elements is not necessarily limited to those steps or elements expressly listed, but may include other steps or elements not expressly listed or inherent to such process, method, product or apparatus.
[0084] In the description of the embodiments of the present invention, the term "correspondence" may indicate a direct or indirect correspondence between the two, or an association relationship between the two, or a relationship between indication and indication, configuration and configuration, etc.
[0085] Before introducing the technical solutions provided by the embodiments of the present invention, a brief introduction to the professional terms involved in the embodiments of the present invention is first given.
[0086] First, the acoustic model (AM) represents differentiated knowledge about acoustics, phonetics, environmental variables, speaker gender, and accent. This includes acoustic models based on the Hidden Markov Model (HMM), such as the Gaussian Mixture Hidden Markov Model (GMM-HMM) and the Deep Neural Network Hidden Markov Model (DNN-HMM). The HMM is a weighted finite state automaton in the discrete time domain. Of course, end-to-end acoustic models can also be included, such as the Continuous Time Classification-Long Short-Term Memory (CTC-LSTM) model and the attention model. Each state of the acoustic model represents the probability distribution of the speech features of a speech unit (such as a word, syllable, or factor) in that state. Transitions between states connect the states into an ordered sequence, resulting in a sequence of speech units represented by a speech signal.
[0087] Second, a language model (LM) represents the knowledge of language structure. This structure can include patterns between words and sentences, such as grammar and common word collocations. Language models can include N-gram models and recurrent neural networks (RNNs). For a text sequence, the language model calculates the probability distribution of the sequence, which can be understood as determining whether a language sequence is a normal sentence.
[0088] Third, the pronunciation dictionary is used to record the correspondence between words and factors and is the hub connecting the acoustic model and the language model.
[0089] Fourth, the word error rate (WER) or character error rate (CER) describes the degree of match between the recognized word sequence and the ground-truth word sequence in a speech recognition task and is an evaluation metric for speech recognition systems. To ensure consistency between the recognized word sequence and the ground-truth word sequence, certain words may need to be replaced, deleted, or inserted. The WER or CER is the percentage of the total number of inserted, replaced, or deleted words divided by the total number of words in the standard word sequence. English speech recognition is typically described using WER, while Chinese speech recognition is described using CER.
[0090] Fifth, federated learning (Federated Learning) refers to a machine learning framework that effectively helps multiple nodes (representing individuals or institutions) jointly train machine learning or deep learning models while ensuring data privacy. As a new machine learning concept, federated learning uses distributed training and encryption technology to ensure maximum protection of user privacy data, thereby enhancing user trust in AI technology. Under the federated learning mechanism, each participant contributes encrypted data models to the alliance, jointly trains a federated model, and then makes this model available to all participants.
[0091] Federated learning combines different participants (participants, or parties, also called data owners, or clients) to perform machine learning modeling. During the federated learning process, participants do not need to expose their own data to other participants and coordinators (also called servers, parameter servers, or aggregation servers). Therefore, federated learning can effectively protect user privacy and data security, and can solve the problem of data silos.
[0092] Sixth, evolutionary algorithms, or evolutionary algorithms, are a family of algorithms. Despite their numerous variations, including different genetic expression methods, crossover and mutation operators, the use of specialized operators, and diverse regeneration and selection methods, they all draw inspiration from biological evolution in nature. Compared to traditional optimization algorithms such as calculus-based methods and exhaustive methods, evolutionary computation is a mature, highly robust, and widely applicable global optimization method. Its self-organizing, self-adaptive, and self-learning properties allow it to effectively address complex problems that are difficult for traditional optimization algorithms to solve, regardless of the nature of the problem.
[0093] Different encoding schemes, selection strategies, and search operations produce different evolutionary algorithms. Evolutionary algorithms include genetic algorithms, evolutionary programming, evolutionary programming, and evolutionary strategies. Genetic algorithms include crossover and mutation operations, with a greater emphasis on crossover, while mutation is a secondary operation. Evolutionary programming and evolutionary strategies focus more on mutation, or even exclusively employ mutation.
[0094] Currently, most speech recognition models are limited to specific scenarios. For example, a speech recognition model for an in-vehicle environment can only recognize audio data related to the vehicle, while a smart TV can only recognize audio data related to a smart home environment. With the continuous advancement of artificial intelligence technology, the integration of speech recognition models for different target scenarios is becoming an inevitable trend. Improving the recognition performance of multi-scenario speech recognition models is a pressing issue.
[0095] To address these issues, the present invention proposes the following technical concepts: Based on the concept of federated learning, a training platform for multi-scenario speech recognition models is established. This platform provides an automated modeling process for different institutions or enterprises, enabling modelers to create speech recognition models for specific scenarios without requiring specialized technical skills. The platform also offers model fusion capabilities. To improve the fusion effect, an evolutionary algorithm is employed to iterate the model parameters of the multi-scenario speech recognition models multiple times. By updating the fused model parameters, the recognition performance of the speech recognition models created by each institution or enterprise is improved.
[0096] Before introducing the technical solution of the present invention, the application scenarios of the technical solution of the present invention are briefly introduced.
[0097] Figure 1 Schematic diagram of a scenario of a method for processing a speech recognition model provided by an embodiment of the present invention. Figure 1 As shown, the scenario includes multiple terminal devices, such as Figure 1 Terminal devices 11 and 12, and server 13. Terminal devices 11 and 12 are respectively connected to server 13 for communication. Server 13 provides modeling services for speech recognition models. Terminal device 11 is a terminal device belonging to a modeler of organization A, and terminal device 12 is a terminal device belonging to a modeler of organization B.
[0098] As an example, modelers from organizations A and B access server 13 through their respective terminal devices to create their respective speech recognition models. It should be understood that different organizations use different modeling data to train their speech recognition models. For example, the modeling data for a car company comes from audio data in the vehicle environment, while the modeling data for smart home appliances comes from audio data in the home environment.
[0099] As an example, server 13 stores the model parameters of speech recognition models created by various organizations. Different organizations can obtain the model parameters of speech recognition models created by other organizations from server 13 based on their needs to optimize their local speech recognition models and enhance their recognition performance. It should be noted that each organization's modeling data is derived from user data from the corresponding scenario. To protect user privacy, server 13 should promptly delete the modeling data after completing the training of the corresponding speech recognition model.
[0100] As an example, a robotic process automation (RPA) device (i.e., an RPA platform) is integrated into server 13. The RPA device has the following functions: first, automatic text annotation of audio data, which can be used as a reference text for text annotation; second, automatic modeling of speech recognition models; and third, visual information display of the modeling process.
[0101] Based on the above scenario, the technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0102] Figure 2 Schematic diagram of the process of the speech recognition model processing method provided by the embodiment of the present invention Figure 1 The technical solution provided in this embodiment can be applied to Figure 1 The server shown. Figure 2 As shown, the method includes:
[0103] Step 201: In response to a model selection operation on the graphical user interface of the RPA platform, speech recognition models for multiple target scenarios are obtained from a database of the RPA platform.
[0104] Among them, the speech recognition model of each target scene is obtained by training the basic speech recognition model through transfer learning based on the audio samples of the target scene and the text annotations corresponding to the audio samples.
[0105] In this step, the graphical user interface of the RPA platform can be referred to Figure 8 or, Figure 9 or Figure 10 ,exist Figure 8 or Figure 9 The graphical user interface shown includes a "Select Model" control, and the user can select a model by clicking on the control; or Figure 10The graphical user interface includes a "Source Model Path" input control, where users can select a model by entering at least two model paths. Based on the speech recognition models for the multiple target scenarios selected by the user, the RPA platform retrieves all information about the speech recognition models, such as model parameters, from its database.
[0106] In this embodiment, the speech recognition models of the multiple target scenes include a speech recognition model of the first target scene, and the first target scene is any one of the multiple target scenes. The training process of the speech recognition model of the first target scene is described in detail below.
[0107] Acquire a speech recognition model for each target scenario through transfer learning, including:
[0108] Step 1: Obtain an audio sample of a first target scene and a text annotation corresponding to the audio sample.
[0109] Step 2: Obtain a basic speech recognition model. The basic speech recognition model includes a basic acoustic model and a basic language model. The output of the basic acoustic model corresponds to the input of the basic language model.
[0110] In this embodiment, the input of the basic speech recognition model is used as the input of the basic acoustic model. After data processing of the basic acoustic model and the basic language model, the output of the basic language model is used as the output of the basic speech recognition model.
[0111] It should be noted that the basic speech recognition model refers to a speech recognition model trained based on audio samples of common scenarios (such as life scenarios), and the target scenario refers to a scenario in a specific field or specific working environment.
[0112] Specifically, the input of the speech recognition model is an audio sample, such as a piece of speech data. The audio sample first enters the acoustic model of the speech recognition model, and the following steps are performed within the acoustic model:
[0113] First, the speech data is segmented and features are extracted from the segmented speech segments to obtain feature vectors (also called feature sequences or pronunciation sequences) corresponding to the multiple speech segments. The acoustic model can output multiple sets of possible feature vectors.
[0114] Subsequently, multiple groups of possible feature vectors are input into the language model of the speech recognition model. The language model is used to estimate the probability (or rationality) of each group of possible feature vectors, and determine the group of feature vectors with the highest probability from the multiple groups of possible feature vectors. The text sequence corresponding to the group of feature vectors with the highest probability best conforms to the grammatical rules.
[0115] Finally, the language model obtains the text sequence corresponding to the optimal feature vector based on the pronunciation dictionary, and uses this text sequence as the text sequence corresponding to the audio sample.
[0116] Optionally, the server provides multiple types of acoustic models, including but not limited to hybrid acoustic models such as GMM-HMM-based acoustic models, DNN-HMM-based acoustic models, RNN-HMM-based acoustic models, and CNN-HMM-based acoustic models, as well as end-to-end acoustic models such as the Continuous Time Classification-Long Short-Term Memory (CTC-LSTM) model and the attention model. This embodiment of the present invention does not impose any restrictions on the selection of acoustic models.
[0117] Optionally, the server provides multiple types of language models, including but not limited to statistical language models and neural network language models. Classic statistical language models include N-gram language models. The present embodiment does not impose any restrictions on the selection of language models.
[0118] Step 3: Use the audio sample of the first target scene as the input of the basic acoustic model, and the text annotation corresponding to the audio sample as the output of the basic language model to train the basic speech recognition model.
[0119] In an optional embodiment, the acoustic model and the language model are trained as a whole.
[0120] In an optional embodiment, based on the audio samples in the first target scene and the annotations corresponding to the audio samples (for example, including phoneme annotations, text annotations, etc.), the acoustic model and the language model can be trained separately. When the evaluation parameters of the acoustic model and the language model are within their respective preset evaluation parameter ranges, the acoustic model and the language model are combined to obtain a language recognition model.
[0121] For example, Figure 3 Schematic diagram of a graphical user interface provided by a server according to an embodiment of the present invention Figure 1 , Figure 4 Schematic diagram of a graphical user interface provided by a server according to an embodiment of the present invention Figure 2 . Figure 3 The graphical user interface shown is the training interface for the acoustic model, which includes options such as hot start model path, training corpus path, acoustic model output path, etc. It also includes a visual display of the training progress, such as start, data preparation, feature extraction, data alignment, model training, and completion. Among them, the hot start model refers to the basic acoustic model, that is, the acoustic model trained based on audio samples of common scenes. The training prediction path stores audio samples of the specific scene currently to be trained. After determining the hot start model path, training prediction path, and acoustic model output path, click Start Training, and you can eventually obtain the acoustic model for identifying the specific scene.
[0122] Figure 4 The graphical user interface shown is a training interface for a language model, which includes options such as uploading text, uploading hot words, merging language models, language model output path, model order, etc. It also includes a visual display of the training progress, such as start, language model training, merging language models, hot word optimization, and completion. Among them, the selection model for the merged language model item refers to the selection of a basic language model, that is, a language model trained based on text samples of common scenarios. Uploading text or uploading hot words can be regarded as uploading text samples of the specific scenario to be trained. After determining the training text, basic language model, and language model output path, click Start Training, and you can eventually obtain a language model for recognizing that specific scenario.
[0123] The above graphical user interface is only an example. New functional items may be added or current functional items may be optimized according to actual needs, and this embodiment of the present invention does not impose any limitation on this.
[0124] Based on the above graphical user interface example, the acoustic model and language model of the speech recognition model for a specific scenario (such as the first target scenario) can be trained. The server for this process already integrates the relevant knowledge of model training, allowing operators to complete the training of the speech recognition model for a specific scenario without requiring advanced expertise.
[0125] Step 4: When the evaluation parameters of the speech recognition model in the first target scene are within a preset evaluation parameter range, obtain the speech recognition model of the first target scene.
[0126] In an optional embodiment, the evaluation parameters of the speech recognition model include word error rate. If the word error rate of the output result of the speech recognition model is less than or equal to a preset threshold value compared with the annotation result, the evaluation parameters of the speech recognition model can be considered to be within the preset evaluation parameter range. Otherwise, the model parameters of the speech recognition model need to be further optimized.
[0127] Based on steps 1 to 4 above, the speech recognition model of the first target scenario is obtained, and the model parameters of the speech recognition model of the first target scenario can be stored in the storage space of the RPA platform for subsequent federated learning of multi-scenario speech recognition models.
[0128] It should be understood that the training data for different target scenarios are different, but their training processes are similar.
[0129] Step 202: In response to a trigger operation of a model training control on a graphical user interface, model parameters of speech recognition models for multiple target scenarios are obtained, and model parameters of a federated speech recognition model are obtained based on the model parameters of the speech recognition models for multiple target scenarios.
[0130] The federated speech recognition model supports the recognition of audio data from multiple target scenarios. The model parameters of the federated speech recognition model are determined by performing N iterations on the model parameters of the speech recognition models for the multiple target scenarios. Where N is a positive integer.
[0131] This embodiment does not impose any specific restrictions on the iteration method of the model parameters, as long as the iteration of the model parameters can improve the recognition effect of the federated speech recognition model.
[0132] In this step, the graphical user interface of the RPA platform can be referred to Figure 8 、 Figure 9 or Figure 10 ,exist Figure 8 、 Figure 9 or Figure 10 The graphical user interface includes a "Start Training" control, which is the model training control. After the user selects speech recognition models for multiple target scenarios, by clicking this control, the RPA platform triggers N iterations of the model parameters of the speech recognition models for multiple target scenarios, thereby obtaining the model parameters of the federated speech recognition model.
[0133] Step 203: Store the model parameters of the federated speech recognition model in a preset path.
[0134] Before the RPA platform performs multi-scenario model training, users can preset the storage path of the federated speech recognition model on the RPA platform. After determining the model parameters of the federated speech recognition model, the RPA platform directly stores all information related to the federated speech recognition model in the specified location.
[0135] As an example, enterprise user A logs in to the RPA platform, inputs data samples from domain A (including audio samples and corresponding text annotations), and trains the basic speech recognition model through transfer learning to obtain enterprise A's speech recognition model A. Enterprise user B logs in to the RPA platform, inputs data samples from domain B, and trains the basic speech recognition model through transfer learning to obtain enterprise B's speech recognition model B. Similarly, enterprise user C logs in to the RPA platform. After completing enterprise C's speech recognition model C, they can select the model parameters of the shared enterprise A's speech recognition model A and enterprise B's speech recognition model B on the RPA platform. Then, based on the federated learning or evolutionary learning functional modules provided by the RPA platform, enterprise C's speech recognition model C is further optimized, enabling speech recognition model C to recognize audio data from domains A, B, and C.
[0136] The method for processing a speech recognition model shown in an embodiment of the present invention responds to a model selection operation on a graphical user interface of an RPA platform, obtains speech recognition models for multiple target scenarios from the database of the RPA platform, iterates the model parameters of the speech recognition models of the multiple target scenarios N times to determine the model parameters of the federated speech recognition model, obtains a final federated speech recognition model, and stores its model parameters in a preset path. The speech recognition model of each target scenario is obtained by training a basic speech recognition model through transfer learning based on the audio samples of the target scenario and the text annotations corresponding to the audio samples, where N is a positive integer. The above scheme is executed by the RPA platform. On the one hand, the user only needs to add a certain number of annotated audio samples to obtain a speech recognition model for a specific scenario. On the other hand, the user can obtain a federated speech recognition model by adding speech recognition models that have been trained for multiple specific scenarios that need to be fused. The federated speech recognition model fuses the model parameters of multiple target scenarios, improves the breadth and depth of the speech recognition model, and achieves better recognition effect.
[0137] Figure 5 Schematic diagram of the process of processing the speech recognition model provided by the embodiment of the present invention Figure 2 The technical solution provided in this embodiment can also be applied to Figure 1 The server shown. Figure 5 As shown, the method includes:
[0138] Step 301: Acquire model parameters of speech recognition models for multiple target scenarios to obtain multiple groups of model parameters.
[0139] Step 302: Use an evolutionary algorithm to adjust multiple groups of model parameters.
[0140] In an optional embodiment, at least two groups of model parameters are randomly selected from the multiple groups of model parameters, and the at least two groups of model parameters are cross-exchanged to obtain multiple groups of adjusted model parameters.
[0141] For example, assume there are three groups of model parameters, each of which includes five parameters, namely {a1, b1, c1, d1, e1}, {a2, b2, c2, d2, e2}, and {a3, b3, c3, d3, e3}. Randomly select the first two groups from these three groups of model parameters, for example, and swap the model parameters. For example, swap the first three parameters of the first and second groups to obtain two new groups of model parameters {a2, b2, c2, d1, e1} and {a1, b1, c1, d2, e2}. After the above operation, the three adjusted groups of model parameters are obtained, namely {a2, b2, c2, d1, e1}, {a1, b1, c1, d2, e2}, and {a3, b3, c3, d3, e3}.
[0142] This embodiment only involves the cross-exchange of model parameters of the multi-scenario speech recognition model. The model parameters to be fused are obtained through this operation, which can improve the recognition accuracy of the federated speech recognition model.
[0143] In an optional embodiment, at least two groups of model parameters are randomly selected from multiple groups of model parameters, the at least two groups of model parameters are cross-interchanged, and the at least two groups of model parameters after the cross-interchange are randomly adjusted according to the coefficient of variation to obtain multiple groups of adjusted model parameters.
[0144] Exemplarily, based on the example of the previous embodiment, two sets of model parameters are randomly selected from the three sets of model parameters and cross-interchange is performed to obtain two new sets of model parameters, namely {a2, b2, c2, d1, e1} and {a1, b1, c1, d2, e2}. Based on a preset coefficient of variation, at least one parameter in the two new sets of model parameters is randomly adjusted, for example, by adding a standard deviation t to c2 to generate a number close to c2, such as c2-t. After the above operation, the three adjusted sets of model parameters are obtained, namely {a2, b2, c2-t, d1, e1}, {a1, b1, c1, d2, e2}, and {a3, b3, c3, d3, e3}.
[0145] This embodiment involves cross-exchange and mutation of model parameters of multi-scenario speech recognition models. By obtaining the model parameters to be fused through this operation, the recognition accuracy of the federated speech recognition model can be improved.
[0146] Step 303: Determine the model parameters of the federated speech recognition model according to the adjusted multiple groups of model parameters.
[0147] Specifically, the evaluation parameters of the federated speech recognition model corresponding to the candidate model parameters are obtained. If the evaluation parameters are within the preset evaluation parameter range (for example, the word error rate is less than or equal to the preset threshold), the candidate model parameters are used as the model parameters of the federated speech recognition model; if the evaluation parameters are outside the preset evaluation parameter range (for example, the word error rate is greater than the preset threshold), the step of obtaining the federated speech recognition model based on the speech recognition models of multiple target scenarios is continued, that is, the evolutionary algorithm is continued to be used to adjust multiple sets of model parameters, and the model parameters of the federated speech recognition model are determined based on the adjusted multiple sets of model parameters until the evaluation parameters corresponding to the candidate model parameters are within the preset evaluation parameter range.
[0148] Figure 6 Schematic diagram of a graphical user interface provided by a server according to an embodiment of the present invention Figure 3 . Figure 6The graphical user interface shown is the model performance evaluation interface, which includes options such as the acoustic model path, language model path, test audio, and annotated text of the speech recognition model. The overall performance of the speech recognition model is comprehensively evaluated through test audio and annotated text to obtain evaluation results, such as the word error rate.
[0149] The above graphical user interface is only used as an example. The performance evaluation of the acoustic model or language model of the speech recognition model can be performed separately according to actual needs. This embodiment of the present invention does not impose any limitation on this.
[0150] In an optional embodiment, corresponding parameters in the adjusted multiple sets of model parameters are weighted and summed to obtain candidate model parameters of the federated speech recognition model. The candidate model parameters of the federated speech recognition model can be determined in the following two ways:
[0151] In one possible implementation, the corresponding parameters in the adjusted sets of model parameters are summed and averaged to determine candidate model parameters for the federated speech recognition model. For example, assuming there are three sets of model parameters, each containing two parameters, {a1, b1}, {a2, b2}, and {a3, b3}, the candidate model parameters for the federated speech recognition model can be expressed as {(a1+a2+a3) / 3, (b1+b2+b3) / 3}.
[0152] In one possible implementation, corresponding parameters in the adjusted multiple sets of model parameters are weighted and summed to determine candidate model parameters for the federated speech recognition model. For example, based on the previous example, different sets of model parameters correspond to different application scenarios. For example, the three sets of model parameters correspond to scenarios A, B, and C, respectively. The weights of scenarios A, B, and C are preset to 0.5, 0.4, and 0.1, respectively. Accordingly, the candidate model parameters for the federated speech recognition model can be expressed as {0.5×a1+0.4×a2+0.1×a3, 0.5×b1+0.4×b2+0.1×b3}.
[0153] Optionally, in some embodiments, the server also provides an auxiliary tagging function for audio samples. Before model training, the server first uses automatic speech recognition (ASR) technology to generate a reference text for the audio sample based on the audio sample of the scene input by the user. In response to the user's modification operation on the reference text, the server obtains the text annotation corresponding to the audio sample. In other words, the text annotation is determined by the user based on the reference text.
[0154] Figure 7 Schematic diagram of a graphical user interface provided by a server according to an embodiment of the present invention Figure 4 . Figure 7The graphical user interface shown is an audio data annotation interface, which includes a visual feature map of the audio data and a reference text of the audio data automatically generated by the server. The user can directly modify the reference text and click Submit after completing the modification.
[0155] Figure 8 Schematic diagram of a graphical user interface provided by a server according to an embodiment of the present invention Figure 5 . Figure 8 The graphical user interface shown is the federated learning interface of the acoustic model. By inputting the acoustic models of multiple scenes that have been trained, the acoustic models of multiple scenes are fused to obtain a federated acoustic model. Among them, the acoustic model of each scene that has been trained can be obtained through transfer learning of the basic acoustic model. For details, please refer to Figure 3 .
[0156] Figure 9 Schematic diagram of a graphical user interface provided by a server according to an embodiment of the present invention Figure 6 . Figure 9 The graphical user interface shown is the federated learning interface of the language model. By inputting the language models of multiple scenarios that have been trained, the language models of multiple scenarios are integrated to obtain the federated language model. Among them, the language model of each scenario that has been trained can be obtained through the transfer learning of the basic language model. For details, please refer to Figure 4 .
[0157] Figure 10 Schematic diagram of a graphical user interface provided by a server according to an embodiment of the present invention Figure 7 . Figure 10 The graphical user interface shown is the evolutionary learning interface of the acoustic model. By inputting at least two acoustic models and the training prediction of the current specific scene, the acoustic models of multiple scenes are fused to obtain the evolved acoustic model. Among them, each acoustic model of at least two acoustic models can be obtained through transfer learning of the basic acoustic model. For details, please refer to Figure 3 .
[0158] It should be noted that the algorithms used in federated learning and evolutionary learning in this embodiment are similar, such as the genetic algorithm or evolutionary algorithm mentioned above, which are essentially for optimizing multi-scenario model parameters.
[0159] Based on the above embodiments, it can be seen that in order to reduce the labor cost of model training, the present invention proposes an RPA device with an automated overall process based on speech recognition, so that users of the RPA device can complete the optimization of the speech recognition model in a specific scenario without having any model experience.
[0160] The RPA device mainly consists of four parts: voice data labeling part, model optimization (training) part, model evaluation part and federated learning part.
[0161] The speech data annotation process uses ASR technology to pre-provide audio reference text to annotators. This significantly improves annotation efficiency and reduces costs compared to starting from scratch. After data annotation is complete, structured data is automatically generated, enabling rapid subsequent model optimization and evaluation.
[0162] For model optimization, transfer learning methods can be used, including optimization of acoustic models and language models. Users only need to configure the scene data and click training to complete the model optimization. The training process can be visualized.
[0163] The model evaluation part supports flexible model configuration, free combination of acoustic models and language models, and efficiently supports model effect evaluation and model iterative optimization.
[0164] In the federated learning section, users can upload model parameters and use federated learning to train data from different target scenarios to optimize the performance of multi-scenario speech recognition models. Evolutionary or genetic algorithms can be used to merge multi-scenario models to achieve even better speech recognition results.
[0165] The embodiment of the present invention can divide the processing device of the speech recognition model into functional modules according to the above-mentioned method embodiment. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present invention is schematic and is only a logical functional division. There may be other division methods in actual implementation. The following is an example of dividing each functional module corresponding to each function.
[0166] Figure 11 The structure block diagram of the processing device of the speech recognition model provided by the embodiment of the present invention. For the convenience of explanation, only the part related to the embodiment of the present invention is shown. Figure 11 As shown, the speech recognition model processing device provided in this embodiment includes: an acquisition module 401 and a processing module 402.
[0167] An acquisition module 401 is configured to, in response to a model selection operation on a graphical user interface of the RPA platform, acquire speech recognition models for multiple target scenarios from a database of the RPA platform, where the speech recognition model for each target scenario is obtained by training a basic speech recognition model through transfer learning based on audio samples of the target scenario and text annotations corresponding to the audio samples;
[0168] A processing module 402 is configured to, in response to a triggering operation of a model training control on the graphical user interface, obtain model parameters of the speech recognition models for the multiple target scenarios, and obtain model parameters of a federated speech recognition model based on the model parameters of the speech recognition models for the multiple target scenarios; the federated speech recognition model is configured to recognize audio data for the multiple target scenarios, and the model parameters of the federated speech recognition model are determined by performing N iterations on the model parameters of the speech recognition models for the multiple target scenarios, where N is a positive integer;
[0169] The storage module is used to store the model parameters of the federated speech recognition model in a preset path.
[0170] In an optional embodiment of this embodiment, the acquisition module 401 is configured to acquire model parameters of the speech recognition models of the multiple target scenes to obtain multiple groups of model parameters;
[0171] A processing module 402 is configured to adjust the plurality of model parameters using an evolutionary algorithm;
[0172] The model parameters of the federated speech recognition model are determined according to the adjusted multiple groups of model parameters.
[0173] In an optional embodiment of this embodiment, the processing module 402 is configured to:
[0174] randomly selecting at least two groups of model parameters from the multiple groups of model parameters;
[0175] The at least two groups of model parameters are cross-exchanged to obtain multiple groups of adjusted model parameters.
[0176] In an optional embodiment of this embodiment, the processing module 402 is configured to:
[0177] randomly selecting at least two groups of model parameters from the multiple groups of model parameters;
[0178] The at least two groups of model parameters are cross-exchanged, and the at least two groups of model parameters after the cross-exchange are randomly adjusted according to the coefficient of variation to obtain multiple groups of adjusted model parameters.
[0179] In an optional embodiment of this embodiment, the processing module 402 is configured to:
[0180] Determining candidate model parameters of the federated speech recognition model based on the adjusted multiple sets of model parameters;
[0181] Obtaining evaluation parameters of the federated speech recognition model corresponding to the candidate model parameters;
[0182] If the evaluation parameter is within a preset evaluation parameter range, the candidate model parameter is used as the model parameter of the federated speech recognition model.
[0183] In an optional embodiment of this embodiment, the processing module 402 is configured to:
[0184] The corresponding parameters in the adjusted multiple sets of model parameters are summed and averaged, or the corresponding parameters in the adjusted multiple sets of model parameters are weighted and summed according to preset weight values of different target scene model parameters to obtain candidate model parameters of the federated speech recognition model.
[0185] In an optional embodiment of this embodiment, the processing module 402 is configured to:
[0186] If the evaluation parameter is outside the preset evaluation parameter range, the step of adjusting the multiple sets of model parameters using the evolutionary algorithm is performed again until the evaluation parameter corresponding to the candidate model parameter is within the preset evaluation parameter range.
[0187] In an optional embodiment of this embodiment, the speech recognition models of the multiple target scenes include a speech recognition model of a first target scene, where the first target scene is any one of the multiple target scenes; the acquisition module 401 is configured to:
[0188] Obtaining an audio sample of the first target scene and a text annotation corresponding to the audio sample;
[0189] Acquire the basic speech recognition model, where the basic speech recognition model includes a basic acoustic model and a basic language model, and the output of the basic acoustic model corresponds to the input of the basic language model;
[0190] A processing module 402 is configured to train the basic speech recognition model by using the audio sample of the first target scene as the input of the basic acoustic model and the text annotation corresponding to the audio sample as the output of the basic language model;
[0191] When the evaluation parameters of the speech recognition model of the first target scene are within a preset evaluation parameter range, the speech recognition model of the first target scene is obtained.
[0192] In an optional embodiment of this embodiment, the acquisition module 401 is used to obtain an audio sample of the first target scene input by a user, and the processing module 402 is used to generate a reference text of the audio sample using automatic speech recognition ASR technology;
[0193] The acquisition module 401 is further configured to acquire the text annotation corresponding to the audio sample in response to a user's modification operation on the reference text.
[0194] The processing device of the speech recognition model provided in the embodiment of the present invention is used to execute the technical solution provided by any of the aforementioned method embodiments. Its implementation principles and technical effects are similar and will not be repeated here.
[0195] Figure 12 This is a structural block diagram of an electronic device provided by an embodiment of the present invention. Figure 12 As shown, the electronic device 500 of this embodiment may include:
[0196] At least one processor 501 ( Figure 12 Only one processor is shown); and
[0197] A memory 502 in communication with the at least one processor; wherein,
[0198] The memory 502 stores a computer program that can be executed by the at least one processor 501. The computer program is executed by the at least one processor 501 to enable the electronic device 500 to execute the technical solution of the first device in any of the aforementioned method embodiments.
[0199] Optionally, the memory 502 may be independent or integrated with the processor 501 .
[0200] When the memory 502 is a device independent of the processor 501 , the electronic device 500 further includes a bus 503 for connecting the memory 502 and the processor 501 .
[0201] The electronic device provided by the embodiment of the present invention can execute the technical solution provided by any of the aforementioned method embodiments, and its implementation principles and technical effects are similar, which will not be repeated here.
[0202] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, it is used to implement the technical solution provided by any of the aforementioned method embodiments.
[0203] An embodiment of the present invention further provides a computer program product, including a computer program, which implements the technical solution provided by any of the aforementioned method embodiments when executed by a processor.
[0204] An embodiment of the present invention further provides a chip, comprising: a processing module and a communication interface, wherein the processing module can execute the technical solution provided by any of the aforementioned method embodiments.
[0205] Furthermore, the chip also includes a storage module (such as a memory), the storage module is used to store instructions, the processing module is used to execute the instructions stored in the storage module, and the execution of the instructions stored in the storage module enables the processing module to execute the technical solution provided by any of the aforementioned method embodiments.
[0206] It should be understood that the processor described above may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), or application-specific integrated circuits (ASICs). A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the present invention may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.
[0207] The memory may include a high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk or an optical disk.
[0208] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, the buses shown in the drawings of the present invention are not limited to just one bus or just one type of bus.
[0209] The storage medium may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0210] An exemplary storage medium is coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an application-specific integrated circuit (ASIC). Of course, the processor and storage medium can also exist as discrete components in an electronic device.
[0211] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for processing a speech recognition model, characterized in that: The method is applied to a Robotic Process Automation (RPA) platform and includes: In response to a model selection operation on the graphical user interface of the RPA platform, multiple speech recognition models for target scenarios are obtained from a database of the RPA platform. The speech recognition model for each target scenario is obtained by training a basic speech recognition model through transfer learning based on audio samples of the target scenario and text annotations corresponding to the audio samples. The RPA platform has automatic text annotation capabilities and automatic modeling capabilities for speech recognition models. In response to a triggering operation of a model training control on the graphical user interface, model parameters of the speech recognition models of the multiple target scenarios are obtained to obtain multiple sets of model parameters; Using an evolutionary algorithm to adjust the multiple sets of model parameters; The corresponding parameters in the adjusted multiple sets of model parameters are summed and averaged, or the corresponding parameters in the adjusted multiple sets of model parameters are weighted and summed according to preset weight values of different target scenario model parameters to obtain candidate model parameters of the federated speech recognition model; Obtaining evaluation parameters of the federated speech recognition model corresponding to the candidate model parameters; If the evaluation parameter is within a preset evaluation parameter range, using the candidate model parameter as the model parameter of the federated speech recognition model; If the evaluation parameter is outside the preset evaluation parameter range, performing the step of adjusting the multiple sets of model parameters using the evolutionary algorithm again until the evaluation parameter corresponding to the candidate model parameter is within the preset evaluation parameter range; the federated speech recognition model is used to recognize the audio data of the multiple target scenes, and the model parameters of the federated speech recognition model are determined by performing N iterations on the model parameters of the speech recognition models of the multiple target scenes, where N is a positive integer; The model parameters of the federated speech recognition model are stored in a preset path.
2. The method according to claim 1, characterized in that The adopting of an evolutionary algorithm to adjust the multiple groups of model parameters comprises: randomly selecting at least two groups of model parameters from the multiple groups of model parameters; The at least two groups of model parameters are cross-exchanged to obtain multiple groups of adjusted model parameters.
3. The method according to claim 1, characterized in that The adopting of an evolutionary algorithm to adjust the multiple groups of model parameters comprises: randomly selecting at least two groups of model parameters from the multiple groups of model parameters; The at least two groups of model parameters are cross-exchanged, and the at least two groups of model parameters after the cross-exchange are randomly adjusted according to the coefficient of variation to obtain multiple groups of adjusted model parameters.
4. The method according to any one of claims 1 to 3, characterized in that The speech recognition models of the multiple target scenes include a speech recognition model of a first target scene, where the first target scene is any one of the multiple target scenes; The training process of the speech recognition model for the first target scenario includes: Obtaining an audio sample of the first target scene and a text annotation corresponding to the audio sample; Acquire the basic speech recognition model, where the basic speech recognition model includes a basic acoustic model and a basic language model, and the output of the basic acoustic model corresponds to the input of the basic language model; Using the audio sample of the first target scene as the input of the basic acoustic model and the text annotation corresponding to the audio sample as the output of the basic language model, to train the basic speech recognition model; When the evaluation parameters of the speech recognition model of the first target scene are within a preset evaluation parameter range, the speech recognition model of the first target scene is obtained.
5. The method according to claim 4, characterized in that Obtaining the text annotation corresponding to the audio sample includes: Acquire an audio sample of the first target scene input by a user, and generate a reference text of the audio sample using automatic speech recognition (ASR) technology; In response to a user's modification operation on the reference text, a text annotation corresponding to the audio sample is obtained.
6. A training device for a speech recognition model, characterized in that: include: An acquisition module, configured to respond to a model selection operation on a graphical user interface of the RPA platform and obtain speech recognition models for multiple target scenarios from a database of the RPA platform, wherein the speech recognition model for each target scenario is obtained by training a basic speech recognition model through transfer learning based on audio samples of the target scenario and text annotations corresponding to the audio samples, wherein the RPA platform has automatic text annotation capabilities and automatic modeling capabilities for speech recognition models; A processing module, configured to respond to a triggering operation of a model training control acting on the graphical user interface, obtain model parameters of the speech recognition models of the multiple target scenarios, and obtain multiple sets of model parameters; adjust the multiple sets of model parameters using an evolutionary algorithm; sum and average the corresponding parameters in the adjusted multiple sets of model parameters, or perform weighted summation on the corresponding parameters in the adjusted multiple sets of model parameters according to preset weight values of the model parameters of different target scenarios to obtain candidate model parameters of the federated speech recognition model; obtain evaluation parameters of the federated speech recognition model corresponding to the candidate model parameters; if the evaluation parameters are within a preset evaluation parameter range, use the candidate model parameters as the model parameters of the federated speech recognition model; if the evaluation parameters are outside the preset evaluation parameter range, perform the step of adjusting the multiple sets of model parameters using an evolutionary algorithm again until the evaluation parameters corresponding to the candidate model parameters are within the preset evaluation parameter range; the federated speech recognition model is configured to recognize audio data of the multiple target scenarios, and the model parameters of the federated speech recognition model are determined by performing N iterative processes on the model parameters of the speech recognition models of the multiple target scenarios, where N is a positive integer; The storage module is used to store the model parameters of the federated speech recognition model in a preset path.
7. An electronic device, characterized in that: The electronic device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the method according to any one of claims 1 to 5 when executed by the processor.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 5.
9. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 5 when executed by a processor.
Citation Information
Patent Citations
Recognition model training method, device and equipment and readable storage medium
CN111428881A
Model training method and device based on federal learning and electronic equipment
CN113052323A
Data recognition method and device, electronic equipment and computer readable storage medium
CN113066486A