Speech recognition system, related methods, devices and equipment
By dynamically adjusting the parameters of the speech recognition model, the problem of maintaining multiple models in the existing technology is solved, and the effects of resource saving, cost reduction and scalability improvement are achieved.
Patent Information
- Application Number
- CN202010701047.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-15
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2040-07-15
AI Technical Summary
Existing voice recognition systems need to maintain multiple models to meet different requirements for computing volume and delay in different applications, resulting in high resource consumption, high maintenance costs and low scalability.
It provides a speech recognition system, which collects voice data through the client and sends it to the server. The server learns a dynamic and variable speech recognition model from the training sample set, and dynamically determines model parameters, including model size and delay value, according to the needs of the target application, to achieve speech recognition that adapts to different application scenarios.
It realizes that different needs of different applications for computing volume and delay are met through a general speech recognition model, saves equipment resources, reduces model maintenance costs, improves the scalability of application scenarios and the model deployment efficiency of new application scenarios.
Smart Images

Figure CN114023309B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and specifically to a speech recognition system, method and device, a speech recognition service upgrade method and device, a speech recognition service test method and device, a smart speaker, a smart television, a food ordering device, a smart mobile device, a vehicle-mounted voice assistant device, a trial device, and an electronic device. Background Art
[0002] Different application scenarios of speech recognition systems have different requirements for computing power and latency. For example, in smart speaker scenarios, speech recognition systems are usually deployed in the cloud. Since cloud devices have better performance, in order to improve speech recognition performance, a speech recognition model with more computing units and higher computing latency can be used; in scenarios such as ordering machines, smart TVs, and court trials, speech recognition systems are usually deployed on the end. Due to the limited performance of end devices and in order to meet the real-time requirements of user interaction, a speech recognition model with fewer computing units and higher latency requirements is usually required; in instant messaging scenarios (such as DingTalk), the latency requirements for speech recognition systems are not high, and a model with a larger amount of computing can be used.
[0003] At present, the speech recognition system mainly meets the different requirements of different applications for computing power and latency by maintaining multiple speech recognition models at the same time. That is, each model has a fixed model size and latency. Different applications use different speech recognition models according to their different requirements for computing power and latency. Different speech recognition models need to be trained and maintained separately.
[0004] However, in the process of implementing the present invention, the inventors found that the technical solution has at least the following problems: 1) Since multiple models need to be maintained at the same time to meet the different requirements of different applications for computing power and latency, more computing resources and storage resources will be consumed, and the model training and maintenance costs are high; 2) When facing the speech recognition needs of new application scenarios, it is necessary to retrain a speech recognition model suitable for the scenario's requirements for computing power and latency, so the scalability of the speech recognition system is low. In summary, how to provide a unified speech recognition model with controllable model parameters to meet the different requirements of different application scenarios for computing power and latency, so as to save equipment resources, improve the scalability of application scenarios, and reduce model maintenance costs has become a problem that technicians in this field urgently need to solve. Summary of the invention
[0005] The present application provides a speech recognition system to solve the problem that the existing technology cannot meet the different requirements of different applications for computing power and latency through a universal speech recognition model. The present application also provides a speech recognition method and device, a speech recognition service upgrade method and device, a speech recognition service test method and device, a smart speaker, a smart TV, a food ordering device, a smart mobile device, a car voice assistant device, a trial device, and an electronic device.
[0006] The present application provides a speech recognition system, comprising:
[0007] The client is used to collect the voice data of the target application and send the voice data to the server;
[0008] The server is used to learn a speech recognition model with dynamically changeable model parameters from a training sample set; determine the target model parameters corresponding to the target application for the speech data sent by the terminal device; and convert the speech data into a text sequence through the speech recognition model based on the target model parameters.
[0009] The present application also provides a speech recognition method, comprising:
[0010] A speech recognition model with dynamically variable model parameters is learned from a training sample set;
[0011] Determining target model parameters corresponding to the target application;
[0012] The speech data of the target application is converted into a text sequence by using the speech recognition model based on the target model parameters.
[0013] Optionally, the model parameters include: model size;
[0014] The model size includes: the number of layers and / or the number of neurons of the neural network;
[0015] The speech recognition model with dynamically variable model parameters obtained by learning from the training sample set includes:
[0016] Iterative training is performed on the model according to the dynamically determined model size.
[0017] Optionally, the dynamically determined model size is determined in the following manner:
[0018] Choose any model size from multiple preset model sizes.
[0019] Optionally, the model includes: a streaming end-to-end speech recognition model;
[0020] The model includes: an audio encoder and a decoder;
[0021] The model size includes: the size of the audio encoder.
[0022] Optionally, the model parameters include: a delay value;
[0023] The speech recognition model with dynamically variable model parameters obtained by learning from the training sample set includes:
[0024] The model is iteratively trained according to the dynamically determined delay value.
[0025] Optionally, the dynamically determined delay value is determined in the following manner:
[0026] Select any delay value from multiple preset delay values;
[0027] The delay value of the target application includes: a delay value other than the preset delay value.
[0028] Optionally, the model includes: a streaming end-to-end speech recognition model;
[0029] The model includes: an audio encoder, a feature data determination module, and a decoder;
[0030] The step of converting the speech data into a text sequence by using the speech recognition model based on the target model parameters comprises:
[0031] Determine audio feature data of the speech data through an audio encoder, and store the audio feature data in a block memory according to a delay value of a target application;
[0032] Determine, by a feature data determination module, feature data corresponding to the words in the speech data according to the audio feature data in the block memory;
[0033] The decoder determines the words in the speech data based on the feature data of the words to form the text sequence.
[0034] Optionally, the determining, by the feature data determination module, of the audio feature data corresponding to the word in the speech data according to the audio feature data in the block memory includes:
[0035] Determine the correspondence between words and block memory;
[0036] According to the corresponding relationship, feature data corresponding to the word is determined.
[0037] Optionally, the feature data determination module includes: a predictor;
[0038] The feature data determining module determines the feature data corresponding to the word in the voice data according to the audio feature data in the block memory, and further includes:
[0039] Determining the length of the text included in each block by the predictor;
[0040] The correspondence between characters and blocks is determined according to the text length.
[0041] Optionally, determining the target model parameters corresponding to the target application includes:
[0042] Determine the speech recognition performance requirements of the target application;
[0043] The target model parameters are determined according to the performance requirement information.
[0044] Optionally, if a first user associated with the target application sends a resource object corresponding to the target model parameters to a second user associated with the model, the speech data is converted into a text sequence through the speech recognition model based on the target model parameters.
[0045] The present application also provides a speech recognition method, comprising:
[0046] Collect voice data of the target application and send the voice data to the server so that the server can learn a voice recognition model with dynamically changeable model parameters from a training sample set; determine the target model parameters corresponding to the target application for the voice data; and convert the voice data into a text sequence through the voice recognition model based on the target model parameters.
[0047] The present application also provides a speech recognition method, comprising:
[0048] A speech recognition model with dynamically variable model parameters is learned from a training sample set;
[0049] Determining target model parameters corresponding to the target application;
[0050] The speech recognition model based on the target model parameters is sent to a target device running a target application, so that the target application converts the speech data into a text sequence through the speech recognition model based on the target model parameters.
[0051] Optionally, determining the target model parameters corresponding to the target application includes:
[0052] Determine the speech recognition performance requirements of the target application;
[0053] The target model parameters are determined according to the performance requirement information.
[0054] Optionally, also include:
[0055] Determining device performance requirement information of the target device according to the performance requirement information;
[0056] Sending the device performance requirement information to a management device associated with the target application so that the management device displays the device performance requirement information;
[0057] The speech recognition model based on the target model parameters is sent to a target device that meets the device performance requirement information.
[0058] Optionally, determining the target model parameters corresponding to the target application includes:
[0059] Determine the performance information of the device running the target application;
[0060] The target model parameters are determined according to the device performance information.
[0061] Optionally, the device performance information includes: computing resource information and storage resource information;
[0062] Determining the target model parameters according to the device performance information includes:
[0063] Determining a model size according to the computing resource information;
[0064] A delay value is determined according to the storage resource information.
[0065] Optionally, also include:
[0066] Determining resource information corresponding to the target model parameters;
[0067] Sending the resource information to a first user associated with the target application;
[0068] If the first user sends the resource object to the second user associated with the model, the speech recognition model based on the target model parameters is sent to the target device.
[0069] Optionally, also include:
[0070] Determining speech recognition performance information according to the target model parameters;
[0071] The performance information is sent to a management device associated with the target application, so that the management device displays the performance information.
[0072] Optionally, the target model parameters include: a delay value;
[0073] Correspondingly, the performance information includes: real-time degree of speech recognition.
[0074] Optionally, the target model parameters include: model size;
[0075] Correspondingly, the performance information includes: speech recognition accuracy.
[0076] The present application also provides a speech recognition method, comprising:
[0077] Send a request to the server to obtain the speech recognition model for the target application;
[0078] A speech recognition model with dynamically variable model parameters based on target model parameters corresponding to a target application sent back by a receiving server;
[0079] The speech data is converted into a text sequence by the speech recognition model based on the target model parameters.
[0080] Optionally, also include:
[0081] Determine the speech recognition performance requirements of the target application;
[0082] The request includes the performance requirement information, so that the server determines the target model parameters according to the performance requirement information.
[0083] Optionally, also include:
[0084] Receiving device performance requirement information for running the target application determined according to the performance requirement information, sent by the server;
[0085] The device performance requirement information is displayed to facilitate determining a target device that meets the device performance requirement information, so that the server sends the speech recognition model based on the target model parameters to the target device.
[0086] Optionally, also include:
[0087] Determine the performance information of the device running the target application;
[0088] The request includes the device performance information, so that the server can determine the target model parameters according to the device performance information.
[0089] Optionally, also include:
[0090] Receiving resource information corresponding to the target model parameters sent by the server;
[0091] The resource object is sent to a second user associated with the model so that the server sends the speech recognition model based on the target model parameters.
[0092] Optionally, also include:
[0093] Receiving speech recognition performance information corresponding to the target model parameters sent by the server;
[0094] The speech recognition performance information is displayed.
[0095] Optionally, also include:
[0096] A test system for receiving a speech recognition model based on multiple sets of model parameters sent by a server;
[0097] The speech data is converted into a text sequence by using a speech recognition model based on each set of model parameters, so as to determine the speech recognition performance of each set of model parameters;
[0098] Determine the target model parameters and send them to the server.
[0099] This application also provides a method for upgrading a speech recognition service, including:
[0100] Determining usage status information of a target application for a speech recognition model whose model parameters are dynamically variable based on the first model parameter;
[0101] Determining a second model parameter of the speech recognition model according to the usage status information;
[0102] The model parameters of the speech recognition model on the device running the target application are configured as second model parameters, so that the device converts the speech data into a text sequence through the speech recognition model based on the second model parameters.
[0103] Optionally, determining a second model parameter of the speech recognition model according to the usage status information includes:
[0104] Determining multiple groups of model parameters of the speech recognition model according to the usage status information;
[0105] A test system of a speech recognition model based on multiple groups of model parameters is sent to the device, so that the target application converts speech data into a text sequence through a speech recognition model based on each group of model parameters, so as to determine the speech recognition performance of each group of model parameters, and determine the second model parameters according to the speech recognition performance.
[0106] Optionally, determining a second model parameter of the speech recognition model according to the usage status information includes:
[0107] Determining speech recognition performance requirement information of the target application according to the usage status information;
[0108] Determine a second model parameter of the speech recognition model according to the performance requirement information.
[0109] This application also provides a method for upgrading a speech recognition service, including:
[0110] Determining usage status information of a target application for a speech recognition model whose model parameters are dynamically variable based on the first model parameter;
[0111] Determining a second model parameter of the speech recognition model according to the usage status information;
[0112] The correspondence between the target application and the second model parameters is stored so that for the speech data to be processed by the target application, the speech data is converted into a text sequence according to the correspondence through the speech recognition model based on the second model parameters.
[0113] The present application also provides a speech recognition service testing method, comprising:
[0114] receiving a speech recognition service test request for a target application;
[0115] For the plurality of sets of model parameters, voice data of a target application is converted into a text sequence by a speech recognition model with dynamically variable model parameters based on each set of model parameters;
[0116] The text sequences corresponding to each set of model parameters are sent back to the requesting party so that the requesting party can determine the speech recognition performance of each set of model parameters and determine the target model parameters corresponding to the target application based on the performance.
[0117] The present application also provides a method for constructing a speech recognition model, comprising:
[0118] Determine a training data set, wherein the training data includes: speech data and text sequence annotation information;
[0119] Constructing a network structure of the model;
[0120] According to the dynamically determined model parameters, iterative training is performed on the model to obtain a speech recognition model with dynamically variable model parameters.
[0121] Optionally, the model includes: a streaming end-to-end speech recognition model;
[0122] The model includes: an audio encoder and a decoder;
[0123] The model parameters include a model size, which includes a size of an audio encoder.
[0124] Optionally, the model includes: a streaming end-to-end speech recognition model;
[0125] The model includes: an audio encoder, a feature data determination module, and a decoder;
[0126] The model parameters include: delay value;
[0127] The audio encoder is used to determine audio feature data of the speech data and store the audio feature data in a block memory according to a delay value of a target application;
[0128] The feature data determination module is used to determine the feature data corresponding to the words in the voice data according to the audio feature data in the block memory;
[0129] The decoder is used to determine the characters in the speech data according to the feature data of the characters to form the text sequence.
[0130] Optionally, the feature data determination module is specifically used to determine the correspondence between characters and blocks; and determine the feature data corresponding to the characters based on the correspondence.
[0131] Optionally, the training data further includes: text length annotation information of each block;
[0132] The data determination module includes: a predictor;
[0133] The predictor is used to determine the length of the text included in each block;
[0134] The feature data determination module is used to determine the correspondence between characters and blocks according to the text length.
[0135] The present application also provides an electronic device, comprising:
[0136] Processor; and
[0137] The memory is used to store a program for implementing a speech recognition method. After the device is powered on and runs the program of the method through the processor, the following steps are performed: a speech recognition model with dynamically variable model parameters is obtained by learning from a training sample set; target model parameters corresponding to a target application are determined; and speech data of the target application is converted into a text sequence through the speech recognition model based on the target model parameters.
[0138] The present application also provides a smart speaker, comprising:
[0139] Processor; and
[0140] The memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: voice data of a target application is collected and sent to a server so that the server can learn a voice recognition model with dynamically variable model parameters from a training sample set; target model parameters corresponding to the target application are determined; and the voice data is converted into a text sequence through the voice recognition model based on the target model parameters.
[0141] The present application also provides a food ordering device, comprising:
[0142] Processor; and
[0143] A memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: voice ordering data is collected, and the voice ordering data is converted into ordering text through a voice recognition model with dynamically variable model parameters based on target model parameters corresponding to the ordering application; and ordering processing is performed according to the ordering text.
[0144] The present application also provides a smart TV, comprising:
[0145] Processor; and
[0146] The memory is used to store a program for implementing the voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: collecting television control voice data, converting the voice data into television control text through a voice recognition model with dynamically variable model parameters based on target model parameters corresponding to the television application; and performing television control processing according to the television control text.
[0147] The present application also provides a smart mobile device, comprising:
[0148] Processor; and
[0149] The memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: voice data is collected, and the voice data is converted into a text sequence through a voice recognition model with dynamically variable model parameters based on target model parameters corresponding to the device; and voice interaction processing is performed according to the text sequence.
[0150] The present application also provides a vehicle-mounted voice assistant device, including:
[0151] Processor; and
[0152] The memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: voice data is collected, and the voice data is converted into a text sequence through a voice recognition model with dynamically variable model parameters based on target model parameters corresponding to the device; and voice interaction processing is performed according to the text sequence.
[0153] The present application also provides a trial device, including:
[0154] Processor; and
[0155] The memory is used to store a program for implementing a speech recognition method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: speech data is collected, and the speech data is converted into a text sequence through a speech recognition model with dynamically variable model parameters based on target model parameters corresponding to the device.
[0156] The present application also provides a speech recognition device, comprising:
[0157] A model building unit, used for learning from a training sample set to obtain a speech recognition model with dynamically variable model parameters;
[0158] A model parameter determination unit, used to determine target model parameters corresponding to a target application;
[0159] A model prediction unit is used to convert the speech data of the target application into a text sequence through the speech recognition model based on the target model parameters.
[0160] The present application also provides an electronic device, comprising:
[0161] Processor; and
[0162] The memory is used to store a program for implementing a speech recognition method. After the device is powered on and runs the program of the method through the processor, the following steps are performed: a speech recognition model with dynamically variable model parameters is obtained by learning from a training sample set; target model parameters corresponding to a target application are determined; and speech data of the target application is converted into a text sequence through the speech recognition model based on the target model parameters.
[0163] The present application also provides a speech recognition device, comprising:
[0164] A model building unit, used for learning from a training sample set to obtain a speech recognition model with dynamically variable model parameters;
[0165] A model parameter determination unit, used to determine target model parameters corresponding to a target application;
[0166] The model sending unit is used to send the speech recognition model based on the target model parameters to a target device running a target application, so that the target application converts speech data into a text sequence through the speech recognition model based on the target model parameters.
[0167] The present application also provides an electronic device, comprising:
[0168] Processor; and
[0169] The memory is used to store a program for implementing a speech recognition method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: a speech recognition model with dynamically variable model parameters is learned from a training sample set; target model parameters corresponding to a target application are determined; and the speech recognition model based on the target model parameters is sent to a target device running a target application, so that the target application converts speech data into a text sequence through the speech recognition model based on the target model parameters.
[0170] The present application also provides a speech recognition device, comprising:
[0171] A request sending unit, used for sending a speech recognition model acquisition request for a target application to a server;
[0172] A model receiving unit, used for receiving a speech recognition model with dynamically variable model parameters based on target model parameters corresponding to a target application, which is sent back by a server;
[0173] A speech recognition unit is used to convert speech data into a text sequence through the speech recognition model based on the target model parameters.
[0174] The present application also provides an electronic device, comprising:
[0175] Processor; and
[0176] The memory is used to store a program for implementing a speech recognition method. After the device is powered on and runs the program of the method through the processor, the following steps are performed: sending a speech recognition model acquisition request for a target application to a server; receiving a speech recognition model with dynamically variable model parameters based on target model parameters corresponding to the target application sent back by the server; and converting speech data into a text sequence through the speech recognition model based on the target model parameters.
[0177] The present application also provides a speech recognition service upgrade device, comprising:
[0178] an application usage status determination unit, configured to determine usage status information of a target application on a speech recognition model having dynamically variable model parameters based on a first model parameter;
[0179] a model parameter determination unit, configured to determine a second model parameter of the speech recognition model according to the usage status information;
[0180] The model parameter updating unit is used to configure the model parameters of the speech recognition model on the device running the target application as second model parameters, so that the device converts the speech data into a text sequence through the speech recognition model based on the second model parameters.
[0181] The present application also provides an electronic device, comprising:
[0182] Processor; and
[0183] The memory is used to store a program for implementing a method for upgrading a speech recognition service. After the device is powered on and runs the program of the method through the processor, the following steps are performed: determining usage status information of a speech recognition model whose model parameters are dynamically variable based on first model parameters by a target application; determining second model parameters of the speech recognition model based on the usage status information; and configuring the model parameters of the speech recognition model on a device running the target application as second model parameters, so that the device converts speech data into a text sequence through the speech recognition model based on the second model parameters.
[0184] The present application also provides a speech recognition service upgrade device, comprising:
[0185] The application usage status determination unit is used to determine the usage status information of the target application on the speech recognition model with dynamically variable model parameters based on the first model parameter.
[0186] A model parameter determination unit is used to determine a second model parameter of the speech recognition model according to the usage status information.
[0187] The model parameter updating unit is used to store the correspondence between the target application and the second model parameter, so that the speech data to be processed for the target application is converted into a text sequence according to the correspondence through the speech recognition model based on the second model parameter.
[0188] The present application also provides an electronic device, comprising:
[0189] Processor; and
[0190] A memory is used to store a program for implementing a method for upgrading a speech recognition service. After the device is powered on and the program of the method is run through the processor, the following steps are performed: determining usage status information of a speech recognition model whose model parameters are dynamically variable based on a first model parameter by a target application; determining a second model parameter of the speech recognition model according to the usage status information; and storing a correspondence between the target application and the second model parameter so that for speech data to be processed by the target application, the speech data is converted into a text sequence according to the correspondence through the speech recognition model based on the second model parameter.
[0191] The present application also provides a speech recognition service testing device, comprising:
[0192] A test request receiving unit, used to receive a speech recognition service test request for a target application;
[0193] A speech recognition test unit, for converting speech data of a target application into a text sequence through a speech recognition model with dynamically variable model parameters based on each set of model parameters, for multiple sets of model parameters;
[0194] The text sequence feedback unit is used to return the text sequence corresponding to each group of model parameters to the requester, so that the requester can determine the speech recognition performance of each group of model parameters and determine the target model parameters corresponding to the target application based on the performance.
[0195] The present application also provides an electronic device, comprising:
[0196] Processor; and
[0197] The memory is used to store a program for implementing a speech recognition service test method. After the device is powered on and runs the program of the method through the processor, the following steps are performed: receiving a speech recognition service test request for a target application; a speech recognition unit is used to convert the speech data of the target application into a text sequence based on a speech recognition model with dynamically variable model parameters based on each group of model parameters for multiple groups of model parameters; a text sequence feedback unit is used to feedback the text sequence corresponding to each group of model parameters to the requester, so that the requester can determine the speech recognition performance of each group of model parameters, and determine the target model parameters corresponding to the target application based on the performance.
[0198] The present application also provides a speech recognition model construction device, comprising:
[0199] A training data determination unit, used to determine a training data set, wherein the training data includes: speech data and text sequence annotation information;
[0200] A network construction unit, used to construct the network structure of the model;
[0201] The network training unit is used to perform iterative training on the model according to dynamically determined model parameters to obtain a speech recognition model with dynamically variable model parameters.
[0202] The present application also provides an electronic device, comprising:
[0203] Processor; and
[0204] The memory is used to store a program for implementing a method for building a speech recognition model. After the device is powered on and runs the program of the method through the processor, the following steps are performed: determining a training data set, the training data including: speech data and text sequence annotation information; building a network structure of the model; and performing iterative training on the model according to dynamically determined model parameters to obtain a speech recognition model with dynamically variable model parameters.
[0205] The present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, and when the computer-readable storage medium is run on a computer, the computer executes the above-mentioned various methods.
[0206] The present application also provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the above-mentioned various methods.
[0207] Compared with the prior art, this application has the following advantages:
[0208] The speech recognition system provided in the embodiment of the present application collects speech data of different applications through multiple clients and sends the speech data to the server; the server learns from the training sample set to obtain a speech recognition model with dynamically changeable model parameters, and determines the model parameters of the model used by each application; determines the model parameters of the target application for the speech data sent by the client; uses the model parameters of the target application as the model parameters of the speech recognition model, and converts the speech data into a text sequence through the speech recognition model based on the model parameters of the target application; adopts this processing method to realize a streaming speech recognition system with controllable model parameters (such as the model size that affects the amount of calculation, and the latency that affects the recognition response speed); during speech recognition, the corresponding model parameters are configured according to the actual application scenario requirements, thereby achieving that a general model can meet the different requirements of different applications for the amount of calculation and latency; therefore, it can effectively save system resources, reduce model maintenance costs, improve the scalability of the model in application scenarios, and improve the model deployment efficiency in new application scenarios. In addition, the performance of this dynamically trained model is better than that of a model trained separately with fixed model parameters. If SCAMA-based streaming end-to-end speech recognition is adopted, the performance of offline speech recognition based on the whole sentence attention mechanism can be achieved, thus effectively improving the speech recognition performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0209] Figure 1 A schematic diagram of the structure of an embodiment of a speech recognition system provided by the present application;
[0210] Figure 2 A schematic diagram of a scenario of an embodiment of a speech recognition system provided by the present application;
[0211] Figure 3 A schematic diagram of device interaction of an embodiment of a speech recognition system provided by the present application;
[0212] Figure 4 A schematic diagram of a model of an embodiment of a speech recognition system provided by the present application;
[0213] Figure 5 The present application provides another model schematic diagram of an embodiment of a speech recognition system. DETAILED DESCRIPTION
[0214] Many specific details are described in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present application, so the present application is not limited by the specific implementation disclosed below.
[0215] In this application, a speech recognition system, method and device, a speech recognition model construction method and device, a smart speaker, a smart TV, a food ordering device, a smart mobile device, a car voice assistant device, a court trial device, and an electronic device are provided. Each solution is described in detail in the following embodiments.
[0216] First embodiment
[0217] Please refer to Figure 1 , which is a schematic diagram of an embodiment of a speech recognition system of the present application. The speech recognition system provided in this embodiment includes: a server 1 and a client 2.
[0218] The server 1 may be a server deployed on a cloud server, or may be a server dedicated to implementing a speech recognition system, which may be deployed in a data center.
[0219] Client 2 includes but is not limited to terminal devices such as smart speakers, smart TVs, ordering equipment, court trial equipment, in-vehicle voice assistant equipment, personal computers, tablet computers, smart phones, etc., and can also be servers within the enterprise local area network, etc.
[0220] Please refer to Figure 2, which is a scenario diagram of the speech recognition system of the present application. The server 1 and the client 2 can be connected through a network, such as the terminal device can be connected to the Internet through WIFI, etc. The user and the terminal device can interact through voice. Taking the smart speaker as an example, the user gives a voice command to the smart speaker (such as what is the weather like today, call someone, etc.), and the smart speaker sends the user's voice data to the server; the server determines the text sequence of the voice data through a speech recognition model with dynamically variable model parameters; and performs voice interaction processing based on the recognized text sequence. In this embodiment, the server can provide speech recognition services for multiple applications through a common speech recognition model. By configuring model parameters for different applications, the model can meet the different requirements of different applications for computing power and latency.
[0221] Please refer to Figure 3 , which is a device schematic diagram of the speech recognition system of the present application. In this embodiment, the client is used to collect speech data of the target application and send the speech data to the server; the server is used to learn a speech recognition model with dynamically variable model parameters from a training sample set and determine the model parameters of the model used by each application; the model parameters of the target application are determined for the speech data sent by the client; the model parameters of the target application are used as the model parameters of the speech recognition model, and the speech data is converted into a text sequence through the speech recognition model based on the model parameters of the target application.
[0222] The speech recognition model is a speech recognition model with dynamically variable model parameters, and its model parameters can change with application requirements. The model can be a general speech recognition model that can provide online speech recognition services for multiple applications, and can meet the different requirements of different applications for computing power and latency. The application can be an application related to smart speakers, a food ordering application, a TV program on demand application, and so on.
[0223] The structure of the speech recognition model can be a non-end-to-end speech recognition model or an end-to-end (End2End) speech recognition model. The non-end-to-end speech recognition model includes an independent acoustic model and a language model. The pronunciation sequence can be first recognized by the acoustic model, and then the text sequence can be determined by the language model. The end-to-end speech recognition model can adopt a speech recognition framework that combines the acoustic model and the language model into one, so that there is no error propagation effect between modules, which can significantly improve the speech recognition performance and greatly reduce the complexity of system training.
[0224] The model parameters may be parameters that affect the computational complexity of the model, including the model size, such as the size of the audio encoder in an end-to-end speech recognition model. The modules in the speech recognition model (such as the audio encoder, decoder, etc.) may adopt a neural network structure, and the model size may be the number of layers of the module neural network, or the number of neurons in a certain layer, or may include both the number of layers and the number of neurons in the neural network.
[0225] In one example, the speech recognition model adopts a streaming end-to-end speech recognition model, which includes: an audio encoder and a decoder. Among them, the audio encoder adopts a neural network structure, and the size of the network is variable, that is, configurable, so it can also be called a dynamic encoder. The configurable model size of the speech recognition model can be the number of neural network layers of the dynamic encoder, or the number of neurons in a certain layer, and can also include both the number of layers and the number of neurons in the neural network. In this case, the server needs to learn from the training sample set to obtain a speech recognition model with dynamically variable model parameters, which can be achieved in the following way: according to the dynamically determined model size, iterative training is performed on the model. Among them, a training sample may include voice data and text annotation information, and the text annotation information can be manually annotated.
[0226] In specific implementation, the dynamically determined model size can be determined in the following manner: a model size is arbitrarily selected from a plurality of preset model sizes. Table 1 shows a model size parameter table in this embodiment.
[0227] Model parameter name Model parameter candidate values Model parameter type Number of neural network layers 3,5,10 Model size Number of neurons 128,256,512,1024 Model size
[0228] Table 1. Model size parameter table
[0229] As can be seen from Table 1, the model size may include two parameters: the number of neural network layers and the number of neurons. Each parameter may be set with multiple candidate values, which are used to arbitrarily select the parameter value for each iteration during the iterative training of the model. For example, during each iteration (such as 100 samples each time) of the model training process, a model size may be randomly selected from the model size candidate table (including: 128, 256, 512, 1024 and other model size values) as the model size of the current iteration. During each iterative training, the model size parameter changes dynamically, and finally a speech recognition model with a dynamically variable model size is obtained through training.
[0230] After training the speech recognition model with dynamically variable model parameters, the corresponding model parameters can be configured for each application according to the actual application scenario requirements. Table 2 shows the model parameter configuration table.
[0231] Application Name Model parameters Smart speaker applications Neural network layers: 10, number of neurons: 1024 Ordering App Neural network layers: 3, number of neurons: 256 Court trial application Neural network layers: 5, number of neurons: 512 In-car voice assistant Neural network layers: 3, number of neurons: 128 Smart TV Apps Neural network layers: 5, number of neurons: 512 Smartphone Apps Neural network layers: 3, number of neurons: 512 …
[0232] Table 2. Model parameter configuration table
[0233] As can be seen from Table 2, different model size values can be set for different applications, such as setting the model size of application 1 and application 2 to 128, the model size of application 3 to 512, and so on. In this way, after the server receives the voice data to be recognized from the client, it can first determine the model parameters of the application corresponding to the voice data, and then convert the voice data into a text sequence through a speech recognition model based on the model parameters.
[0234] The model parameter may also be a latency value that affects the speech recognition reaction speed. For example, when performing speech recognition online, if the speech recognition latency value is set to 150 milliseconds, speech recognition is performed every 150 milliseconds. In this way, the user can perceive that the speech recognition reaction speed is only 150 milliseconds later than the actual speaking, rather than waiting for the entire sentence to be finished before performing speech recognition.
[0235] In one example, the speech recognition model includes: an audio encoder with a fixed model size, a feature data determination module and a decoder. The audio encoder is used to determine the audio feature data of the speech data, and store the audio feature data in a chunk memory (Chunk Memory) according to the delay value of the target application; the feature data determination module is used to determine the feature data corresponding to the word in the speech data according to the audio feature data in the chunk memory; the decoder is used to determine the word in the speech data according to the feature data of the word to form the text sequence.
[0236] The configurable model parameters of the speech recognition model include the delay value. For different applications using the model, the model determines the size of each block according to the delay value of the application, and the block size is related to the delay value. For example, if the frame length of the audio frame is 60ms, if the delay value of application A is 300ms, the number of audio frames memorized in one block is 5 frames; if the delay value of application B is 600ms, the number of audio frames memorized in one block is 10 frames.
[0237] In a segment of speech data (e.g., 600ms of speech data), the sounds of multiple words may be included, or the sounds of no words may be included, but noise or background music may be included instead. The decoder can sequentially identify each word in the speech data, and when identifying a word, the word can be determined based on the relevant feature data. For the convenience of description in this embodiment, the feature data (input data of the decoder) used to determine the word is referred to as the feature data of the word.
[0238] Each character in a piece of speech data may have different feature data. The feature data of a character may include acoustic information related to the pronunciation of the character, and may also include contextual semantic information of the character. The contextual semantic information of a character may affect the recognition of the character. For example, the second character in "happy" and "lucky" has the same pronunciation, "xing", but the contextual semantics of these two characters are different. Therefore, the characters with the same pronunciation are recognized as two different characters, which can improve the recognition accuracy of the characters. After each character in a piece of speech data is recognized, the text sequence of this piece of speech data is obtained.
[0239] In the case where the delay value is adjustable, the server needs to learn from the training sample set to obtain a speech recognition model with dynamically variable model parameters, which can be implemented in the following manner: according to the dynamically determined delay value, iterative training is performed on the model. In specific implementation, the dynamically determined delay value can be determined in the following manner: arbitrarily selecting a delay value from a plurality of preset delay values. Table 3 shows a delay value parameter table in this embodiment.
[0240] Model parameter name Model parameter candidate values Delay value 150ms, 300ms, 600ms, 900ms, 1200ms…
[0241] Table 3. Delay value parameter table
[0242] As can be seen from Table 3, multiple candidate values can be set for the delay value parameter, and the delay value of each iteration can be arbitrarily selected from them during the iterative training of the model. For example, at each iteration (such as 100 samples each time) during the model training process, a delay value can be randomly selected from the delay value candidate table (including: 300ms, 600ms, 900ms, 1200ms, etc.) as the delay of the current iteration. With each iterative training, the delay parameter changes dynamically, and finally a speech recognition model with a dynamically variable delay value is trained.
[0243] After training the speech recognition model with dynamically variable delay parameters, corresponding delay parameters can be configured for each application according to the actual application scenario requirements. Table 2 shows a delay parameter configuration table.
[0244]
[0245]
[0246] Table 4. Delay parameter configuration table
[0247] As can be seen from Table 4, different delay values can be set for different applications, such as setting the delay of application 1 and application 2 to 150ms, the delay of application 3 to 900ms, etc. In this way, after the server receives the voice data to be recognized from the client, it can first determine the delay value of the application corresponding to the voice data, and then convert the voice data into a text sequence through a voice recognition model based on the delay value.
[0248] When the delay value is adjustable, in this embodiment, the server stores the audio feature data output by the audio encoder into a block memory, and the size of the block may be related to the delay value. Through the feature data determination module, the target block related to the word to be recognized can be determined, and the feature data of the word can be determined at least according to the audio feature data of the target block.
[0249] In a specific implementation, the feature data determination module may adopt the following processing method: determine the correspondence between the word and the block memory, that is, determine which word is in which block; determine the target block related to the word to be recognized based on the correspondence, such as to recognize the 12th word, which is in the 3rd block, then the blocks related to the word may include the 1st block, the 2nd block and the 3rd block; determine the feature data corresponding to the word to be recognized based on the audio feature data of the target block, so that the recognized word is not affected by the contextual semantic information; or, determine the feature data corresponding to the word to be recognized based on the audio feature data of the target block and the contextual information of the word to be recognized, so that the recognized word will be affected by the contextual semantic information, so the recognition accuracy of the word is higher.
[0250] In this embodiment, the feature data determination module may include: a predictor for determining the length of the text included in each block, so that the correspondence between the word and the block memory can be more accurately determined based on the text length. For example, in the model use stage, the delay value of application A is 300ms. For the voice data to be recognized (300ms), the audio feature data (also called audio feature encoding data) of the voice data can be determined through the audio encoder and stored in the block memory; the audio feature data of 5 audio frames with a block memory length of 300ms is input into the predictor, and the predictor is used to determine how many words this section of voice data includes.
[0251] In specific implementation, other methods may be used to determine the correspondence between words and blocks, such as determining the correspondence by identifying a terminal symbol, etc. Experiments have shown that the correspondence between words and blocks can be determined more accurately by using a predictor.
[0252] In this embodiment, the attention module is used to determine the feature data of the word according to the text length of each block and the audio feature data in the block memory. For example, the attention module is used to determine the correspondence between the word and the block according to the text length of each block; then, the corresponding relationship between the word and the block is determined according to the relationship with the word to be recognized. l+1 The relevant target block determines the key-value pair, where the key can be semantic information and the value can be audio coding feature information; for y l+1 , first calculate y lThe similarity between the contextual semantic information and each key is used to determine the weight of each key, so that we can determine which information is important and which is not important. Then, we perform a weighted sum operation on the value to determine the value that is similar to y. l+1 Corresponding feature data; finally, through the decoder, according to y l+1 Corresponding feature data, determine y l+1 Among them, the weighted sum of feature data may include y l+1 c before 1 to c m (c m Represents y l+1 The feature data after weighted summation is the data output by the feature data determination module, that is, the input data of the decoder.
[0253] The predictor can be trained at the same time as the speech recognition model is trained. The training data of the speech recognition model may include speech data and text annotation information, and the text annotation information may be annotated manually. In this case, in order to train the predictor, the training data also includes text length annotation information corresponding to each block. The entire model needs to calculate two loss values during the training process, one is the loss value of the model output data, and the other is the loss value of the text length output by the predictor.
[0254] During the model training process, the speech data in the training data is usually much longer than the candidate delay value. For example, assuming the delay value is 600ms, a frame of speech is 60ms, so the size of a block is 10 frames; during the model training stage, the feature input to the model is long speech, which belongs to the case of pseudo-stream decoding. For example, if the input speech length is 15 seconds, that is, 250 frames (15s*1000 / 60ms=250 frames), a total of 250 / 10=25 blocks. At this time, semantic information can be calculated within each block, and the text length of each of the 25 blocks can be calculated at one time. At this time, the predictor outputs the number of output characters contained in each block to know the attention module, which block of memory needs to be paid attention to when decoding the current character (which character). At the same time, historical characters can also be used as input to predict the current text. For example, the predictor outputs that block 1 includes 15 characters, block 2 includes 18 characters, block 3 includes 20 characters, ..., block 10 includes 13 characters. To decode the 51st character, the attention module needs to pay attention to the audio feature data of blocks 1, 2, and 3.
[0255] However, since manual annotation is usually performed on a relatively complete speech data (such as the speech data of a sentence), the speech data in the training data is usually much longer than the candidate delay value, so it is impossible to manually annotate the text length of each block. For example, a long speech of the training sample includes 25 blocks, and it is impossible to manually divide the text sequence of the long speech into sub-texts corresponding to each block, and then determine the text length of each block. In order to solve this problem, this embodiment automatically determines the annotation data of the text length of each block through the traditional CDC method.
[0256] During the model usage phase, the input feature is a block, such as a 600ms block; the first 600ms is used to calculate the first memory block, the second 600ms is used to get the second memory block, and so on; at the same time, the predictor predicts the number of characters in each memory block. If there are characters, the attention module can be used to pay attention to this memory block (there can be multiple, including historical memory blocks) to predict the current text. If there are no characters, wait for the next memory block to arrive, and then see if there are characters in this memory block, and repeat the above process.
[0257] exist Figure 4 In, c 1 Indicates the first chunk, c 2 Indicates the second block; n 1 、n 2 Respectively represent the number of characters included in the first block and the second block; y l+1 ∈c m Indicates that the l+1th output word is in the mth block, from which we can determine which block the word to be processed is in, and then we can use the context y l Calculate y together with the calculated mth block l+1 The output text of . Experiments show that the performance of the model trained with dynamic delay values is better than the model trained with fixed delay values because the predictor can determine the text length in each block more accurately.
[0258] It should be noted that during the model usage phase, the latency value of the target application can be set to a latency value different from that of the training phase. For example, if the latency value during training is 150ms, 300ms, 600ms, etc., then the latency value during the usage phase only needs to be greater than the shortest latency value of 150ms during the training phase, specifically 200ms, 320ms, etc. Experiments have shown that even if the latency value during the usage phase is different from that during the training phase, the same speech recognition performance can be obtained, thus effectively reducing the number of speech recognition models with different latency values.
[0259] like Figure 5As shown, in another example, the model is a streaming end-to-end speech recognition model, which can improve online speech recognition services for various applications; the model includes: a dynamic encoder, a block memory, an attention network, and a decoder. Among them, the attention network can realize the function of the above-mentioned feature data determination module, that is, the structure of the data determination module is an attention network. In this model, the dynamically adjustable parameters (configurable parameters) include both model size and delay value. At each iteration in the model training process, a model size can be randomly selected from the model size candidate table as the model size of the current iteration; a delay value can be randomly selected from the delay value candidate table as the delay of the current iteration; each iteration training, the model parameters change dynamically; when the model is used, according to the actual application scenario requirements, the corresponding model parameters (model size and delay value) are configured, so that a general speech recognition model can be adopted. In the actual scenario, the model parameters are configured according to the application requirements, which can not only improve the scalability of the model in the application scenario, and the model deployment efficiency in the new application scenario, but also reduce the number of models, training costs and maintenance costs, and save system resources.
[0260] Use Figure 5 The model shown, the server needs to perform speech recognition through the speech recognition model based on the target model parameters, and the following processing can be specifically adopted. First, the audio feature data of the speech data is determined by the audio encoder and stored in the block memory; then, the target block related to the word can be determined by the attention network according to the delay value of the target application; finally, the text sequence can be determined by the decoder according to the audio feature data of the target block and the historical text.
[0261] like Figure 5 As shown, in this embodiment, the speech recognition model includes four modules: 1) Dynamic Encoder; 2) Chunk Memory; 3) Attention Block; 4) Decoder. Figure 5 The structure and working mode of these modules are explained in detail.
[0262] 1) Dynamic Encoder: It can be a multi-layer neural network. There are many choices of neural networks, such as DFSMN, CNN, BLSTM, Transformer, etc. The size of each layer can be randomly selected from the model size candidate table shown in Table 1 above. Figure 4In the example, taking one layer as an example, the candidate values of the number of neurons can be [128, 256, 512, 1024]. During training, at each iteration, one of the numbers is randomly selected as the model size of the current iteration, and the above process is repeated in the next iteration. During decoding, the corresponding model size is configured according to the actual application scenario requirements.
[0263] 2) Chunk Memory: for the input T-frame acoustic features (X 1 , X 2 ,…,X T ), after passing through the dynamic encoder, block memory is generated. The size of the block is expressed as a latency value (latency size), as shown in the dotted rectangular box in the figure. During the latency training process, each iteration changes dynamically; during decoding, the corresponding latency size is configured according to the actual application scenario requirements.
[0264] 3) Attention Block, including Predictor and Attention module. The function of Predictor is to predict how many outputs need to be predicted in each block. During the training process, the number of words actually contained in each block can be marked to determine, and the Predictor is trained through this marking. In this way, during the prediction process, the number of words that may be contained in each block can be obtained through the Predictor; the Attention module is guided by the Predictor to determine which block the current attention needs to be in, so that the output can be predicted without the information of the entire sentence, so that streaming recognition can be achieved.
[0265] 4) Decoder: It can also be a multi-layer neural network, whose functions include: receiving historical prediction output and attention information to predict the next output target, similar to the language model. Figure 5 It can be seen that the word y 1 In C 1 In the block, the word y 2 ,y 3 ,y 4 In C 2 In the block, the word y l ,y l+1 In C m In the block, the word y L In C M block.
[0266] In specific implementations, the network structure of the module is variable, such as which network structure the encoder and decoder use is variable, and only one of the network options is shown.
[0267] Figure 5The system provided is a streaming speech recognition system with controllable dynamic model size and delay. During training, at each iteration, a number is randomly selected from the candidate table as the model size of the current iteration. The delay is similar, and it changes dynamically at each iteration training. During decoding, the corresponding model and delay size are configured according to the actual application scenario requirements. After research, it is found that the performance of this dynamically trained model is better than that of a model trained separately with a fixed model size and delay. In this way, a model can be used and configured according to needs in actual scenarios to reduce maintenance costs.
[0268] In one example, the system may adopt a SCAMA streaming solution. Experiments have shown that the performance of SCAMA-based streaming speech recognition and the offline speech recognition based on the whole sentence attention mechanism are substantially intact.
[0269] The voice interaction system provided in the embodiment of the present application collects voice data of different applications through multiple clients and sends the voice data to the server; the server learns from the training sample set to obtain a voice recognition model with dynamically changeable model parameters, and determines the model parameters of the model used by each application; determines the model parameters of the target application for the voice data sent by the client; uses the model parameters of the target application as the model parameters of the voice recognition model, and converts the voice data into a text sequence through the voice recognition model based on the model parameters of the target application; adopts this processing method to realize a streaming voice recognition system with controllable model parameters (such as the model size that affects the amount of calculation, and the latency that affects the recognition reaction speed); during voice recognition, the corresponding model parameters are configured according to the actual application scenario requirements, thereby achieving that a general model can meet the different requirements of different applications for the amount of calculation and latency; therefore, it can effectively save system resources, reduce model maintenance costs, improve the scalability of the model in application scenarios, and improve the model deployment efficiency in new application scenarios. In addition, the performance of this dynamically trained model is better than that of a model trained separately with fixed model parameters. If SCAMA-based streaming end-to-end speech recognition is adopted, the performance of offline speech recognition based on the whole sentence attention mechanism can be achieved, thus effectively improving the speech recognition performance.
[0270] Second embodiment
[0271] In the above embodiment, a speech recognition system is provided. Correspondingly, the present application also provides a speech recognition method, and the execution subject of the method can be a smart speaker, a voice vending machine, a voice ticket machine, a chat robot and other devices. The method corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, please refer to the corresponding parts in the first embodiment.
[0272] The method may include the following steps: collecting voice data of a target application, and sending the voice data to a server, so that the server learns a voice recognition model with dynamically changeable model parameters from a training sample set; determining target model parameters corresponding to the target application for the voice data; and converting the voice data into a text sequence through the voice recognition model based on the target model parameters.
[0273] Third embodiment
[0274] In the above embodiment, a speech recognition method is provided, and correspondingly, the present application also provides a speech recognition device. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here, please refer to the corresponding parts in the first embodiment.
[0275] A speech recognition device provided by the present application comprises:
[0276] A voice data collection unit, used to collect voice data of a target application;
[0277] A voice data sending unit is used to send the voice data to a server so that the server can learn a voice recognition model with dynamically changeable model parameters from a training sample set; determine target model parameters corresponding to the target application for the voice data; and convert the voice data into a text sequence through the voice recognition model based on the target model parameters.
[0278] Fourth embodiment
[0279] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0280] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: voice data of a target application is collected, and the voice data is sent to a server so that the server learns a voice recognition model with dynamically variable model parameters from a training sample set; target model parameters corresponding to the target application are determined; and the voice data is converted into a text sequence through the voice recognition model based on the target model parameters.
[0281] The electronic device may be a smart speaker, a smart phone, a smart TV, a voice ordering machine, a voice vending machine, a voice ticketing machine, a chat robot, or other device that has a voice recognition service requirement.
[0282] Fifth embodiment
[0283] In the above embodiment, a speech recognition system is provided. Correspondingly, the present application also provides a speech recognition method, and the execution subject of the method can be a server, etc. The method corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, and please refer to the corresponding parts in the first embodiment.
[0284] A speech recognition method provided in an embodiment of the present application may include the following steps:
[0285] Step 1: Learn a speech recognition model with dynamically variable model parameters from the training sample set.
[0286] In one example, the model parameters include: model size; the model size can be the number of layers of the neural network, the number of neurons, or the number of layers and the number of neurons of the neural network. Accordingly, step 1 can be implemented as follows: according to the dynamically determined model size, iterative training is performed on the model.
[0287] In a specific implementation, the dynamically determined model size may be determined in the following manner: a model size is arbitrarily selected from a plurality of preset model sizes, such as traversing each preset model size, or a preset model size is randomly selected.
[0288] The model may be a streaming end-to-end speech recognition model or a non-end-to-end speech recognition model. The streaming end-to-end speech recognition model may include: an audio encoder and a decoder; and the model size includes: the size of the audio encoder.
[0289] In another example, the model parameters include: a delay value; accordingly, step 1 can be implemented in the following manner: performing iterative training on the model according to the dynamically determined delay value.
[0290] During specific implementation, the dynamically determined delay value may be determined in the following manner: a delay value is arbitrarily selected from a plurality of preset delay values.
[0291] In the case where the controllable model parameter is a delay value, the model may be a streaming end-to-end speech recognition model; the model may include: an audio encoder, a feature data determination module, and a decoder.
[0292] It should be noted that, in the model application stage, the delay value of the target application can be the same as the preset delay value in the training stage, or a delay value other than the preset delay value. This processing method can effectively reduce the number of models, training costs and maintenance costs.
[0293] Step 2: Determine the target model parameters corresponding to the target application.
[0294] In this embodiment, the execution subject of the method is a server, which can receive voice data of a target application sent by a terminal device, and determine target model parameters corresponding to the target application for the voice data to be processed.
[0295] The model can meet the different requirements of different applications for computing capacity and latency. In specific implementation, the corresponding relationship between the application and the model parameters can be pre-stored, and the model parameters of the target application can be determined based on the relationship.
[0296] In specific implementation, the model parameters of each application can also be determined according to the performance requirements of each application for speech recognition. The performance requirements can be the response speed (real-time) of speech recognition, the accuracy of speech recognition, etc.
[0297] In specific implementation, the model parameters of the target application may also be determined based on the device performance information (such as storage resources, computing resources, etc.) on which the target application is deployed. In this case, it is usually necessary to deploy the speech recognition model and the target application in the same device.
[0298] Step 3: Convert the speech data into a text sequence through the speech recognition model based on the target model parameters.
[0299] When the controllable model parameter is a delay value, the model can be a streaming end-to-end speech recognition model; the model may include: an audio encoder, a feature data determination module, and a decoder; accordingly, step 3 may include the following sub-steps: 3.1) determining the audio feature data of the speech data through the audio encoder, and storing the audio feature data in a block memory according to the delay value of the target application; 3.2) determining the feature data corresponding to the characters in the speech data according to the audio feature data in the block memory through the feature data determination module; 3.3) determining the characters in the speech data according to the feature data of the characters through the decoder to form the text sequence.
[0300] In this embodiment, step 3.2 may include the following sub-steps: 3.2.1) determining the correspondence between the word and the block memory; 3.2.2) determining the feature data corresponding to the word according to the correspondence. In specific implementation, it may be that the target block related to the word to be recognized is first determined according to the correspondence; then the feature data corresponding to the word to be recognized is determined according to the audio feature data of the target block; or the feature data corresponding to the word to be recognized is determined according to the audio feature data of the target block and the context information of the word to be recognized.
[0301] In this embodiment, the feature data determination module may further include a predictor; step 3.2 may further include the following sub-steps: 3.2.0) determining the text length included in each block through the predictor; correspondingly, the correspondence between characters and blocks may be determined based on the text length.
[0302] In one example, if a first user associated with a target application sends a resource object corresponding to the target model parameters to a second user associated with the model, the voice data is converted into a text sequence through the voice recognition model based on the target model parameters. The first user may be a user of the voice recognition model, such as a developer of the target application. The second user may be a developer or manager of the voice recognition model. The resource object may be currency or virtual currency, etc. If the resource object is currency, the currency may be transferred from the first user to the second user through a third-party payment platform. By adopting this processing method, when different applications use the same voice recognition model for voice recognition, the resource object corresponding to the configured model parameters may be sent to the second user, so that the voice data of the application can be processed using the voice recognition model based on the configured parameters.
[0303] Sixth embodiment
[0304] In the above embodiment, a speech recognition method is provided, and correspondingly, the present application also provides a speech recognition device. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the first embodiment are not repeated here, please refer to the corresponding parts in the first embodiment.
[0305] A speech recognition device provided by the present application comprises:
[0306] A model building unit, used for learning from a training sample set to obtain a speech recognition model with dynamically variable model parameters;
[0307] A model parameter determination unit, used to determine target model parameters corresponding to a target application;
[0308] A model prediction unit is used to convert the speech data of the target application into a text sequence through the speech recognition model based on the target model parameters.
[0309] Seventh embodiment
[0310] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0311] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a speech recognition method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: a speech recognition model with dynamically variable model parameters is learned from a training sample set; target model parameters corresponding to a target application are determined; and speech data of the target application is converted into a text sequence through the speech recognition model based on the target model parameters.
[0312] Eighth embodiment
[0313] In the above embodiment, a speech recognition system is provided. Correspondingly, the present application also provides a speech recognition method, and the execution subject of the method can be a server, etc. The method corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, and please refer to the corresponding parts in the first embodiment.
[0314] A speech recognition method provided in an embodiment of the present application may include the following steps:
[0315] Step 1: Learn a speech recognition model with dynamically variable model parameters from the training sample set.
[0316] The similarities between the method provided in this embodiment and the method executed by the server in the above-mentioned embodiment 2 include: both can build the speech recognition model on the server; the differences include: after building the model, the method provided in this embodiment sends the model to other devices (such as the client), and the other devices can perform speech recognition processing locally without calling the speech recognition service provided by the server, while in embodiment 2, the server provides speech recognition services to multiple clients through the speech recognition model.
[0317] Step 2: Determine the target model parameters corresponding to the target application.
[0318] The target application may be an application deployed on a server or an application deployed on a terminal device. For example, the target application may be a self-service ordering application deployed on a food ordering machine, an automatic vending application on a vending machine, a TV program on-demand application on a smart TV, an automatic question-and-answer service on a smart speaker, and so on.
[0319] In one example, step 2 may include the following sub-steps: 2.1A) determining speech recognition performance requirement information of the target application; 2.2A) determining the target model parameters based on the performance requirement information.
[0320] The performance requirement information may be speech recognition accuracy, speech recognition reaction speed (also known as latency, real-time), etc. The model size may be determined according to the speech recognition accuracy. Table 5 lists the corresponding relationship between speech recognition accuracy and model size.
[0321] Speech recognition accuracy Model size parameters 98% Neural network layers: 10, number of neurons: 1024 95% Neural network layers: 3, number of neurons: 256 90% Neural network layers: 5, number of neurons: 512 85% Neural network layers: 3, number of neurons: 128 …
[0322] Table 5. Correspondence between speech recognition accuracy and model size
[0323] It can be seen from Table 5 that different model size parameters correspond to different speech recognition accuracy. According to the speech recognition accuracy requirement information of the target application, the table can be queried to determine the corresponding target model parameters, and then the speech data can be converted into a text sequence through the speech recognition model based on the model parameters.
[0324] In a specific implementation, the execution subject of the method receives a speech recognition model acquisition request for a target application sent by other devices; the request may include speech recognition performance requirement information of the target application.
[0325] In specific implementation, the method may further include the following steps: determining the device performance requirement information of the target device according to the performance requirement information; sending the device performance requirement information to a management device related to the target application so that the management device displays the device performance requirement information; correspondingly, sending the speech recognition model based on the target model parameters to the target device that meets the device performance requirement information. Table 6 lists the correspondence between speech recognition performance and device performance.
[0326]
[0327]
[0328] Table 6. Correspondence between speech recognition performance and device performance
[0329] It can be seen from Table 6 that different speech recognition performances correspond to different device performance parameters. According to the speech recognition performance requirement information of the target application, the table can be queried to determine the corresponding device performance, and then the device performance requirement information is sent to a management device related to the target application (such as a personal computer, etc.) so that the management device displays the device performance requirement information. The manager of the target application can configure the target device according to the device performance requirement information so that the performance of the target device can ensure the normal operation of the speech recognition model.
[0330] In another example, step 2 may include the following sub-steps: 2.1B) determining the performance information of the device running the target application; 2.2B) determining the target model parameters based on the device performance information. This processing method allows the appropriate model parameters to be determined based on the performance of the existing device of the model application party to ensure that the target device can run the speech recognition model normally.
[0331] In specific implementation, the speech recognition performance corresponding to the target device performance can be determined according to Table 6, and then the model size corresponding to the recognition accuracy can be determined according to Table 5.
[0332] In specific implementation, the device performance information includes: computing resource information and storage resource information; step 2.2B may include the following sub-steps: 2.2B.1) determining the model size based on the computing resource information; 2.2B.2) determining the delay value based on the storage resource information.
[0333] In this case, the method may further include the following steps: determining speech recognition performance information based on the target model parameters; sending the performance information to a management device related to the target application so that the management device displays the performance information. The performance information may include the real-time degree of speech recognition; accordingly, the target model parameters include: delay value. The performance information may also include speech recognition accuracy; accordingly, the target model parameters include: model size. By adopting this processing method, the first user of the target application can know the speech recognition performance that can be achieved by the speech recognition model running on the target device.
[0334] Step 3: Send the speech recognition model based on the target model parameters to a target device running a target application, so that the target application converts speech data into a text sequence through the speech recognition model based on the target model parameters.
[0335] In this embodiment, the method may also include the following steps: determining resource information corresponding to the target model parameters; sending the resource information to a first user related to the target application; if the first user sends the resource object to a second user related to the model, then sending the speech recognition model based on the target model parameters to the target device.
[0336] Ninth embodiment
[0337] In the above embodiment, a speech recognition method is provided, and correspondingly, the present application also provides a speech recognition device. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the eighth embodiment are not repeated here, please refer to the corresponding parts in the eighth embodiment.
[0338] A speech recognition device provided by the present application comprises:
[0339] A model building unit, used for learning from a training sample set to obtain a speech recognition model with dynamically variable model parameters;
[0340] A model parameter determination unit, used to determine target model parameters corresponding to a target application;
[0341] The model sending unit is used to send the speech recognition model based on the target model parameters to a target device running a target application, so that the target application converts speech data into a text sequence through the speech recognition model based on the target model parameters.
[0342] Tenth embodiment
[0343] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0344] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a speech recognition method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: a speech recognition model with dynamically changeable model parameters is learned from a training sample set; target model parameters corresponding to a target application are determined; and the speech recognition model based on the target model parameters is sent to a target device running a target application, so that the target application converts speech data into a text sequence through the speech recognition model based on the target model parameters.
[0345] Eleventh Embodiment
[0346] In the above-mentioned Embodiment 8, a speech recognition method is provided. Correspondingly, the present application also provides a speech recognition method, and the execution subject of the method can be a server or a terminal device, etc. The method corresponds to the embodiment of the above-mentioned method. The parts of this embodiment that are the same as the eighth embodiment are not repeated here, and please refer to the corresponding parts in Embodiment 8.
[0347] A speech recognition method provided in an embodiment of the present application may include the following steps:
[0348] Step 1: Send a request to the server to obtain the speech recognition model for the target application;
[0349] Step 2: receiving a speech recognition model with dynamically variable model parameters based on target model parameters corresponding to the target application sent back by the server;
[0350] Step 3: Convert the speech data into a text sequence through the speech recognition model based on the target model parameters.
[0351] In one example, the method may further include the following steps: determining speech recognition performance requirement information of the target application; the request includes the performance requirement information, so that the server determines the target model parameters according to the performance requirement information.
[0352] In specific implementation, the method may also include the following steps: receiving device performance requirement information for running the target application determined according to the performance requirement information and sent by the server; displaying the device performance requirement information to facilitate determining the target device that meets the device performance requirement information, so that the server sends the speech recognition model based on the target model parameters to the target device.
[0353] In another example, the method may further include the following steps: determining performance information of a device running the target application; the request includes the device performance information, so that the server determines the target model parameters according to the device performance information.
[0354] During specific implementation, the method may further include the following steps: receiving speech recognition performance information corresponding to the target model parameters sent by the server; and displaying the speech recognition performance information.
[0355] In one example, the method may also include the following steps: receiving resource information corresponding to the target model parameters sent by the server; sending the resource object to a second user related to the model, so that the server sends the speech recognition model based on the target model parameters.
[0356] In one example, the method may also include the following steps: receiving a test system of a speech recognition model based on multiple groups of model parameters sent by a server; converting speech data into a text sequence through a speech recognition model based on each group of model parameters, so as to determine the speech recognition performance of each group of model parameters; determining the target model parameters, and sending the target model parameters to the server. This processing method allows the target application to test the performance of speech recognition using speech recognition models with different model parameters, such as recognition accuracy and delay, so that users can determine the required target model parameters based on the actual perceived speech recognition performance; therefore, the user experience can be effectively improved.
[0357] Twelfth Embodiment
[0358] In the above embodiment, a speech recognition method is provided, and correspondingly, the present application also provides a speech recognition device. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the eleventh embodiment are not repeated here, please refer to the corresponding parts in the eleventh embodiment.
[0359] A speech recognition device provided by the present application comprises:
[0360] A request sending unit, used for sending a speech recognition model acquisition request for a target application to a server;
[0361] A model receiving unit, used for receiving a speech recognition model with dynamically variable model parameters based on target model parameters corresponding to a target application, which is sent back by a server;
[0362] A speech recognition unit is used to convert speech data into a text sequence through the speech recognition model based on the target model parameters.
[0363] Thirteenth Embodiment
[0364] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0365] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a speech recognition method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: a speech recognition model acquisition request for a target application is sent to a server; a speech recognition model with dynamically variable model parameters based on target model parameters corresponding to the target application is received from the server; and speech data is converted into a text sequence through the speech recognition model based on the target model parameters.
[0366] Fourteenth Embodiment
[0367] In the above embodiment, a speech recognition system is provided. Correspondingly, the present application also provides a speech recognition service upgrade method, and the execution subject of the method can be a server, etc. The method corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, and please refer to the corresponding parts in the first embodiment.
[0368] A method for upgrading a speech recognition service provided in this application may include the following steps:
[0369] Step 1: Determine usage status information of a target application for a speech recognition model whose model parameters are dynamically variable based on a first model parameter.
[0370] The usage status information may include information such as the user's behavior in using the target application.
[0371] For example, if the target application is a voice ordering application deployed on the ordering device, if it is found based on the user's ordering behavior data that the same user often repeats the ordering voice several times when ordering, it means that the voice recognition accuracy may not be enough, so you can go to step 2. For another example, if the target application is a voice transcription application deployed in the trial device, if it is found that the text transcribed by the model is greatly changed during the later correction, it means that the voice recognition accuracy may not be enough, so you can go to step 2.
[0372] For another example, users often leave after completing only part of the voice ordering operation. This means that the voice recognition speed is too slow and the user does not have the patience to wait, so you can go to step 2.
[0373] Step 2: Determine a second model parameter of the speech recognition model based on the usage status information.
[0374] In one example, step 2 may include the following sub-steps: 2.1B) determining the speech recognition performance requirement information of the target application based on the usage status information; 2.2B) determining the second model parameter of the speech recognition model based on the performance requirement information.
[0375] For example, the target application is a voice ordering application deployed on an ordering device. If the usage information shows that the user often needs to repeat the ordering voice several times, it means that the voice recognition accuracy may not be enough and the voice recognition accuracy needs to be improved. For example, based on the accuracy that can be achieved by the first model parameters, it is necessary to increase the accuracy to a higher level. At this time, the second model parameters can be determined based on the higher level of accuracy.
[0376] For another example, the target application is a speech transcription application deployed in a court trial device. If the usage status information shows that the user makes significant changes to the text transcribed by the model during later corrections, it means that the speech recognition accuracy may be insufficient and it is necessary to improve the speech recognition accuracy. For example, based on the accuracy achieved by the first model parameters, the accuracy can be improved by one level. At this time, the second model parameters can be determined based on the higher level of accuracy.
[0377] For example, users often leave after completing part of the voice ordering operation. This means that the voice recognition speed is too slow and the real-time performance of voice recognition needs to be improved. For example, based on the delay achieved by the first model parameter, the delay value can be reduced. At this time, the second model parameter can be determined according to the higher-level delay value.
[0378] During specific implementation, the second model parameters may be determined based on Table 5 of the above embodiment and the re-determined speech recognition performance.
[0379] In one example, step 2 may include the following sub-steps: 2.1A) determining multiple groups of model parameters of the speech recognition model based on the usage status information; 2.2A) sending a test system of the speech recognition model based on multiple groups of model parameters to the device, so that the target application converts the speech data into a text sequence through the speech recognition model based on each group of model parameters, so as to determine the speech recognition performance of each group of model parameters, and determine the second model parameters according to the speech recognition performance. This processing method makes it possible to test the performance of speech recognition of the target application using speech recognition models with different model parameters, such as recognition accuracy and delay, so that the user can determine the required second model parameters according to the speech recognition performance corresponding to the various model parameters actually perceived (which may include speech recognition accuracy, speed, delay, etc.); therefore, the user experience can be effectively improved.
[0380] In specific implementation, not only can the multiple groups of model parameters of the speech recognition model be re-determined according to the usage status information, but the resource information corresponding to each group of model parameters can also be re-determined, so that the user can know the resources required to use various model parameters, and assist the user in determining the second model parameters; if the first user sends the resource object of the second model parameter to the second user, the speech recognition model based on the first model parameter on the device will be updated to a speech recognition model based on the second model parameter.
[0381] Step 3: Configure the model parameters of the speech recognition model on the device running the target application as second model parameters, so that the device converts the speech data into a text sequence through the speech recognition model based on the second model parameters.
[0382] In one example, the controllable model parameters of the model may be set in an encrypted model parameter configuration file, and when the target application calls the model, the model performs speech recognition processing according to the second model parameters in the configuration file.
[0383] In another example, the controllable model parameters and uncontrollable model parameters of the model may be packaged together and completely updated as a whole.
[0384] It can be seen from the above embodiments that the speech recognition service upgrade method provided in the embodiments of the present application determines the usage status information of the target application for the speech recognition model whose model parameters are dynamically changeable based on the first model parameters; determines the second model parameters of the speech recognition model according to the usage status information; configures the model parameters of the speech recognition model on the device running the target application as the second model parameters, so that the device converts the speech data into a text sequence through the speech recognition model based on the second model parameters; this processing method makes it possible to update the controllable model parameters of the speech recognition model whose model parameters are dynamically changeable according to the actual usage of the application, so that the model can meet the speech recognition requirements of the application; therefore, it can effectively ensure the normal operation of the application and improve the availability and practicality of the application.
[0385] Fifteenth Embodiment
[0386] In the above embodiment, a method for upgrading a speech recognition service is provided. Correspondingly, the present application also provides a device for upgrading a speech recognition service. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the fourteenth embodiment are not repeated here. Please refer to the corresponding parts in the fourteenth embodiment.
[0387] A speech recognition service upgrade device provided by the present application includes:
[0388] an application usage status determination unit, configured to determine usage status information of a target application on a speech recognition model having dynamically variable model parameters based on a first model parameter;
[0389] a model parameter determination unit, configured to determine a second model parameter of the speech recognition model according to the usage status information;
[0390] The model parameter updating unit is used to configure the model parameters of the speech recognition model on the device running the target application as second model parameters, so that the device converts the speech data into a text sequence through the speech recognition model based on the second model parameters.
[0391] Sixteenth Embodiment
[0392] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0393] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a method for upgrading a speech recognition service. After the device is powered on and the program of the method is run through the processor, the following steps are performed: determining usage status information of a speech recognition model whose model parameters are dynamically variable based on first model parameters by a target application; determining second model parameters of the speech recognition model based on the usage status information; configuring the model parameters of the speech recognition model on the device running the target application to the second model parameters, so that the device converts speech data into a text sequence through the speech recognition model based on the second model parameters.
[0394] Seventeenth Embodiment
[0395] In the above-mentioned embodiment, a method for upgrading a speech recognition service is provided. Correspondingly, the present application also provides a method for upgrading a speech recognition service, and the execution subject of the method may be a server, etc. The method corresponds to the embodiment of the above-mentioned system. The parts of this embodiment that are the same as those in the fourteenth embodiment are not repeated here, and please refer to the corresponding parts in the fourteenth embodiment.
[0396] A method for upgrading a speech recognition service provided in this application may include the following steps:
[0397] Step 1: Determine usage status information of a target application for a speech recognition model whose model parameters are dynamically variable based on a first model parameter.
[0398] The similarities between the method provided in this embodiment and the method executed by the server in the above-mentioned embodiment fourteen include: both can build the speech recognition model on the server and redefine the model parameters for the target application; the differences include: after building the model, the method provided in embodiment fourteen sends the model to the device on the target application side, so that the target application can perform speech recognition processing locally without calling the speech recognition service provided by the server, while in this embodiment, the server provides speech recognition services to multiple clients through the speech recognition model.
[0399] Step 2: Determine a second model parameter of the speech recognition model based on the usage status information.
[0400] Step 3: Store the correspondence between the target application and the second model parameters, so that for the speech data to be processed for the target application, according to the correspondence, the speech data is converted into a text sequence through the speech recognition model based on the second model parameters.
[0401] In specific implementation, the model parameters in Table 2 and Table 4 in the above embodiments may be changed.
[0402] It can be seen from the above embodiments that the speech recognition service upgrade method provided in the embodiments of the present application determines the usage status information of the target application for the speech recognition model whose model parameters are dynamically changeable based on the first model parameters; determines the second model parameters of the speech recognition model according to the usage status information; stores the correspondence between the target application and the second model parameters, so that for the to-be-processed speech data of the target application, according to the correspondence, the speech recognition model based on the second model parameters converts the speech data into a text sequence; this processing method makes it possible to update the controllable model parameters of the speech recognition model whose model parameters are dynamically changeable according to the actual usage of the application, so that the model can meet the speech recognition requirements of the application; therefore, it can effectively ensure the normal operation of the application and improve the availability and practicality of the application.
[0403] Eighteenth Embodiment
[0404] In the above embodiment, a method for upgrading a speech recognition service is provided. Correspondingly, the present application also provides a device for upgrading a speech recognition service. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the fourteenth embodiment are not repeated here. Please refer to the corresponding parts in the fourteenth embodiment.
[0405] A speech recognition service upgrade device provided by the present application includes:
[0406] The application usage status determination unit is used to determine the usage status information of the target application on the speech recognition model with dynamically variable model parameters based on the first model parameter.
[0407] A model parameter determination unit is used to determine a second model parameter of the speech recognition model according to the usage status information.
[0408] The model parameter updating unit is used to store the correspondence between the target application and the second model parameter, so that the speech data to be processed for the target application is converted into a text sequence according to the correspondence through the speech recognition model based on the second model parameter.
[0409] Nineteenth Embodiment
[0410] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0411] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a method for upgrading a speech recognition service. After the device is powered on and the program of the method is run through the processor, the following steps are performed: determining usage status information of a speech recognition model whose model parameters are dynamically variable based on a first model parameter by a target application; determining a second model parameter of the speech recognition model based on the usage status information; and storing a correspondence between the target application and the second model parameter so that, for speech data to be processed by the target application, the speech data can be converted into a text sequence based on the correspondence through the speech recognition model based on the second model parameter.
[0412] Twentieth Embodiment
[0413] In the above embodiment, a speech recognition system is provided. Correspondingly, the present application also provides a speech recognition service test method, and the execution subject of the method can be a server, etc. The method corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, and please refer to the corresponding parts in the first embodiment.
[0414] A speech recognition service testing method provided in this application may include the following steps:
[0415] Step 1: Receive a speech recognition service test request for the target application.
[0416] In one example, a developer of a voice service application wants to release a voice service application that will be used on a smartphone and will recognize user voices through a server-side voice recognition model. At this time, the application developer needs to conduct a real-machine test on the mobile phone and determine which voice recognition model to use through the real-machine test. Therefore, a voice recognition service test request for the target application can be sent to the server through the smartphone to test which model parameters of the voice recognition model deployed on the server can meet the application's requirements for voice recognition performance.
[0417] Step 2: For multiple groups of model parameters, the speech data of the target application is converted into a text sequence through a speech recognition model with dynamically variable model parameters based on each group of model parameters.
[0418] In this embodiment, the server converts the speech data of the target application into a text sequence through a speech recognition model with dynamically variable model parameters based on each group of model parameters.
[0419] Step 3: Send back the text sequence corresponding to each set of model parameters to the requester so that the requester can determine the speech recognition performance of each set of model parameters and determine the target model parameters corresponding to the target application based on the performance.
[0420] It can be seen from the above embodiments that the speech recognition service testing method provided in the embodiments of the present application receives a speech recognition service test request for a target application; for multiple groups of model parameters, the speech data of the target application is converted into a text sequence through a speech recognition model whose model parameters are dynamically variable based on each group of model parameters; a text sequence corresponding to each group of model parameters is sent back to the requesting party so that the requesting party can determine the speech recognition performance of each group of model parameters, and determine the target model parameters corresponding to the target application based on the performance; this processing method makes it possible to test the performance of speech recognition using speech recognition models with different model parameters of the target application, such as recognition accuracy and delay, so that the user can determine the required model parameters based on the speech recognition performance corresponding to the various model parameters actually perceived (which may include speech recognition accuracy, speed, delay, etc.), so that the model can meet the speech recognition requirements of the application; therefore, it can effectively ensure the normal operation of the application, improve the availability and practicality of the application, and improve the user experience.
[0421] Twenty-first embodiment
[0422] In the above embodiment, a speech recognition service test method is provided, and correspondingly, the present application also provides a speech recognition service test device. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as the twentieth embodiment are not repeated here, please refer to the corresponding parts in the twentyth embodiment.
[0423] A speech recognition service testing device provided by the present application comprises:
[0424] A test request receiving unit, used to receive a speech recognition service test request for a target application;
[0425] A speech recognition test unit, for converting speech data of a target application into a text sequence through a speech recognition model with dynamically variable model parameters based on each set of model parameters, for multiple sets of model parameters;
[0426] The text sequence feedback unit is used to return the text sequence corresponding to each group of model parameters to the requester, so that the requester can determine the speech recognition performance of each group of model parameters and determine the target model parameters corresponding to the target application based on the performance.
[0427] Twenty-second embodiment
[0428] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0429] An electronic device of the present embodiment includes: a processor and a memory; the memory is used to store a program for implementing a speech recognition service test method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: receiving a speech recognition service test request for a target application; a speech recognition unit is used to convert speech data of the target application into a text sequence for multiple groups of model parameters through a speech recognition model with dynamically variable model parameters based on each group of model parameters; a text sequence feedback unit is used to feedback text sequences corresponding to each group of model parameters to the requester, so that the requester can determine the speech recognition performance of each group of model parameters, and determine the target model parameters corresponding to the target application based on the performance.
[0430] Twenty-third embodiment
[0431] In the above embodiment, a speech recognition system is provided. Correspondingly, the present application also provides a speech recognition model construction method, and the execution subject of the method can be a server, etc. The method corresponds to the embodiment of the above system. The parts of this embodiment that are the same as the first embodiment are not repeated here, please refer to the corresponding parts in the first embodiment.
[0432] A method for constructing a speech recognition model provided in this application may include the following steps:
[0433] Step 1: Determine a training data set, where the training data includes speech data and text sequence annotation information.
[0434] Step 2: Construct the network structure of the model.
[0435] Step 3: Perform iterative training on the model according to the dynamically determined model parameters to obtain a speech recognition model with dynamically variable model parameters.
[0436] In one example, the model includes: a streaming end-to-end speech recognition model; the model includes: an audio encoder, a decoder; the model parameters include a model size, and the model size includes the size of the audio encoder.
[0437] In another example, the model includes: a streaming end-to-end speech recognition model; the model includes: an audio encoder, a feature data determination module, and a decoder. The audio encoder is used to determine the audio feature data of the speech data and store the audio feature data in a block memory according to the delay value of the target application; the feature data determination module is used to determine the feature data corresponding to the word in the speech data according to the audio feature data in the block memory; the decoder is used to determine the word in the speech data according to the feature data of the word to form the text sequence.
[0438] In this embodiment, the feature data determination module is further used to determine the correspondence between the word and the block memory, and determine the feature data corresponding to the word according to the correspondence.
[0439] In a specific implementation, the training data may also include: text length annotation information of each block; the data determination module includes: a predictor; the predictor is used to determine the text length included in each block; the feature data determination module is used to determine the correspondence between characters and blocks based on the text length.
[0440] It can be seen from the above embodiments that the speech recognition model construction method provided in the embodiments of the present application determines a training data set, wherein the training data includes: speech data and text sequence annotation information; constructs the network structure of the model; and performs iterative training on the model according to dynamically determined model parameters to obtain a speech recognition model with dynamically variable model parameters. This processing method enables the construction of a speech recognition model with dynamically variable model parameters, and this general model can provide speech recognition services for applications with different speech recognition performance requirements. Therefore, it can effectively reduce the number of models, training costs and maintenance costs, and provide a basis for improving the scalability of application scenarios of speech recognition models.
[0441] Twenty-fourth embodiment
[0442] In the above embodiment, a method for constructing a speech recognition model is provided. Correspondingly, the present application also provides a device for constructing a speech recognition model. The device corresponds to the embodiment of the above method. The parts of this embodiment that are the same as those of the twenty-third embodiment are not repeated here. Please refer to the corresponding parts in the twenty-third embodiment.
[0443] A speech recognition model construction device provided in this application includes:
[0444] A training data determination unit, used to determine a training data set, wherein the training data includes: speech data and text sequence annotation information;
[0445] A network construction unit, used to construct the network structure of the model;
[0446] The network training unit is used to perform iterative training on the model according to dynamically determined model parameters to obtain a speech recognition model with dynamically variable model parameters.
[0447] Twenty-fifth embodiment
[0448] The present application also provides an electronic device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0449] An electronic device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a method for building a speech recognition model. After the device is powered on and runs the program of the method through the processor, the following steps are performed: determining a training data set, the training data including: speech data and text sequence annotation information; building a network structure of the model; performing iterative training on the model according to dynamically determined model parameters to obtain a speech recognition model with dynamically variable model parameters.
[0450] Twenty-sixth embodiment
[0451] The present application also provides a smart speaker. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0452] A smart speaker in this embodiment, the electronic device includes: a processor and a memory; the memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: voice data of a target application is collected, and the voice data is sent to a server so that the server learns a voice recognition model with dynamically variable model parameters from a training sample set; target model parameters corresponding to the target application are determined; and the voice data is converted into a text sequence through the voice recognition model based on the target model parameters.
[0453] The target application may be such speaker skills, such as weather forecast, health check, song on demand, etc.
[0454] Twenty-seventh embodiment
[0455] The present application also provides a food ordering device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0456] A food ordering device according to the present embodiment includes: a processor and a memory; the memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: voice ordering data is collected, and the voice ordering data is converted into ordering text through a speech recognition model with dynamically variable model parameters based on target model parameters corresponding to the ordering application; and ordering processing is performed according to the ordering text.
[0457] In this embodiment, since users use the ordering device to place voice orders and have a high demand for real-time voice recognition, in order to avoid the reduction of real-time voice recognition due to network delay, the voice recognition model is usually deployed locally on the ordering device instead of calling the voice recognition model deployed on the server. At the same time, since the hardware configuration of the ordering device is usually lower than that of the server device, it is impossible to run a voice recognition model with high computational complexity on the ordering device. Therefore, the model size can be set smaller to ensure that the ordering device can run the voice recognition model normally. In addition, since users use the ordering device to place voice orders and have a high demand for real-time voice recognition, in order to avoid long waiting time and queuing problems caused by slow ordering, the model delay value can be set lower to ensure a higher voice recognition feedback speed, such as setting the delay value to 150ms.
[0458] Twenty-eighth Embodiment
[0459] The present application also provides a smart TV. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0460] A smart TV in this embodiment, the electronic device includes: a processor and a memory; the memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run by the processor, the following steps are performed: TV control voice data is collected, and the voice data is converted into TV control text through a voice recognition model with dynamically variable model parameters based on target model parameters corresponding to a TV application; and TV control processing is performed according to the TV control text.
[0461] Twenty-ninth embodiment
[0462] The present application also provides a smart mobile device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0463] An intelligent mobile device of the present embodiment includes: a processor and a memory; the memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: voice data is collected, and the voice data is converted into a text sequence through a voice recognition model with dynamically variable model parameters based on target model parameters corresponding to the device; and voice interaction processing is performed according to the text sequence.
[0464] The intelligent mobile device may be a terminal device such as a smart phone or a PAD.
[0465] Thirtieth Embodiment
[0466] The present application also provides a vehicle-mounted voice assistant device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0467] An in-vehicle voice assistant device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a voice interaction method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: voice data is collected, and the voice data is converted into a text sequence through a voice recognition model with dynamically variable model parameters based on target model parameters corresponding to the device; and voice interaction processing is performed according to the text sequence.
[0468] In this embodiment, since the in-vehicle voice assistant device uses a non-220V power supply, higher requirements are placed on voice real-time performance. Therefore, the voice recognition model is usually deployed locally on the in-vehicle voice assistant device instead of calling the voice recognition model deployed on the server.
[0469] Thirty-first embodiment
[0470] The present application also provides a court trial device. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0471] A court trial device of this embodiment includes: a processor and a memory; the memory is used to store a program for implementing a speech recognition method. After the device is powered on and the program of the method is run through the processor, the following steps are performed: voice data is collected, and the voice data is converted into a text sequence through a speech recognition model with dynamically variable model parameters based on target model parameters corresponding to the device.
[0472] In this embodiment, since the court trial equipment usually has high-performance computing resources but cannot accept network delays and has high requirements for real-time speech recognition, the speech recognition model is usually deployed locally on the court trial equipment rather than calling the speech recognition model deployed on the server.
[0473] Although the present application is disclosed as above in the form of a preferred embodiment, it is not intended to limit the present application. Any technical personnel in this field may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.
[0474] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0475] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0476] 1. Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include non-transitory media such as modulated data signals and carrier waves.
[0477] 2. Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
Claims
1. A speech recognition system, characterized in that: include: The client is used to collect the voice data of the target application and send the voice data to the server; The server is used to learn from a training sample set to obtain a speech recognition model with dynamically variable model parameters, including: performing iterative training on the model according to the model parameters; The model parameters are different for each iterative training, and the model parameters include a delay value; for the voice data sent by the terminal device, the target model parameters corresponding to the target application are determined; and the voice data is converted into a text sequence through the speech recognition model based on the target model parameters.
2. A speech recognition method, characterized in that: include: Learning from a training sample set to obtain a speech recognition model with dynamically variable model parameters includes: performing iterative training on the model according to the model parameters; each time the iterative training has different model parameters, the model parameters include a delay value; Determining target model parameters corresponding to the target application; The speech data of the target application is converted into a text sequence by using the speech recognition model based on the target model parameters.
3. The method according to claim 2, characterized in that The model parameters include: model size; The model size includes: the number of layers and / or the number of neurons of the neural network; The speech recognition model with dynamically variable model parameters obtained by learning from the training sample set includes: According to the model size, iterative training is performed on the model; the model size is different in each iterative training.
4. The method according to claim 3, characterized in that The model size is determined in the following manner: Choose any model size from multiple preset model sizes.
5. The method according to claim 3, characterized in that: The model includes: a streaming end-to-end speech recognition model; The model includes: an audio encoder and a decoder; The model size includes: the size of the audio encoder.
6. The method according to claim 2, characterized in that The delay value is determined in the following manner: Select any delay value from multiple preset delay values; The delay value of the target application includes: a delay value other than the preset delay value.
7. The method according to claim 2, characterized in that The model includes: a streaming end-to-end speech recognition model; The model includes: an audio encoder, a feature data determination module, and a decoder; The step of converting the speech data into a text sequence by using the speech recognition model based on the target model parameters comprises: Determine audio feature data of the speech data through an audio encoder, and store the audio feature data in a block memory according to a delay value of a target application; Determine, by a feature data determination module, feature data corresponding to the words in the speech data according to the audio feature data in the block memory; The decoder determines the words in the speech data based on the feature data of the words to form the text sequence.
8. The method according to claim 7, characterized in that The feature data determination module determines the audio feature data corresponding to the word in the speech data according to the audio feature data in the block memory, including: Determine the correspondence between words and block memory; According to the corresponding relationship, feature data corresponding to the word is determined.
9. The method according to claim 8, characterized in that The feature data determination module includes: a predictor; The feature data determining module determines the feature data corresponding to the word in the voice data according to the audio feature data in the block memory, and further includes: Determining the length of the text included in each block by the predictor; The correspondence between characters and blocks is determined according to the text length.
10. The method according to claim 2, characterized in that The determining of the target model parameters corresponding to the target application includes: Determine the speech recognition performance requirements of the target application; The target model parameters are determined according to the performance requirement information.
11. The method according to claim 2, characterized in that If a first user associated with the target application sends a resource object corresponding to the target model parameters to a second user associated with the model, the speech data is converted into a text sequence through the speech recognition model based on the target model parameters.
12. A speech recognition method, characterized in that: include: Collecting voice data of a target application and sending the voice data to a server, so that the server can learn a voice recognition model with dynamically variable model parameters from a training sample set, including: performing iterative training on the model according to the model parameters; The model parameters are different for each iterative training, and the model parameters include a delay value; for the speech data, the target model parameters corresponding to the target application are determined; and the speech data is converted into a text sequence through the speech recognition model based on the target model parameters.
13. A speech recognition method, characterized in that: include: Learning from a training sample set to obtain a speech recognition model with dynamically variable model parameters includes: performing iterative training on the model according to the model parameters; each time the iterative training has different model parameters, the model parameters include a delay value; Determining target model parameters corresponding to the target application; The speech recognition model based on the target model parameters is sent to a target device running a target application, so that the target application converts the speech data into a text sequence through the speech recognition model based on the target model parameters.
14. The method according to claim 13, characterized in that The determining of the target model parameters corresponding to the target application includes: Determine the speech recognition performance requirements of the target application; The target model parameters are determined according to the performance requirement information.
15. The method according to claim 13, characterized in that The determining of the target model parameters corresponding to the target application includes: Determine the performance information of the device running the target application; The target model parameters are determined according to the device performance information.
16. The method according to claim 15, characterized in that The device performance information includes: computing resource information and storage resource information; Determining the target model parameters according to the device performance information includes: Determining a model size according to the computing resource information; A delay value is determined according to the storage resource information.
17. The method according to claim 13, characterized in that Also includes: Determining resource information corresponding to the target model parameters; Sending the resource information to a first user associated with the target application; If the first user sends the resource object to the second user associated with the model, the speech recognition model based on the target model parameters is sent to the target device.
18. The method according to claim 13, characterized in that Also includes: Determining speech recognition performance information according to the target model parameters; The performance information is sent to a management device associated with the target application, so that the management device displays the performance information.
19. A speech recognition method, characterized in that: include: Send a request to the server to obtain the speech recognition model for the target application; A speech recognition model with dynamically variable model parameters based on target model parameters corresponding to a target application sent back by a receiving server; The model is learned from a training sample set, including: performing iterative training on the model according to model parameters; the model parameters are different in each iterative training, and the model parameters include a delay value; The speech data is converted into a text sequence by the speech recognition model based on the target model parameters.
20. The method according to claim 19, characterized in that Also includes: Determine the speech recognition performance requirements of the target application; The request includes the performance requirement information, so that the server determines the target model parameters according to the performance requirement information.
21. The method according to claim 20, characterized in that Also includes: Receiving device performance requirement information for running the target application determined according to the performance requirement information, sent by the server; The device performance requirement information is displayed to facilitate determining a target device that meets the device performance requirement information, so that the server sends the speech recognition model based on the target model parameters to the target device.
22. The method according to claim 19, characterized in that Also includes: Determine the performance information of the device running the target application; The request includes the device performance information, so that the server can determine the target model parameters according to the device performance information.
23. The method according to claim 19, characterized in that Also includes: Receiving resource information corresponding to the target model parameters sent by the server; The resource object is sent to a second user associated with the model so that the server sends the speech recognition model based on the target model parameters.
24. The method according to claim 19, characterized in that Also includes: Receiving speech recognition performance information corresponding to the target model parameters sent by the server; The speech recognition performance information is displayed.
25. The method according to claim 19, characterized in that Also includes: A test system for receiving a speech recognition model based on multiple sets of model parameters sent by a server; The speech data is converted into a text sequence by using a speech recognition model based on each set of model parameters, so as to determine the speech recognition performance of each set of model parameters; Determine the target model parameters and send them to the server.
26. A method for upgrading a speech recognition service, characterized in that: include: Determining usage status information of a target application for a speech recognition model whose model parameters are dynamically variable based on the first model parameter; The model is learned from a training sample set, including: performing iterative training on the model according to the first model parameter; the first model parameter is different in each iterative training, and the first model parameter is a delay value; Determining a second model parameter of the speech recognition model according to the usage status information, wherein the second model parameter is a delay value; The model parameters of the speech recognition model on the device running the target application are configured as second model parameters, so that the device converts the speech data into a text sequence through the speech recognition model based on the second model parameters.
27. A method for upgrading a speech recognition service, characterized in that: include: Determining usage status information of a target application for a speech recognition model whose model parameters are dynamically variable based on the first model parameter; The model is learned from a training sample set, including: performing iterative training on the model according to the first model parameter; the first model parameter is different in each iterative training, and the first model parameter is a delay value; Determining a second model parameter of the speech recognition model according to the usage status information, wherein the second model parameter is a delay value; The correspondence between the target application and the second model parameters is stored so that for the speech data to be processed by the target application, the speech data is converted into a text sequence according to the correspondence through the speech recognition model based on the second model parameters.
28. A speech recognition service testing method, characterized in that: include: receiving a speech recognition service test request for a target application; For the plurality of sets of model parameters, voice data of a target application is converted into a text sequence by a speech recognition model with dynamically variable model parameters based on each set of model parameters; The model is learned from a training sample set, including: performing iterative training on the model according to the model parameters; the model parameters are different in each iterative training, and the model parameters include a delay value; The text sequences corresponding to each set of model parameters are sent back to the requesting party so that the requesting party can determine the speech recognition performance of each set of model parameters and determine the target model parameters corresponding to the target application based on the performance.
29. A method for constructing a speech recognition model, characterized in that: include: Determine a training data set, wherein the training data includes: speech data and text sequence annotation information; Constructing a network structure of the model; According to the model parameters, iterative training is performed on the model to obtain a speech recognition model with dynamically variable model parameters; each time the iterative training, the model parameters are different, and the model parameters include a delay value.
Citation Information
Patent Citations
Audio corpus screening method and device for speech recognition and computer equipment
CN110263322A
Speech recognition model training method and device
CN110364144A