Voice interaction and voice recognition method, device, equipment and storage medium
By combining the streaming speech recognition model and the offline speech recognition model, the problem of difficulty in improving delay and accuracy in streaming speech recognition is solved, and efficient speech recognition processing is achieved, ensuring real-time and accuracy.
Patent Information
- Application Number
- CN202011112060.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-16
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2040-10-16
AI Technical Summary
The prior art is difficult to ensure both delay and recognition accuracy in streaming speech recognition, and usually requires sacrifice on one hand in exchange for improvement on the other.
Using a method combining the streaming speech recognition model and the offline speech recognition model, the speech signal is encoded and decoded in real time through the streaming speech recognition model to obtain the initial text, and the acoustic features and semantic vectors of the speech signal are spliced and input into the offline speech recognition model to obtain the corrected text.
It realizes that the delay and recognition accuracy are guaranteed in streaming speech recognition at the same time. With the assistance of offline speech recognition model, the recognition results of the streaming speech recognition model are corrected, and the overall performance of speech recognition is improved.
Smart Images

Figure CN114446280B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a voice interaction and voice recognition method, device, equipment and storage medium. Background Art
[0002] Speech recognition technology can convert human speech into machine-recognizable text. At present, there are some end-to-end speech recognition models that are used to provide speech recognition services. In contrast to end-to-end speech recognition models, non-end-to-end speech recognition models include independent acoustic models and language models. The pronunciation sequence can be first recognized by the acoustic model, and then the language model is combined with a preset pronunciation dictionary to determine the text sequence. The end-to-end speech recognition model can adopt a speech recognition framework that combines the acoustic model and the language model into one, and does not require a pronunciation dictionary. In this way, there is no error propagation effect between modules, which can significantly improve speech recognition performance and reduce training complexity.
[0003] From the perspective of application scenarios, at present, end-to-end speech recognition models can be divided into two categories. One category is the model suitable for offline speech recognition, called offline speech recognition model (or offline end-to-end speech recognition model), and the other category is the model suitable for streaming speech or real-time speech recognition, called streaming speech recognition model (or streaming end-to-end speech recognition model). Simply put, the so-called streaming speech recognition refers to the process of performing speech recognition in real time (that is, within a very short delay) with the user's output speech, that is, performing speech recognition while speaking; the so-called offline speech recognition refers to the process of performing speech recognition on the collected user speech after the user's speech output is completed.
[0004] In the process of speech recognition for streaming output speech, on the one hand, there are strong requirements for latency, and on the other hand, there are also high requirements for recognition accuracy. At present, latency and recognition accuracy are often a pair of opposing indicators, that is, recognition accuracy is often sacrificed in exchange for lower latency, or latency is often sacrificed in exchange for higher recognition accuracy. Therefore, how to ensure both latency and recognition accuracy is an urgent problem to be solved. Summary of the invention
[0005] The embodiments of the present invention provide a voice interaction and voice recognition method, apparatus, device and storage medium, which can simultaneously ensure the delay and recognition accuracy of voice recognition.
[0006] In a first aspect, an embodiment of the present invention provides a speech recognition method, the method comprising:
[0007] Encoding the acoustic features of the currently generated speech signal block through a first encoding network in the streaming speech recognition model to sequentially obtain first semantic vectors corresponding to each of the plurality of speech signal blocks, wherein the plurality of speech signal blocks correspond to a continuous speech segment, and each speech signal block has a preset duration;
[0008] Decoding the first semantic vector corresponding to the currently generated speech signal block through the first decoding network in the streaming speech recognition model to sequentially output the first text corresponding to the multiple speech signal blocks;
[0009] Inputting the concatenation result of the acoustic features and the first semantic vector corresponding to each of the plurality of speech signal blocks into an offline speech recognition model, so as to output second texts corresponding to the plurality of speech signal blocks through the offline speech recognition model;
[0010] The first text corresponding to the multiple speech signal blocks output by the streaming speech recognition model is updated according to the second text.
[0011] In a second aspect, an embodiment of the present invention provides a speech recognition device, the device comprising:
[0012] A streaming encoding module, used for encoding the acoustic features of the currently generated speech signal block through a first encoding network in a streaming speech recognition model to sequentially obtain first semantic vectors corresponding to each of the plurality of speech signal blocks, wherein the plurality of speech signal blocks correspond to a continuous speech segment, and each speech signal block has a preset duration;
[0013] A streaming decoding module, used for decoding the first semantic vector corresponding to the currently generated speech signal block through the first decoding network in the streaming speech recognition model, so as to sequentially output the first text corresponding to the plurality of speech signal blocks;
[0014] An offline recognition module, used for inputting the concatenation result of the acoustic features corresponding to each of the plurality of speech signal blocks and the first semantic vector into an offline speech recognition model, so as to output second text corresponding to the plurality of speech signal blocks through the offline speech recognition model;
[0015] An output update module is used to update the first text corresponding to the multiple speech signal blocks output by the streaming speech recognition model according to the second text.
[0016] In a third aspect, an embodiment of the present invention provides an electronic device, comprising: a memory and a processor; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor can at least implement the speech recognition method described in the first aspect.
[0017] In a fourth aspect, an embodiment of the present invention provides a non-temporary machine-readable storage medium having executable code stored thereon. When the executable code is executed by a processor of an electronic device, the processor can at least implement the speech recognition method described in the first aspect.
[0018] In a fifth aspect, an embodiment of the present invention provides a voice interaction method, the method comprising:
[0019] Collecting speech signal blocks in a speech signal stream, each speech signal block having a preset duration;
[0020] The collected speech signal blocks are uploaded to the server, so that the server encodes the acoustic features of the currently generated speech signal blocks through the first encoding network in the streaming speech recognition model, so as to obtain first semantic vectors corresponding to each of the multiple speech signal blocks in sequence, decode the first semantic vectors corresponding to the currently generated speech signal blocks through the first decoding network in the streaming speech recognition model, so as to output first texts corresponding to the multiple speech signal blocks in sequence, and input the splicing results of the acoustic features and the first semantic vectors corresponding to each of the multiple speech signal blocks into the offline speech recognition model, so as to output second texts corresponding to the multiple speech signal blocks through the offline speech recognition model; the multiple speech signal blocks correspond to a continuous speech;
[0021] Displaying the first text received from the server;
[0022] Update the first text according to the second text received from the server.
[0023] In a sixth aspect, an embodiment of the present invention provides a voice interaction device, the device comprising:
[0024] A collection module, used for collecting voice signal blocks in a voice signal stream, each voice signal block having a preset duration;
[0025] A sending module, used for uploading the collected speech signal blocks to a server, so that the server encodes the acoustic features of the currently generated speech signal blocks through a first encoding network in a streaming speech recognition model, so as to sequentially obtain first semantic vectors corresponding to each of the multiple speech signal blocks, decodes the first semantic vectors corresponding to the currently generated speech signal blocks through a first decoding network in the streaming speech recognition model, so as to sequentially output first texts corresponding to the multiple speech signal blocks, and inputs the splicing results of the acoustic features and the first semantic vectors corresponding to each of the multiple speech signal blocks into an offline speech recognition model, so as to output second texts corresponding to the multiple speech signal blocks through the offline speech recognition model; the multiple speech signal blocks correspond to a continuous speech;
[0026] A display module is used to display the first text received from the server; and to update the first text according to the second text received from the server.
[0027] In the seventh aspect, an embodiment of the present invention provides a voice interaction device, comprising: a memory and a processor; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor can at least implement the voice interaction method as described in the fifth aspect.
[0028] In the eighth aspect, an embodiment of the present invention provides a non-temporary machine-readable storage medium, on which executable code is stored. When the executable code is executed by a processor of a voice interaction device, the processor can at least implement the voice interaction method described in the fifth aspect.
[0029] In a ninth aspect, an embodiment of the present invention provides a voice interaction method, the method comprising:
[0030] Acquire voice signal blocks in the conference voice signal stream, each voice signal block having a preset duration;
[0031] Encoding the acoustic features of the currently generated speech signal block through a first encoding network in the streaming speech recognition model to sequentially obtain first semantic vectors corresponding to each of the plurality of speech signal blocks, wherein the plurality of speech signal blocks correspond to a continuous speech segment;
[0032] Decoding the first semantic vector corresponding to the currently generated speech signal block through the first decoding network in the streaming speech recognition model to sequentially output the first text corresponding to the multiple speech signal blocks;
[0033] Inputting the concatenation result of the acoustic features and the first semantic vector corresponding to each of the plurality of speech signal blocks into an offline speech recognition model, so as to output second texts corresponding to the plurality of speech signal blocks through the offline speech recognition model;
[0034] Update the first text corresponding to the plurality of speech signal blocks output by the streaming speech recognition model according to the second text;
[0035] Generate meeting minutes based on the second text.
[0036] In a tenth aspect, an embodiment of the present invention provides a voice interaction device, the device comprising:
[0037] An acquisition module, used to acquire voice signal blocks in the conference voice signal stream, each voice signal block having a preset duration;
[0038] A streaming encoding and decoding module, used for encoding the acoustic features of the currently generated speech signal block through a first encoding network in a streaming speech recognition model, so as to sequentially obtain first semantic vectors corresponding to each of a plurality of speech signal blocks, wherein the plurality of speech signal blocks correspond to a continuous speech; decoding the first semantic vector corresponding to the currently generated speech signal block through a first decoding network in the streaming speech recognition model, so as to sequentially output first texts corresponding to the plurality of speech signal blocks;
[0039] An offline processing module, configured to input the concatenation result of the acoustic features and the first semantic vector corresponding to each of the plurality of speech signal blocks into an offline speech recognition model, so as to output second characters corresponding to the plurality of speech signal blocks through the offline speech recognition model; and to update the first characters corresponding to the plurality of speech signal blocks output by the streaming speech recognition model according to the second characters;
[0040] A generating module is used to generate a meeting record according to the second text.
[0041] In the eleventh aspect, an embodiment of the present invention provides a voice interaction device, comprising: a memory and a processor; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor can at least implement the voice interaction method as described in the ninth aspect.
[0042] In the twelfth aspect, an embodiment of the present invention provides a non-temporary machine-readable storage medium, on which executable code is stored. When the executable code is executed by a processor of a voice interaction device, the processor can at least implement the voice interaction method described in the ninth aspect.
[0043] In a thirteenth aspect, an embodiment of the present invention provides a voice interaction method, the method comprising:
[0044] Acquire voice signal blocks in the host's voice signal stream, each voice signal block having a preset duration;
[0045] Encoding the acoustic features of the currently generated speech signal block through a first encoding network in the streaming speech recognition model to sequentially obtain first semantic vectors corresponding to each of the plurality of speech signal blocks, wherein the plurality of speech signal blocks correspond to a continuous speech segment;
[0046] Decoding the first semantic vector corresponding to the currently generated speech signal block through the first decoding network in the streaming speech recognition model to sequentially output the first text corresponding to the multiple speech signal blocks;
[0047] Displaying the first text in the live broadcast interface;
[0048] Inputting the concatenation result of the acoustic features and the first semantic vector corresponding to each of the plurality of speech signal blocks into an offline speech recognition model, so as to output second texts corresponding to the plurality of speech signal blocks through the offline speech recognition model;
[0049] The first text is replaced with the second text in the live broadcast interface.
[0050] In a fourteenth aspect, an embodiment of the present invention provides a voice interaction device, the device comprising:
[0051] An acquisition module, used to acquire voice signal blocks in the host's voice signal stream, each voice signal block having a preset duration;
[0052] A streaming encoding and decoding module, used for encoding the acoustic features of the currently generated speech signal block through a first encoding network in a streaming speech recognition model, so as to sequentially obtain first semantic vectors corresponding to each of a plurality of speech signal blocks, wherein the plurality of speech signal blocks correspond to a continuous speech; decoding the first semantic vector corresponding to the currently generated speech signal block through a first decoding network in the streaming speech recognition model, so as to sequentially output first texts corresponding to the plurality of speech signal blocks;
[0053] A display module, used for displaying the first text in a live broadcast interface;
[0054] An offline processing module, used for inputting the concatenation result of the acoustic features corresponding to each of the plurality of speech signal blocks and the first semantic vector into an offline speech recognition model, so as to output second text corresponding to the plurality of speech signal blocks through the offline speech recognition model;
[0055] The display module is also used to replace the first text with the second text in the live broadcast interface.
[0056] In the fifteenth aspect, an embodiment of the present invention provides a voice interaction device, comprising: a memory and a processor; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor can at least implement the voice interaction method as described in the thirteenth aspect.
[0057] In the sixteenth aspect, an embodiment of the present invention provides a non-temporary machine-readable storage medium having executable code stored thereon. When the executable code is executed by a processor of a voice interaction device, the processor can at least implement the voice interaction method as described in the thirteenth aspect.
[0058] In an embodiment of the present invention, a solution is provided for performing speech recognition processing on a speech signal stream by combining a streaming speech recognition model and an offline speech recognition model, so as to obtain good delay performance and recognition accuracy.
[0059] Specifically, for the voice signal stream, the voice signal stream is cut into voice signal blocks with a preset time length as the delay value, that is, whenever a voice signal block in the voice signal stream is received, the acoustic features of the voice signal block are input into the encoding network (referred to as the first encoding network) in the streaming speech recognition model for encoding to obtain the semantic vector (referred to as the first semantic vector) corresponding to the voice signal block. Then, the first semantic vector corresponding to the voice signal block is input into the decoding network (referred to as the first decoding network) in the streaming speech recognition model to decode and output the text contained in the voice signal block. In this way, real-time speech recognition processing can be achieved based on the streaming speech recognition model, that is, the text corresponding to the currently generated voice signal block can be output in real time and in sequence.
[0060] In addition, assuming that it is determined based on voice activity detection (Voice Activity Detection, referred to as VAD) that the multiple voice signal blocks generated continuously correspond to a continuous speech, after the text corresponding to the last voice signal block in the multiple voice signal blocks is decoded and output by the above-mentioned first decoding network, the acoustic features and the first semantic vector corresponding to each of the multiple voice signal blocks can be obtained, and for each voice signal block, the corresponding acoustic features and the first semantic vector are spliced together to obtain the splicing result corresponding to the voice signal block. Afterwards, the splicing results corresponding to each of the multiple voice signal blocks are input into the offline speech recognition model, so that for each voice signal block, the offline speech recognition model can obtain context information within a longer distance range, so that the text sequence corresponding to the multiple voice signal blocks output by the offline speech recognition model is more accurate. Finally, the text sequence output by the offline speech recognition model replaces the output text of the streaming speech recognition model to achieve the correction of the output result of the streaming speech recognition model. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0062] Figure 1 A schematic diagram of the composition of a speech recognition system provided by an embodiment of the present invention;
[0063] Figure 2 A flow chart of a speech recognition method provided by an embodiment of the present invention;
[0064] Figure 3a-3d A schematic diagram of the working process of a speech recognition system provided by an embodiment of the present invention;
[0065] Figure 4 A flowchart of another speech recognition method provided by an embodiment of the present invention;
[0066] Figure 5 A schematic diagram of another speech recognition system provided by an embodiment of the present invention;
[0067] Figure 6 A schematic diagram of a scenario of a voice interaction method provided by an embodiment of the present invention;
[0068] Figure 7 A schematic diagram of a scenario of a voice interaction method provided by an embodiment of the present invention;
[0069] Figure 8 A schematic diagram of a scenario of a voice interaction method provided by an embodiment of the present invention;
[0070] Fig. 9 A schematic diagram of the structure of a speech recognition device provided by an embodiment of the present invention;
[0071] Fig.10 For Fig. 9 A schematic diagram of the structure of an electronic device corresponding to the speech recognition device provided in the illustrated embodiment;
[0072] Fig.11 A schematic diagram of the structure of a voice interaction device provided by an embodiment of the present invention;
[0073] Fig.12 For Fig.11 A schematic diagram of the structure of a voice interaction device corresponding to the voice interaction apparatus provided in the illustrated embodiment;
[0074] Fig.13 A schematic diagram of the structure of a voice interaction device provided by an embodiment of the present invention;
[0075] Fig.14 For Fig.13 A schematic diagram of the structure of a voice interaction device corresponding to the voice interaction apparatus provided in the illustrated embodiment;
[0076] Fig.15 A schematic diagram of the structure of a voice interaction device provided by an embodiment of the present invention;
[0077] Fig.16 For Fig.15 A schematic diagram of the structure of a voice interaction device corresponding to the voice interaction apparatus provided in the illustrated embodiment. DETAILED DESCRIPTION
[0078] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0079] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings, and "multiple" generally includes at least two.
[0080] As used herein, the words "if" and "if" may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0081] In addition, the step sequence in the following method embodiments is only an example and not a strict limitation.
[0082] The speech recognition method provided in the embodiment of the present invention can be executed by an electronic device, which can be a terminal device such as a PC, a laptop, a smart phone, an intelligent robot, or a server in the cloud. The server can be a physical server including an independent host, or a virtual server, or a cloud server.
[0083] The speech recognition method provided by the embodiment of the present invention can be applied to speech recognition processing on a speech signal stream. In simple terms, it can be applied to speech recognition processing on a speech signal stream generated in real time.
[0084] When real-time speech recognition processing is required on the speech signal stream in certain application scenarios, a streaming speech recognition model (streaming end-to-end speech recognition model) is needed. Although the streaming speech recognition model can achieve good latency performance, that is, it can ensure that there is only a short latency between the output of the text sequence and the input speech signal stream, the streaming speech recognition model can only learn contextual information within a limited distance range, resulting in the need to improve the recognition accuracy.
[0085] However, the offline speech recognition model can learn the complete context information of a continuous speech, which can effectively ensure the recognition accuracy. Based on this, in an embodiment of the present invention, a speech recognition system combining a streaming speech recognition model and an offline speech recognition model is provided, so that the speech recognition system can be used to complete the real-time and accurate speech recognition processing of the speech signal generated by the streaming.
[0086] First, let's briefly introduce the composition of the above speech recognition system. Figure 1 As shown, the speech recognition system includes a streaming speech recognition model and an offline speech recognition model. The streaming speech recognition model includes a first encoding network and a first decoding network. The offline speech recognition model includes a second encoding network and a second decoding network.
[0087] Optionally, the first encoding network, the second encoding network, the first decoding network, and the second decoding network can all be implemented using multi-layer neural networks. For example, the structure of the first encoding network can be any of the following: Feedforward Sequential Memory Networks (FSMN), Deep Feedforward Sequential Memory Networks (DFSMN), Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM), Bi-directional Long Short-Term Memory (BiLSTM), etc. The structure of the second encoding network can be BiLSTM, CNN, DFSMN, etc.
[0088] exist Figure 1 In the composition of the speech recognition system shown, the combination of the streaming speech recognition model and the offline speech recognition model is mainly reflected in: the output of the first encoding network in the streaming speech recognition model is used as the input of the second encoding network in the offline speech recognition model.
[0089] The working process of the above-mentioned speech recognition system during speech recognition will be described in detail below in conjunction with the following embodiments.
[0090] Figure 2 A flow chart of a speech recognition method provided by an embodiment of the present invention is as follows: Figure 2 As shown, the method comprises the following steps:
[0091] 201. Encode the acoustic features of the currently generated speech signal block through the first encoding network in the streaming speech recognition model to obtain first semantic vectors corresponding to each of the multiple speech signal blocks in sequence, wherein the multiple speech signal blocks correspond to a continuous speech, and each speech signal block has a preset duration.
[0092] 202. Decode the first semantic vector corresponding to the currently generated speech signal block through the first decoding network in the streaming speech recognition model to output the first text corresponding to the multiple speech signal blocks in sequence.
[0093] 203. Input the concatenation result of the acoustic features and the first semantic vector corresponding to each of the multiple speech signal blocks into an offline speech recognition model, so as to output second text corresponding to the multiple speech signal blocks through the offline speech recognition model.
[0094] 204. Update the first text output by the streaming speech recognition model and corresponding to the multiple speech signal blocks according to the second text output by the offline speech recognition model.
[0095] In the embodiment of the present invention, in order to realize real-time speech recognition processing on a speech signal stream, the following concept is defined: speech signal segmentation.
[0096] Among them, the speech signal block is defined by a preset duration. For example, if the preset duration is 100 milliseconds (ms), then from the moment the speech signal stream starts to be output, a speech signal block is cut out every 100ms, and each time a speech signal block is cut out, the currently generated speech signal block is immediately subjected to speech recognition processing through the streaming speech recognition model. Therefore, the preset duration can actually reflect the speech recognition delay of the speech signal stream. In the above example, the latency is 100ms, that is, speech recognition processing is performed every 100ms. In this way, the user can perceive that the text output by the speech recognition is only 100ms later than the actual speaking, and there is no need to wait for the user to finish the whole sentence before performing speech recognition.
[0097] Different speech recognition application scenarios have different requirements for the delay value of speech recognition processing. Therefore, in practical applications, the delay value corresponding to the speech recognition application scenario that currently triggers the speech recognition requirement can be determined so that the delay value can be used to divide the speech signal stream into blocks.
[0098] It can be seen that a streaming speech recognition model that can be applied to different delay values can be pre-trained, so that the streaming speech recognition model can be applied to different speech recognition application scenarios. Simply put, during the training process, a variety of delay values can be pre-set, such as 100ms, 150ms, 300ms, and so on. During the iterative training process of the streaming speech recognition model, the delay value used in each iteration process can be arbitrarily selected from these preset delay values, that is, the speech signal used as a training sample is cut into speech signal blocks of corresponding lengths according to the selected delay value, and then the various speech signal blocks obtained by cutting are respectively input into the streaming speech recognition model for training.
[0099] In actual applications, a correspondence between different speech recognition application scenarios and delay values can be pre-established, so as to determine the delay value corresponding to the current speech recognition application scenario based on the correspondence. Different speech recognition application scenarios can be embodied as different application programs.
[0100] For the current speech recognition application scenarios (such as the scenario where the user is giving a speech), assuming that the delay value is 100ms, and assuming that the user starts to output the speech signal from time T0, then from time T0, every time the timing reaches 100ms, the speech signal generated within this 100ms is intercepted as a speech signal block, and the speech signal block is input into the streaming speech recognition model for real-time speech recognition to output the text corresponding to the speech signal block.
[0101] For ease of description and understanding, combined Figure 3a-3d To illustrate the process of speech recognition for streaming speech, Figure 1 The working process of the streaming speech recognition model and the offline speech recognition model in the speech recognition system is shown.
[0102] Assume that starting from the initial time T0, the speech signal blocks that can be obtained in sequence are speech signal block 1, speech signal block 2, ... speech signal block N. And assume that based on VAD, it is determined that speech signal block 1 to speech signal block N are a continuous speech, such as a whole sentence spoken by the user. Figure 3a As shown in , assuming that the currently generated speech signal block is speech signal block 1, by extracting the acoustic feature X1 of speech signal block 1, the acoustic feature X1 of speech signal block 1 is input into the first encoding network in the streaming speech recognition model, so as to encode the acoustic feature of speech signal block 1 through the first encoding network, and obtain the first semantic vector corresponding to speech signal block 1, Figure 3a In the figure, the first semantic vector corresponding to speech signal block 1 is represented as C1. After that, the semantic vector C1 is input into the first decoding network, and the text contained in speech signal block 1 is output through decoding by the first decoding network. Assume that speech signal block 1 contains two texts, which are represented as y1 and y2 respectively.
[0103] Afterwards, when the second speech signal block, namely speech signal block 2, is obtained, similarly, the acoustic feature X2 of speech signal block 2 is extracted. At this time, the acoustic features of speech signal block 2 and speech signal block 1 can be spliced, and the spliced acoustic features are input into the first encoding network to output the first semantic vector corresponding to speech signal block 2 through the first encoding network, represented as C2. The semantic vector C2 is input into the first decoding network to output the text contained in speech signal block 2 through the first decoding network, represented as y3.
[0104] Similarly, the above-mentioned real-time speech recognition processing is performed on the subsequent speech signal blocks obtained one by one through the streaming speech recognition model.
[0105] Among them, the acoustic feature can be any of the following: Mel-Frequency Cepstral Coefficients (MFCC), Linear Predictive Cepstral Coefficients (LPCC), short-time average energy, average amplitude change rate, Fbank features, etc.
[0106] The working process of the streaming speech recognition model is summarized as follows: when receiving the currently generated first speech signal block, obtain at least one second speech signal block generated before the first speech signal block (wherein, the at least one second speech signal block is a speech signal block in the same continuous speech as the first speech signal block); obtain the acoustic features of the first speech signal block, and the acoustic features of the at least one second speech signal block; encode the splicing result of the acoustic features of the first speech signal block and the acoustic features of the at least one second speech signal block through the first encoding network to obtain the first semantic vector corresponding to the first speech signal block. Afterwards, the first semantic vector corresponding to the first speech signal block is input into the first decoding network to output the text contained therein through the first decoding network.
[0107] For any speech signal block, the acoustic features of the speech signal block can be obtained in the following manner:
[0108] Performing frame processing on any speech signal block to obtain multiple frames of speech signals;
[0109] Extracting acoustic features corresponding to each of the multiple frames of speech signals;
[0110] The acoustic features of any speech signal block are determined according to the acoustic features corresponding to each of the multiple frames of speech signals.
[0111] The duration of each frame of speech signal can be preset, for example, 10ms. Assuming that the duration of a speech signal block is 100ms, this means that a speech signal block can be divided into 10 frames of speech signals. After obtaining the acoustic features corresponding to each of the multiple frames of speech signals, the obtained acoustic features are spliced together as the acoustic features of the corresponding speech signal block.
[0112] In an alternative embodiment, if Figure 3c As shown, the streaming speech recognition model may also include a prediction network and an attention network, and the prediction network, the attention network and the first decoding network cooperate with each other to complete the decoding process of the semantic vector output by the first encoding network.
[0113] At this time, the first semantic vector corresponding to the currently generated speech signal block is decoded by the first decoding network, which can be specifically implemented as follows:
[0114] For the first speech signal block, predicting the first semantic vector corresponding to the first speech signal block through a prediction network to obtain the number of characters contained in the first speech signal block;
[0115] The weights of the first semantic vectors corresponding to the first speech signal block and the at least one second speech signal block and the weighted summation result are determined by the attention network in each weight calculation process, and the weighted summation result is input into the first decoding network to output the text corresponding to the first speech signal block through the first decoding network. The number of characters is used to constrain the number of weight calculations, that is, the number of characters is used to constrain the number of weight calculations of the first semantic vectors corresponding to the first speech signal block and the at least one second speech signal block by the attention network.
[0116] Among them, the above-mentioned at least one second speech signal block refers to the speech signal blocks generated before the first speech signal block and belonging to the same continuous speech segment as the first speech signal block, and the semantic vectors corresponding to each speech signal block are obtained through the first encoding network.
[0117] In the embodiment of the present invention, the prediction network and the attention network are similar to the first decoding network and can also be a structure composed of multi-layer neural networks.
[0118] The prediction network is used to determine the number of words included in the speech signal block so as to control the operation of the attention network based on the guidance of the number of words.
[0119] Specifically, after obtaining the first semantic vector corresponding to the first speech signal block, the first semantic vector is input into the prediction network, and the prediction network outputs a prediction sequence corresponding to the first semantic vector. The prediction sequence is used to indicate whether each of the multiple frames of speech signals contained in the first speech signal block corresponds to a character, so that the number of characters contained in the first speech signal block can be determined according to the number of speech signal frames corresponding to the characters in the prediction sequence.
[0120] In simple terms, the prediction network can be trained in the following way: obtain a number of speech signal block samples, annotate the number of characters contained in each speech signal block sample, and use this annotated information to train the prediction network. The trained prediction network has the function of predicting the number of characters contained in the speech signal block.
[0121] In summary, the guiding role of the number of words output by the above-mentioned prediction network on the attention network is mainly reflected in: constraining the number of times the attention network calculates the weights of the first semantic vectors corresponding to the first speech signal block and the at least one second speech signal block.
[0122] Among them, since the second speech signal block is generated before the first speech signal block, after the acoustic features corresponding to the second speech signal block are calculated and the corresponding first semantic vector is calculated based on the acoustic features, the acoustic features and first semantic vector corresponding to the second speech signal block can be stored for subsequent use.
[0123] Under the guidance of the number of characters N (N is greater than or equal to 1) contained in the first speech signal block predicted by the prediction network, the specific working process of the attention network is as follows:
[0124] Initialize the weight calculation times to the number of characters N, and iterate the following process until the weight calculation times is reduced to 0:
[0125] If the number of weight calculations is not 0, the first semantic vectors corresponding to the first speech signal block and the at least one second speech signal block and the last word output by the first decoding network are input into the attention network, so that the attention network determines the weights of the first semantic vectors corresponding to the first speech signal block and the at least one second speech signal block based on the last word, and determines the weighted summation result of the first semantic vectors corresponding to the first speech signal block and the at least one second speech signal block according to the calculated weights, and inputs the weighted summation result into the first decoding network;
[0126] Obtain the currently output text through the first decoding network and input the text into the attention network;
[0127] Subtract one from the number of weight calculations.
[0128] Among them, the process of calculating the above weights by the attention network mainly adopts some existing attention mechanisms (attention), which will not be elaborated.
[0129] Combine the following Figure 3c To illustrate the above decoding process, Figure 3c In the example, it is assumed that the first speech signal is divided into blocks Figure 3aThe speech signal block m shown in the figure includes 10 frames of speech signals. After the encoding process of the first encoding network, the semantic vector Cm is obtained, and the semantic vector Cm is input into the prediction network. The prediction sequence output by the prediction network is a sequence of 0 and 1 with a total length of 10, wherein 0 indicates that the corresponding frame of speech signal does not contain text, and 1 indicates that the corresponding frame of speech signal contains text. For example, if the output prediction sequence is [0, 0, 0, 1, 1, 0, 0, 0, 0, 0], it means that there are texts in the fourth and fifth frames of speech signals among the 10 frames of speech signals, that is, the first speech signal block includes two characters. This means that the attention network needs to perform weight calculations on the semantic vector Cm twice.
[0130] It can be understood that since the speech signal block m is the mth speech signal block generated by streaming, at this time, the acquisition result of the second speech signal block in the above text is speech signal block 1 to speech signal block m-1m. In addition, in the initial case, it is assumed that the output result of the first decoding network is the default character, represented by y0.
[0131] Based on the above assumptions, Figure 3c In the example, since the prediction network predicts that the speech signal block m includes two characters, the weight calculation times N=2 are initialized. In the first weight calculation process, the input of the attention network is the semantic vector Cm and the previous character output by the first decoding network: Y t-1 , through Y t-1 By performing correlation calculation with the semantic vector Cm, the weights corresponding to each element in the semantic vector Cm can be obtained. Among them, each element in the semantic vector Cm corresponds to the encoding result corresponding to each of the 10 frames of speech signals contained in the speech signal block m. Based on the calculated weights, each element in the semantic vector Cm is weighted and summed, and the weighted summation result is input into the first decoding network. At this time, the first decoding network will output the first text contained in the speech signal block m: Y t .Y t Input to the attention network. At this time, the weight calculation times N=1 is updated, and the attention network is based on Y t The weights corresponding to the elements in the semantic vector Cm are calculated again, and the weighted sum of the elements in the semantic vector Cm is calculated according to the weights determined again. The weighted sum result is then input into the first decoding network, and the first decoding network outputs the second word contained in the speech signal block m at this time: Y t+1 At this time, the update weight calculation times N = 0. Since the weight calculation times have been updated to 0, the attention network stops working and waits for the arrival of the semantic vector corresponding to the next speech signal block m+1.
[0132] In summary, in an embodiment of the present invention, the prediction network outputs the number of characters contained in a certain speech signal block, which can be used to guide the attention network to pay attention to which previous semantic vectors in the process of decoding the current character. At the same time, the historical character (previous character) finally output by the first decoding network will also be used as the input of the attention network to predict the current character.
[0133] Specifically, it is assumed that the current first speech signal block is speech signal block 10, and the second speech signal block is speech signal block 1 to speech signal block 9 generated previously, and it is assumed that the semantic vector corresponding to speech signal block 10 is represented as C10, and the semantic vectors corresponding to speech signal block 1 to speech signal block 9 are represented as C1 to C9 respectively. It is assumed that the prediction network determines that speech signal block 10 includes two characters. Since the prediction result input by the prediction network to the attention network is the prediction result corresponding to speech signal block 10, the attention network needs to pay attention to the semantic vectors of each speech signal block in the same continuous speech segment that has been generated before speech signal block 10: C1 to C9. In addition, it is assumed that the prediction network determines that the corresponding speech signal block does not contain characters when predicting C3. Then, optionally, at this time, the attention network can only pay attention to the semantic vectors corresponding to speech signal block 1, speech signal block 2, and speech signal block 4 to speech signal block 9, without paying attention to the semantic vector corresponding to speech signal block 3. That is to say, optionally, the attention network can only focus on the semantic vectors corresponding to each speech signal block containing text.
[0134] Depend on Figure 3a to Figure 3c From the working process of the illustrated streaming speech recognition model, it can be seen that the streaming speech recognition model predicts the text contained in the current speech signal block by combining the historical speech signal blocks generated before the currently generated speech signal block each time a speech signal block is generated, thereby ensuring real-time or short-delay speech recognition effect. However, for the currently generated speech signal block, the visible future information is very limited, which is not conducive to the recognition accuracy of the text contained in the speech signal block.
[0135] The so-called limited visible future information means that, assuming that the currently generated speech signal block is the speech signal block 2 in the above example, and assuming that the speech signal block 2 includes 10 frames of speech signals, then for the first frame of speech signal, only the subsequent nine frames of speech signals are visible. Similarly, for the ninth frame of speech signal, only the subsequent tenth frame of speech signal is visible.
[0136] Therefore, in order to ensure the short latency performance of speech recognition and the accuracy of speech recognition results, the streaming speech recognition model is combined with the assistance of the offline speech recognition model to obtain a good recognition accuracy. The auxiliary role of the offline speech recognition model is mainly due to the fact that the offline speech recognition model can learn the complete context information of a continuous speech. In other words, for a speech signal block, the offline speech recognition model can see the future information of a longer distance.
[0137] Combine the following Figure 3d The process of combining the offline speech recognition model with the streaming speech recognition model is illustrated by way of example.
[0138] exist Figure 3d In the example, it is assumed that a continuous speech segment defined based on VAD (such as a whole sentence spoken by the user) includes speech signal block 1 to speech signal block 10, and it is assumed that the acoustic features of these 10 speech signal blocks are represented as: X1 to X10. It is assumed that the first semantic vectors corresponding to speech signal block 1 to speech signal block 10 are represented as: C1 to C10. In addition, it is assumed that the first text outputted by the first decoding network in sequence is represented as [Y1, Y2, ..., Y15].
[0139] Based on the above assumptions, Figure 3d In the embodiment, after obtaining the text corresponding to the speech signal block 10 through the streaming speech recognition model, the acoustic features and semantic vectors corresponding to the above 10 speech signal blocks can be spliced together accordingly, so that the spliced feature vectors corresponding to the 10 speech signal blocks can be obtained, and the spliced feature vectors corresponding to the 10 speech signal blocks can be input into the second encoding network of the offline speech recognition model for encoding.
[0140] Among them, correspondingly splicing together the acoustic features and semantic vectors corresponding to the above 10 speech signal blocks means: taking speech signal block 1 as an example, splicing its acoustic feature X1 and semantic vector C1 together, and the splicing result is called the spliced feature vector. The same is true for other speech signal blocks. It can be understood that the acoustic features are also expressed in the form of vectors.
[0141] The output result of the second encoding network is input to the second decoding network in the offline speech recognition model, so that the second decoding network can output the second text corresponding to the above 10 speech signal blocks, assuming that it is [W1, W2, ... W18]. Finally, the first text output by the previous streaming speech recognition model can be replaced by the second text to correct the recognition result of the streaming speech recognition model.
[0142] As can be seen from the above example, taking the voice signal block 5 of the above 10 voice signal blocks as an example, assuming that each voice signal block includes 10 frames of voice signals, taking the first frame of voice signal in the voice signal block 5 as an example, in the streaming voice recognition model, the visible future information of the voice signal frame is only the future nine frames of voice signals in the same voice signal block, but in the offline voice recognition model, the visible future information of the voice signal frame includes the future nine frames of voice signals in the same voice signal block and the voice signals of all frames contained in the 5 voice signal blocks generated subsequently. Due to the existence of richer context information, the recognition accuracy can be well guaranteed.
[0143] In summary, combining the advantages of the streaming speech recognition model in terms of latency with the advantages of the offline speech recognition model in terms of recognition accuracy can achieve better recognition results for speech generated by streaming.
[0144] In the above embodiments, the combination of the streaming speech recognition model and the offline speech recognition model is mainly reflected in that after the semantic vectors corresponding to multiple speech signal blocks (corresponding to a continuous speech) are obtained in sequence through the streaming speech recognition model, the semantic vectors of the multiple speech signal blocks are used as an input of the second encoding network in the offline speech recognition model to be spliced corresponding to the acoustic features of the multiple speech signal blocks input into the second encoding network.
[0145] In addition, to obtain a better recognition effect, optionally, the texts corresponding to the multiple speech signal blocks sequentially output by the streaming speech recognition model can also be used as an input of the second decoding network in the offline speech recognition model. This situation is described in conjunction with the following embodiments.
[0146] Figure 4 A flowchart of another speech recognition method provided by an embodiment of the present invention is shown in FIG. Figure 4 As shown, the method comprises the following steps:
[0147] 401. Encode the acoustic features of the currently generated speech signal block through a first encoding network in a streaming speech recognition model to sequentially obtain first semantic vectors corresponding to each of the multiple speech signal blocks, wherein the multiple speech signal blocks correspond to a continuous speech segment, and each speech signal block has a preset duration.
[0148] 402. Decode the first semantic vector corresponding to the currently generated speech signal block through the first decoding network in the streaming speech recognition model to sequentially output first texts corresponding to the plurality of speech signal blocks.
[0149] 403. Determine a character sequence consisting of first characters output in sequence by the first decoding network.
[0150] 404. Input the concatenation result of the acoustic features corresponding to each of the multiple speech signal blocks and the first semantic vector into the second encoding network in the offline speech recognition model, so as to output the second semantic vector corresponding to each of the multiple speech signal blocks through the second encoding network.
[0151] 405. Input the text sequence into a third encoding network in an offline speech recognition model to output a third semantic vector corresponding to the text sequence through the third encoding network.
[0152] 406. Input the second semantic vectors and the third semantic vector corresponding to each of the multiple speech signal blocks into a second decoding network in the offline speech recognition model, so as to output second text corresponding to the multiple speech signal blocks through the second decoding network.
[0153] 407. Update the first text with the second text.
[0154] In this embodiment, after the characters corresponding to the multiple speech signal blocks are obtained in sequence through the first decoding network in the streaming speech recognition model, a character sequence composed of the characters sequentially output by the first decoding network can be determined. After that, the concatenation results of the acoustic features and the first semantic vector corresponding to each of the multiple speech signal blocks and the character sequence are input into the offline speech recognition model, so that the offline speech recognition model outputs the second characters corresponding to the multiple speech signal blocks.
[0155] For ease of description and understanding, combined Figure 5 To illustrate the composition and working process of the speech recognition system.
[0156] like Figure 5 As shown in , different from the aforementioned embodiment, in addition to the second encoding network and the second decoding network described in the aforementioned embodiment, the offline speech recognition model can also include a third encoding network (also called a text encoder), which can be implemented as a structure composed of one or more layers of neural networks.
[0157] exist Figure 5 In Figure 3d For example, assume that a continuous speech segment defined based on VAD (such as a whole sentence spoken by the user) includes speech signal block 1 to speech signal block 10, and assume that the acoustic features of these 10 speech signal blocks are represented as: X1~X10. Assume that the semantic vectors corresponding to speech signal 1 to speech signal block 10 are represented as: C1~C10. In addition, assume that the text sequence composed of the first text outputted by the first decoding network is represented as [Y1, Y2, ····Y15].
[0158] Based on the above assumptions, Figure 5In the embodiment, after obtaining the text corresponding to the speech signal block 10 through the streaming speech recognition model, the acoustic features and semantic vectors corresponding to the above 10 speech signal blocks can be correspondingly spliced together, so that the spliced feature vectors corresponding to the 10 speech signal blocks can be obtained. The second encoding network of the offline speech recognition model encodes the spliced feature vectors corresponding to the 10 speech signal blocks to obtain the second semantic vector corresponding to each speech signal block.
[0159] The text sequence [Y1, Y2, ..···Y15] is input into the third encoding network in the offline speech recognition model to obtain the corresponding third semantic vector.
[0160] The above-mentioned second semantic vector and third semantic vector are input into the second decoding network in the offline speech recognition model, so that the second decoding network can more accurately recognize and output the second text [W1, W2, ..W18] with the assistance of the third semantic vector, and replace the first text with the second text to realize the correction of the recognition result of the streaming speech recognition model.
[0161] The speech recognition method provided in the above embodiment of the present invention can be applied to an architecture composed of a client and a server, wherein the client runs in a terminal device on the user side, and the server can be composed of a server or a server cluster in the cloud, and the server provides a speech recognition service. In actual applications, the client will change accordingly depending on the application scenario. For example, in a conference application scenario, the client can be a client that provides conference-related functions; in a live broadcast scenario, the client can be a host client.
[0162] Under the architecture of the client and the server, an embodiment of the present invention provides a voice interaction method, which can be executed by the client. The voice interaction method may include the following steps:
[0163] Collecting speech signal blocks in a speech signal stream, each speech signal block having a preset duration;
[0164] The collected speech signal blocks are uploaded to the server, so that the server encodes the acoustic features of the currently generated speech signal blocks through the first encoding network in the streaming speech recognition model, so as to obtain first semantic vectors corresponding to each of the multiple speech signal blocks in sequence, decode the first semantic vectors corresponding to the currently generated speech signal blocks through the first decoding network in the streaming speech recognition model, so as to output first texts corresponding to the multiple speech signal blocks in sequence, and input the splicing results of the acoustic features and the first semantic vectors corresponding to each of the multiple speech signal blocks into the offline speech recognition model, so as to output second texts corresponding to the multiple speech signal blocks through the offline speech recognition model; the multiple speech signal blocks correspond to a continuous speech;
[0165] Display the first text received from the server;
[0166] The first text is updated according to the second text received from the server.
[0167] Combine the following Figure 6 To exemplify the execution process of the voice interaction method provided by the embodiment of the present invention under the above-mentioned client and server architecture.
[0168] exist Figure 6 In the embodiment, it is assumed that user A is using a certain client. During the use of the client, user A outputs a voice signal stream. The client collects voice signal blocks in the voice signal stream output by user A, and uploads each voice signal block collected in sequence to the server in real time, wherein each voice signal block has a preset duration.
[0169] For ease of description, the following assumptions are made here: Assume that every time user A finishes a sentence, the offline speech recognition model will be triggered to perform speech recognition processing on the sentence that has been spoken, and while user A is speaking the sentence, every time a speech signal block is collected, the streaming speech recognition model will be triggered to perform speech recognition processing on the speech signal block. In addition, assume that a sentence spoken by user A includes m speech signal blocks V1 to Vm.
[0170] Based on the above assumptions, if Figure 6 As shown in , the client uploads each collected speech signal block to the server, and Vi represents any collected speech signal block, where Vi is one of the m speech signal blocks. The server maintains two trained models, namely, a streaming speech recognition model and an offline speech recognition model. The streaming speech recognition model includes a first encoding network and a first decoding network. Figure 6 They are represented as: streaming encoder and streaming decoder respectively.
[0171] When the server receives the speech signal block Vi, it extracts its corresponding acoustic features Xi, and inputs the extracted acoustic features Xi into the streaming encoder to obtain the corresponding first semantic vector Ci. It is worth noting that in actual applications, if the speech signal block Vi is not the first speech signal block, the acoustic features corresponding to the speech signal block Vi and the previously collected speech signal blocks can be input into the streaming encoder to obtain the first semantic vector corresponding to the speech signal block Vi. In this way, the first semantic vector will contain the above information of the speech signal block Vi.
[0172] After obtaining the first semantic vector Ci corresponding to the voice signal block Vi, the server inputs the first semantic vector Ci into the streaming decoder to decode and output the text included in the voice signal block Vi, assumed to be "thin waist".
[0173] In Figure 6 , assume that the output texts obtained by the above processing process for the two voice signal blocks before the voice signal block Vi and the three voice signal blocks after it are respectively: "It's an honor", "to", "explain", "blockchain", "to everyone". Thus, it can be seen that the streaming speech recognition model can output the text corresponding to each voice signal block in real time, which is called the first text.
[0174] In addition, in the above example of the output results of the first texts corresponding to multiple voice signal blocks, assume that due to limited context information, the speech recognition result of the voice signal block Vi by the streaming speech recognition model has a misspelled word: "thin waist". The correct recognition result should be "invited".
[0175] To correct the above error, after the server outputs the text corresponding to the voice signal block Vm through the streaming speech recognition model, it can input the acoustic features and the first semantic vector corresponding to each voice signal block in these m voice signal blocks V1 to Vm into the offline speech recognition model. The acoustic features corresponding to these m voice signal blocks are represented as X1 to Xm, and the first semantic vectors corresponding to these m voice signal blocks are represented as C1 to Cm. Specifically, in the offline speech recognition model, feature vector encoding and decoding processing will be performed based on the concatenation result of the acoustic features and the first semantic vector of each voice signal block. In this way, as Figure 6 shown, the offline speech recognition model will finally output the text "It's an honor to be invited to explain blockchain to everyone", and replace the output result of the streaming speech recognition model with this text, so as to correct the output result of the streaming speech recognition model.
[0176] Through the above solution, the real-time output of text can be ensured by the streaming speech recognition model, and the accuracy of the final text output result can be ensured by the offline speech recognition model.
[0177] As described above, the server obtains the second text according to the following process:
[0178] Determine the text sequence composed of the first texts sequentially output by the first decoding network;
[0179] Input the concatenation result of the acoustic features and the first semantic vector corresponding to each of the multiple voice signal blocks and the text sequence into the offline speech recognition model, so as to output the second text corresponding to the multiple voice signal blocks through the offline speech recognition model.
[0180] Optionally, the offline speech recognition model includes a second encoding network, a third encoding network, and a second decoding network. Based on this, the server can obtain the second text according to the following process:
[0181] Inputting the concatenation result of the acoustic features corresponding to each of the plurality of speech signal blocks and the first semantic vector into the second encoding network, so as to output the second semantic vector corresponding to each of the plurality of speech signal blocks through the second encoding network;
[0182] Inputting the text sequence into the third encoding network to output a third semantic vector corresponding to the text sequence through the third encoding network;
[0183] The second semantic vectors corresponding to each of the multiple speech signal blocks and the third semantic vector are input into the second decoding network, so as to output second texts corresponding to the multiple speech signal blocks through the second decoding network.
[0184] It is worth noting that the steps executed by the above-mentioned server can also be transferred to the client for execution locally when the client has sufficient computing resources.
[0185] The voice interaction method provided in the above embodiments of the present invention can be applied to any scenario where voice recognition of a voice signal stream is required, such as a conference scenario or a live broadcast scenario.
[0186] Taking a conference scenario as an example, an embodiment of the present invention can provide the following voice interaction method that can be applied to a conference scenario:
[0187] Acquire voice signal blocks in the conference voice signal stream, each voice signal block having a preset duration;
[0188] Encoding the acoustic features of the currently generated speech signal block through a first encoding network in the streaming speech recognition model to sequentially obtain first semantic vectors corresponding to each of the plurality of speech signal blocks, wherein the plurality of speech signal blocks correspond to a continuous speech segment;
[0189] Decoding the first semantic vector corresponding to the currently generated speech signal block through a first decoding network in the streaming speech recognition model to sequentially output first texts corresponding to the plurality of speech signal blocks;
[0190] Inputting the concatenation result of the acoustic features corresponding to each of the plurality of speech signal blocks and the first semantic vector into the offline speech recognition model, so as to output the second text corresponding to the plurality of speech signal blocks through the offline speech recognition model;
[0191] Update the first text corresponding to the plurality of speech signal blocks output by the streaming speech recognition model according to the second text;
[0192] A meeting record is generated based on the second text.
[0193] Optionally, the offline speech recognition model may include a second encoding network, a third encoding network and a second decoding network. Based on this, the process of acquiring the second text may include:
[0194] Determine a character sequence consisting of first characters sequentially output by the first decoding network;
[0195] Inputting the concatenation result of the acoustic features corresponding to each of the plurality of speech signal blocks and the first semantic vector into the second encoding network, so as to output the second semantic vector corresponding to each of the plurality of speech signal blocks through the second encoding network;
[0196] Inputting the text sequence into a third encoding network to output a third semantic vector corresponding to the text sequence through the third encoding network;
[0197] The second semantic vectors and the third semantic vectors corresponding to the multiple speech signal blocks are input into the second decoding network, so as to output the second text corresponding to the multiple speech signal blocks through the second decoding network.
[0198] like Figure 7 As shown, in a conference scenario, the above-mentioned voice interaction method can be executed by a voice interaction device, which can be a terminal device in a conference scenario, referred to as a conference terminal. When the conference terminal has sufficient computing power, the above-mentioned scheme can be completely executed locally in the conference terminal. When the computing power of the conference terminal is insufficient, the core steps can be executed by a server in the cloud. At this time, the conference terminal can only complete the collection and upload of voice signal blocks in the voice signal stream and the display and recording of text recognition results.
[0199] Assume that in a conference scenario, it is necessary to record the speech of the speaker and display it on the screen in real time. At this time, based on the solution provided by the embodiment of the present invention, the speaker's speech content can be displayed on the screen in real time in a companion manner, and after each speaker finishes a whole sentence (defined by VAD), the above-mentioned display content can be corrected based on the output of the offline speech recognition model, and the corrected text sequence can be displayed. In addition, the generation of the meeting record can be completed based on the text output by the offline speech recognition model. For example, suppose that along with the speaker's voice output, the text "I am honored to be able to explain blockchain to you" is displayed on the screen one by one based on the output of the streaming speech recognition model. After this sentence is finished, the offline speech recognition model will obtain the text "I am honored to be invited to explain blockchain to you", and the text output by the offline speech recognition model will replace the text with typos that has been displayed on the screen, and the text output by the offline speech recognition model will be recorded in the meeting record.
[0200] Taking the live broadcast scenario as an example, the embodiment of the present invention can provide the following voice interaction method that can be applied to the live broadcast scenario:
[0201] Acquire voice signal blocks in the host's voice signal stream, each voice signal block having a preset duration;
[0202] Encoding the acoustic features of the currently generated speech signal block through a first encoding network in the streaming speech recognition model to sequentially obtain first semantic vectors corresponding to each of the plurality of speech signal blocks, wherein the plurality of speech signal blocks correspond to a continuous speech segment;
[0203] Decoding the first semantic vector corresponding to the currently generated speech signal block through a first decoding network in the streaming speech recognition model to sequentially output first texts corresponding to the plurality of speech signal blocks;
[0204] Display the first text in the live broadcast interface;
[0205] Inputting the concatenation result of the acoustic features corresponding to each of the plurality of speech signal blocks and the first semantic vector into the offline speech recognition model, so as to output the second text corresponding to the plurality of speech signal blocks through the offline speech recognition model;
[0206] Replace the first text with the second text in the live broadcast interface.
[0207] Optionally, the offline speech recognition model includes a second encoding network, a third encoding network and a second decoding network. Based on this, the process of obtaining the second text may include:
[0208] Determine a character sequence consisting of first characters sequentially output by the first decoding network;
[0209] Inputting the concatenation result of the acoustic features corresponding to each of the plurality of speech signal blocks and the first semantic vector into the second encoding network, so as to output the second semantic vector corresponding to each of the plurality of speech signal blocks through the second encoding network;
[0210] Inputting the text sequence into a third encoding network to output a third semantic vector corresponding to the text sequence through the third encoding network;
[0211] The second semantic vectors and the third semantic vectors corresponding to the multiple speech signal blocks are input into the second decoding network, so as to output the second text corresponding to the multiple speech signal blocks through the second decoding network.
[0212] like Figure 8 As shown in , the anchor's terminal device can segment the voice signal stream spoken by the anchor into voice signal blocks, and transmit the obtained voice signal blocks to the cloud server in real time. The server displays the anchor's speech content in the form of subtitles in the live broadcast interface, so that the audience who pulls the live video stream can see the anchor's speech content through the live broadcast interface. Of course, the anchor's terminal device can also directly upload the collected voice signal stream to the cloud server, and the server will complete the segmentation of the voice signal blocks and subsequent processing. The detailed processing process of the server can refer to the relevant description in the aforementioned embodiment, which will not be repeated here.
[0213] The following will describe in detail the speech recognition device and speech interaction device of one or more embodiments of the present invention. Those skilled in the art will appreciate that these speech recognition devices and speech interaction devices can be configured using commercially available hardware components through the steps taught in this solution.
[0214] Fig. 9 A schematic diagram of the structure of a speech recognition device provided by an embodiment of the present invention is shown in FIG. Fig. 9 As shown, the device includes: a streaming encoding module 11, a streaming decoding module 12, an offline identification module 13, and an output updating module 14.
[0215] The streaming encoding module 11 is used to encode the acoustic features of the currently generated speech signal block through the first encoding network in the streaming speech recognition model to obtain first semantic vectors corresponding to each of the multiple speech signal blocks in turn, wherein the multiple speech signal blocks correspond to a continuous speech, and each speech signal block has a preset duration.
[0216] The streaming decoding module 12 is used to decode the first semantic vector corresponding to the currently generated speech signal block through the first decoding network in the streaming speech recognition model, so as to output the first text corresponding to the multiple speech signal blocks in sequence.
[0217] The offline recognition module 13 is used to input the concatenation result of the acoustic features corresponding to each of the multiple speech signal blocks and the first semantic vector into the offline speech recognition model, so as to output the second text corresponding to the multiple speech signal blocks through the offline speech recognition model.
[0218] The output updating module 14 is used to update the first text corresponding to the multiple speech signal blocks output by the streaming speech recognition model according to the second text.
[0219] Optionally, the streaming encoding module 11 may also be used to: determine a delay value corresponding to a current speech recognition application scenario, the preset duration being the delay value.
[0220] Optionally, the streaming encoding module 11 is specifically used to: when receiving the first speech signal block currently generated, obtain at least one second speech signal block generated before the first speech signal block, and use the first speech signal block and the at least one second speech signal block as the multiple speech signal blocks; obtain the acoustic features of the first speech signal block and the acoustic features of the at least one second speech signal block; encode the concatenation result of the acoustic features of the first speech signal block and the acoustic features of the at least one second speech signal block through the first encoding network to obtain the first semantic vector corresponding to the first speech signal block.
[0221] Among them, optionally, for any speech signal block among the multiple speech signal blocks, the streaming encoding module 11 can obtain the acoustic features of any speech signal block in the following manner: performing frame processing on any speech signal block to obtain multiple frames of speech signals; extracting the acoustic features corresponding to each of the multiple frames of speech signals; and determining the acoustic features of any speech signal block based on the acoustic features corresponding to each of the multiple frames of speech signals.
[0222] Optionally, the streaming speech recognition model includes a prediction network and an attention network. Based on this, the streaming decoding module 12 can be specifically used to: predict the first semantic vector corresponding to the first speech signal block through the prediction network to obtain the number of characters contained in the first speech signal block; determine the weights and weighted summation results of the first semantic vectors corresponding to the first speech signal block and the at least one second speech signal block in each weight calculation process through the attention network, and input the weighted summation result into the first decoding network to output the characters corresponding to the first speech signal block through the first decoding network, and the number of characters is used to constrain the number of weight calculations.
[0223] Wherein, optionally, the number of characters contained in the first speech signal block is greater than or equal to 1, and the streaming decoding module 12 can be specifically used for:
[0224] Initialize the weight calculation times to the number of characters, and iterate the following process until the weight calculation times is reduced to 0:
[0225] If the number of weight calculations is not 0, the first semantic vectors corresponding to the first speech signal block and the at least one second speech signal block and the last word output by the first decoding network are input into the attention network, so that the attention network determines the weights of the first semantic vectors corresponding to the first speech signal block and the at least one second speech signal block based on the last word, and determines the weighted summation result of the first semantic vectors corresponding to the first speech signal block and the at least one second speech signal block according to the weight, and inputs the weighted summation result into the first decoding network;
[0226] Obtaining the currently output text through the first decoding network, and inputting the text into the attention network;
[0227] The number of weight calculations is reduced by one.
[0228] Optionally, the offline recognition module 13 is specifically used to: determine a text sequence consisting of first text output in sequence by the first decoding network; input the concatenation results of the acoustic features and the first semantic vector corresponding to each of the multiple speech signal blocks and the text sequence into an offline speech recognition model, so as to output second text corresponding to the multiple speech signal blocks through the offline speech recognition model.
[0229] Optionally, the offline speech recognition model includes a second encoding network, a third encoding network and a second decoding network. Based on this, the offline recognition module 13 is specifically used to: input the concatenation result of the acoustic features and the first semantic vector corresponding to each of the multiple speech signal blocks into the second encoding network, so as to output the second semantic vector corresponding to each of the multiple speech signal blocks through the second encoding network; input the text sequence into the third encoding network, so as to output the third semantic vector corresponding to the text sequence through the third encoding network; input the second semantic vector corresponding to each of the multiple speech signal blocks and the third semantic vector into the second decoding network, so as to output the second text corresponding to the multiple speech signal blocks through the second decoding network.
[0230] Fig. 9 The device shown can perform the aforementioned Figures 1 to 5The speech recognition method provided in the illustrated embodiment, the detailed execution process and technical effects are described in the aforementioned embodiment and will not be repeated here.
[0231] In one possible design, the above Fig. 9 The structure of the speech recognition device shown can be implemented as an electronic device, such as Fig.10 As shown, the electronic device may include: a processor 21 and a memory 22. The memory 22 stores executable code, and when the executable code is executed by the processor 21, the processor 21 can at least implement the above-mentioned Figures 1 to 5 The speech recognition method provided in the illustrated embodiment.
[0232] Optionally, the electronic device may further include a communication interface 23 for communicating with other devices.
[0233] In addition, an embodiment of the present invention provides a non-transitory machine-readable storage medium, wherein an executable code is stored on the non-transitory machine-readable storage medium. When the executable code is executed by a processor of an electronic device, the processor can at least implement the above-mentioned Figures 1 to 5 The speech recognition method provided in the illustrated embodiment.
[0234] Fig.11 A schematic diagram of the structure of a voice interaction device provided by an embodiment of the present invention is shown in FIG. Fig.11 As shown, the device includes: a collection module 31, a sending module 32, and a display module 33.
[0235] The acquisition module 31 is used to acquire speech signal blocks in the speech signal stream, each speech signal block having a preset duration.
[0236] The sending module 32 is used to upload the collected speech signal blocks to the server, so that the server encodes the acoustic features of the currently generated speech signal blocks through the first encoding network in the streaming speech recognition model, so as to obtain the first semantic vectors corresponding to the multiple speech signal blocks in turn, decode the first semantic vectors corresponding to the currently generated speech signal blocks through the first decoding network in the streaming speech recognition model, so as to output the first texts corresponding to the multiple speech signal blocks in turn, and input the splicing results of the acoustic features and the first semantic vectors corresponding to the multiple speech signal blocks into the offline speech recognition model, so as to output the second texts corresponding to the multiple speech signal blocks through the offline speech recognition model; the multiple speech signal blocks correspond to a continuous speech.
[0237] The display module 33 is configured to display the first text received from the server; and to update the first text according to the second text received from the server.
[0238] Optionally, the server obtains the second text according to the following process:
[0239] Determine a character sequence consisting of first characters sequentially output by the first decoding network;
[0240] The concatenation results of the acoustic features and the first semantic vector corresponding to each of the multiple speech signal blocks and the text sequence are input into an offline speech recognition model, so as to output second text corresponding to the multiple speech signal blocks through the offline speech recognition model.
[0241] Optionally, the offline speech recognition model includes a second encoding network, a third encoding network, and a second decoding network. Based on this, the server obtains the second text according to the following process:
[0242] Inputting the concatenation result of the acoustic features corresponding to each of the plurality of speech signal blocks and the first semantic vector into the second encoding network, so as to output the second semantic vector corresponding to each of the plurality of speech signal blocks through the second encoding network;
[0243] Inputting the text sequence into the third encoding network to output a third semantic vector corresponding to the text sequence through the third encoding network;
[0244] The second semantic vectors corresponding to each of the multiple speech signal blocks and the third semantic vector are input into the second decoding network, so as to output second texts corresponding to the multiple speech signal blocks through the second decoding network.
[0245] In one possible design, the above Fig.11 The structure of the voice interaction device shown can be implemented as a voice interaction device, such as Fig.12 As shown, the voice interaction device may include: a processor 41, a memory 42, and a display 43; wherein the memory 42 stores executable code, and when the executable code is executed by the processor 41, the processor 41 performs the following steps:
[0246] Collecting speech signal blocks in a speech signal stream, each speech signal block having a preset duration;
[0247] The collected speech signal blocks are uploaded to the server, so that the server encodes the acoustic features of the currently generated speech signal blocks through the first encoding network in the streaming speech recognition model, so as to obtain first semantic vectors corresponding to each of the multiple speech signal blocks in sequence, decode the first semantic vectors corresponding to the currently generated speech signal blocks through the first decoding network in the streaming speech recognition model, so as to output first texts corresponding to the multiple speech signal blocks in sequence, and input the splicing results of the acoustic features and the first semantic vectors corresponding to each of the multiple speech signal blocks into the offline speech recognition model, so as to output second texts corresponding to the multiple speech signal blocks through the offline speech recognition model; the multiple speech signal blocks correspond to a continuous speech;
[0248] The first text received from the server is displayed via the display 43; and the first text is updated according to the second text received from the server.
[0249] The processor 41 may also execute other related steps in the voice interaction method provided in the aforementioned related embodiments, which will not be described here.
[0250] Optionally, the voice interaction device may further include a communication interface 44 for communicating with other devices.
[0251] In addition, an embodiment of the present invention further provides a non-transitory machine-readable storage medium, wherein the non-transitory machine-readable storage medium stores executable code. Fig.12 When the processor of the voice interaction device shown is executed, the processor executes the corresponding voice interaction method.
[0252] Fig.13 A schematic diagram of the structure of a voice interaction device provided by an embodiment of the present invention is shown in FIG. Fig.13 As shown, the device includes: an acquisition module 51, a streaming encoding and decoding module 52, an offline processing module 53, and a generation module 54.
[0253] The acquisition module 51 is used to acquire voice signal blocks in the conference voice signal stream, each voice signal block having a preset duration.
[0254] The streaming encoding and decoding module 52 is used to encode the acoustic features of the currently generated speech signal block through the first encoding network in the streaming speech recognition model to obtain first semantic vectors corresponding to each of the multiple speech signal blocks in sequence, and the multiple speech signal blocks correspond to a continuous speech; and decode the first semantic vector corresponding to the currently generated speech signal block through the first decoding network in the streaming speech recognition model to output first text corresponding to the multiple speech signal blocks in sequence.
[0255] The offline processing module 53 is used to input the concatenation results of the acoustic features and the first semantic vector corresponding to each of the multiple speech signal blocks into the offline speech recognition model, so as to output the second text corresponding to the multiple speech signal blocks through the offline speech recognition model; and to update the first text corresponding to the multiple speech signal blocks output by the streaming speech recognition model according to the second text.
[0256] The generating module 54 is used to generate a meeting record according to the second text.
[0257] Optionally, the offline speech recognition model includes a second encoding network, a third encoding network and a second decoding network. Based on this, the offline processing module 53 can be specifically used to: determine a text sequence consisting of the first text outputted in sequence by the first decoding network; input the concatenation result of the acoustic features and the first semantic vector corresponding to each of the multiple speech signal blocks into the second encoding network, so as to output the second semantic vector corresponding to each of the multiple speech signal blocks through the second encoding network; input the text sequence into the third encoding network, so as to output the third semantic vector corresponding to the text sequence through the third encoding network; input the second semantic vector corresponding to each of the multiple speech signal blocks and the third semantic vector into the second decoding network, so as to output the second text corresponding to the multiple speech signal blocks through the second decoding network.
[0258] In one possible design, the above Fig.13 The structure of the voice interaction device shown can be implemented as a voice interaction device, such as Fig.14 As shown, the voice interaction device may include: a processor 61, a memory 62, and a display 63; wherein the memory 62 stores executable code, and when the executable code is executed by the processor 61, the processor 61 performs the following steps:
[0259] Acquire voice signal blocks in the conference voice signal stream, each voice signal block having a preset duration;
[0260] Encoding the acoustic features of the currently generated speech signal block through a first encoding network in the streaming speech recognition model to sequentially obtain first semantic vectors corresponding to each of the plurality of speech signal blocks, wherein the plurality of speech signal blocks correspond to a continuous speech segment;
[0261] Decoding the first semantic vector corresponding to the currently generated speech signal block through the first decoding network in the streaming speech recognition model to sequentially output the first text corresponding to the multiple speech signal blocks;
[0262] Inputting the concatenation result of the acoustic features and the first semantic vector corresponding to each of the plurality of speech signal blocks into an offline speech recognition model, so as to output second texts corresponding to the plurality of speech signal blocks through the offline speech recognition model;
[0263] Update the first text corresponding to the plurality of speech signal blocks output by the streaming speech recognition model according to the second text;
[0264] Generate meeting minutes based on the second text.
[0265] The processor 61 can also execute other related steps in the voice interaction method provided in the aforementioned related embodiments, which are not repeated here.
[0266] Optionally, the voice interaction device may further include a communication interface 64 for communicating with other devices.
[0267] In addition, an embodiment of the present invention further provides a non-transitory machine-readable storage medium, wherein the non-transitory machine-readable storage medium stores executable code. Fig.14 When the processor of the voice interaction device shown is executed, the processor executes the corresponding voice interaction method.
[0268] Fig.15 A schematic diagram of the structure of a voice interaction device provided by an embodiment of the present invention is shown in FIG. Fig.15 As shown, the device includes: an acquisition module 71, a streaming encoding and decoding module 72, a display module 73, and an offline processing module 74.
[0269] The acquisition module 71 is used to acquire voice signal blocks in the host's voice signal stream, each voice signal block having a preset duration.
[0270] The streaming encoding and decoding module 72 is used to encode the acoustic features of the currently generated speech signal block through the first encoding network in the streaming speech recognition model to obtain first semantic vectors corresponding to each of the multiple speech signal blocks in sequence, and the multiple speech signal blocks correspond to a continuous speech; and decode the first semantic vector corresponding to the currently generated speech signal block through the first decoding network in the streaming speech recognition model to output first text corresponding to the multiple speech signal blocks in sequence.
[0271] The display module 73 is used to display the first text in the live broadcast interface.
[0272] The offline processing module 74 is used to input the concatenation result of the acoustic features corresponding to each of the multiple speech signal blocks and the first semantic vector into the offline speech recognition model, so as to output the second text corresponding to the multiple speech signal blocks through the offline speech recognition model.
[0273] The display module 73 is further used to replace the first text with the second text in the live broadcast interface.
[0274] Optionally, the offline speech recognition model includes a second encoding network, a third encoding network and a second decoding network. Based on this, the offline processing module 74 can be specifically used to: determine a text sequence consisting of the first text outputted in sequence by the first decoding network; input the concatenation result of the acoustic features and the first semantic vector corresponding to each of the plurality of speech signal blocks into the second encoding network, so as to output the second semantic vector corresponding to each of the plurality of speech signal blocks through the second encoding network; input the text sequence into the third encoding network, so as to output the third semantic vector corresponding to the text sequence through the third encoding network; input the second semantic vector corresponding to each of the plurality of speech signal blocks and the third semantic vector into the second decoding network, so as to output the second text corresponding to the plurality of speech signal blocks through the second decoding network.
[0275] In one possible design, the above Fig.15 The structure of the voice interaction device shown can be implemented as a voice interaction device, such as Fig.16 As shown, the voice interaction device may include: a processor 81, a memory 82, and a display 83; wherein the memory 82 stores executable code, and when the executable code is executed by the processor 81, the processor 81 performs the following steps:
[0276] Acquire voice signal blocks in the host's voice signal stream, each voice signal block having a preset duration;
[0277] Encoding the acoustic features of the currently generated speech signal block through a first encoding network in the streaming speech recognition model to sequentially obtain first semantic vectors corresponding to each of the plurality of speech signal blocks, wherein the plurality of speech signal blocks correspond to a continuous speech segment;
[0278] Decoding the first semantic vector corresponding to the currently generated speech signal block through the first decoding network in the streaming speech recognition model to sequentially output the first text corresponding to the multiple speech signal blocks;
[0279] Displaying the first text in the live broadcast interface;
[0280] Inputting the concatenation result of the acoustic features and the first semantic vector corresponding to each of the plurality of speech signal blocks into an offline speech recognition model, so as to output second texts corresponding to the plurality of speech signal blocks through the offline speech recognition model;
[0281] The first text is replaced with the second text in the live broadcast interface.
[0282] The processor 81 can also execute other related steps in the voice interaction method provided in the aforementioned related embodiments, which are not repeated here.
[0283] Optionally, the voice interaction device may further include a communication interface 84 for communicating with other devices.
[0284] In addition, an embodiment of the present invention further provides a non-transitory machine-readable storage medium, wherein the non-transitory machine-readable storage medium stores executable code. Fig.16 When the processor of the voice interaction device shown is executed, the processor executes the corresponding voice interaction method.
[0285] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. Those of ordinary skill in the art may understand and implement the present invention without creative effort.
[0286] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by adding a necessary general hardware platform, and of course can also be implemented by combining hardware and software. Based on such an understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a computer product, and the present invention can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0287] The speech recognition method provided in the embodiment of the present invention can be executed by a certain program / software, and the program / software can be provided by the network side. The electronic device mentioned in the above embodiment can download the program / software to a local non-volatile storage medium, and when it needs to execute the above speech recognition method, the program / software is read into the memory by the CPU, and then the CPU executes the program / software to implement the speech recognition method provided in the above embodiment. The execution process can refer to the above Figures 1 to 5 Instructions in .
[0288] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A speech recognition method, It is characterized in that The streaming speech recognition model includes a first encoding network, a prediction network, an attention network and a first decoding network, and the method includes: Encoding the acoustic features of the currently generated speech signal block through the first encoding network to sequentially obtain first semantic vectors corresponding to each of the plurality of speech signal blocks, wherein the plurality of speech signal blocks correspond to a continuous speech segment, and each speech signal block has a preset duration; Predicting the first semantic vector corresponding to the currently generated first speech signal block through the prediction network to obtain the number of characters contained in the first speech signal block, wherein the number of characters is used to constrain the number of calculations for determining the weight of the first semantic vector corresponding to the first speech signal block through the attention network; Outputting the text corresponding to the first speech signal block through the first decoding network according to the weight of the first semantic vector corresponding to the first speech signal block, so as to obtain the first text corresponding to the plurality of speech signal blocks sequentially output by the first decoding network; Inputting the concatenation result of the acoustic features and the first semantic vector corresponding to each of the plurality of speech signal blocks into an offline speech recognition model, so as to output second texts corresponding to the plurality of speech signal blocks through the offline speech recognition model; The first text corresponding to the multiple speech signal blocks output by the streaming speech recognition model is updated according to the second text.
2. The method according to claim 1, It is characterized in that The step of encoding the acoustic features of the currently generated speech signal block by the first encoding network includes: When receiving the currently generated first speech signal block, obtaining at least one second speech signal block generated before the first speech signal block, and using the first speech signal block and the at least one second speech signal block as the multiple speech signal blocks; Acquiring acoustic features of the first speech signal block and acoustic features of the at least one second speech signal block; The first encoding network encodes a concatenation result of the acoustic features of the first speech signal block and the acoustic features of the at least one second speech signal block to obtain a first semantic vector corresponding to the first speech signal block.
3. The method according to claim 1, It is characterized in that For any speech signal block among the multiple speech signal blocks, the acoustic features of the any speech signal block are obtained according to the following method: Performing frame processing on any of the speech signals to obtain multiple frames of speech signals; Extracting acoustic features corresponding to each of the multiple frames of speech signals; The acoustic features of any speech signal block are determined according to the acoustic features corresponding to each of the multiple frames of speech signals.
4. The method according to claim 2, It is characterized in that The step of outputting text corresponding to the first speech signal block through the first decoding network according to the weight of the first semantic vector corresponding to the first speech signal block includes: The attention network is used to determine the weights of the first semantic vectors corresponding to the first speech signal block and the at least one second speech signal block in each weight calculation process, as well as the weighted summation result, and the weighted summation result is input into the first decoding network to output the text corresponding to the first speech signal block through the first decoding network, and the number of characters is used to constrain the number of weight calculations.
5. The method according to claim 4, It is characterized in that The number of characters contained in the first speech signal block is greater than or equal to 1; The step of determining the weights of the first semantic vectors corresponding to the first speech signal block and the at least one second speech signal block and the weighted summation result in each weight calculation process through the attention network, and inputting the weighted summation result into the first decoding network, so as to output the text corresponding to the first speech signal block through the first decoding network, includes: Initialize the weight calculation times to the number of characters, and iterate the following process until the weight calculation times is reduced to 0: If the number of weight calculations is not 0, the first semantic vectors corresponding to the first speech signal block and the at least one second speech signal block and the last word output by the first decoding network are input into the attention network, so that the attention network determines the weights of the first semantic vectors corresponding to the first speech signal block and the at least one second speech signal block based on the last word, and determines the weighted summation result of the first semantic vectors corresponding to the first speech signal block and the at least one second speech signal block according to the weight, and inputs the weighted summation result into the first decoding network; Obtaining the currently output text through the first decoding network, and inputting the text into the attention network; The number of weight calculations is reduced by one.
6. The method according to claim 1, It is characterized in that The method further comprises: Determine a delay value corresponding to the current speech recognition application scenario, the preset duration being the delay value.
7. The method according to claim 1, It is characterized in that The step of inputting the concatenation result of the acoustic features and the first semantic vector corresponding to each of the plurality of speech signal blocks into the offline speech recognition model comprises: Determine a character sequence consisting of first characters sequentially output by the first decoding network; The concatenation results of the acoustic features and the first semantic vector corresponding to each of the multiple speech signal blocks and the text sequence are input into an offline speech recognition model, so as to output second text corresponding to the multiple speech signal blocks through the offline speech recognition model.
8. The method according to claim 7, It is characterized in that The offline speech recognition model includes a second encoding network, a third encoding network and a second decoding network; The step of inputting the concatenation result of the acoustic features and the first semantic vector corresponding to each of the plurality of speech signal blocks and the text sequence into an offline speech recognition model, so as to output second text corresponding to the plurality of speech signal blocks through the offline speech recognition model, comprises: Inputting the concatenation result of the acoustic features corresponding to each of the plurality of speech signal blocks and the first semantic vector into the second encoding network, so as to output the second semantic vector corresponding to each of the plurality of speech signal blocks through the second encoding network; Inputting the text sequence into the third encoding network to output a third semantic vector corresponding to the text sequence through the third encoding network; The second semantic vectors corresponding to each of the multiple speech signal blocks and the third semantic vector are input into the second decoding network, so as to output second texts corresponding to the multiple speech signal blocks through the second decoding network.
9. A speech recognition device, It is characterized in that The streaming speech recognition model includes a first encoding network, a prediction network, an attention network and a first decoding network, and the device includes: A streaming encoding module, configured to encode the acoustic features of the currently generated speech signal block through the first encoding network to sequentially obtain first semantic vectors corresponding to each of the plurality of speech signal blocks, wherein the plurality of speech signal blocks correspond to a continuous speech segment, and each speech signal block has a preset duration; A streaming decoding module, configured to predict the first semantic vector corresponding to the currently generated first speech signal block through the prediction network to obtain the number of characters contained in the first speech signal block, the number of characters being used to constrain the number of calculations for determining the weight of the first semantic vector corresponding to the first speech signal block through the attention network, and outputting the characters corresponding to the first speech signal block through the first decoding network according to the weight of the first semantic vector corresponding to the first speech signal block, so as to obtain the first characters corresponding to the multiple speech signal blocks sequentially output by the first decoding network; An offline recognition module, used for inputting the concatenation result of the acoustic features corresponding to each of the plurality of speech signal blocks and the first semantic vector into an offline speech recognition model, so as to output second text corresponding to the plurality of speech signal blocks through the offline speech recognition model; An output update module is used to update the first text corresponding to the multiple speech signal blocks output by the streaming speech recognition model according to the second text.
10. The device according to claim 9, It is characterized in that The streaming encoding module is specifically used for: When receiving the currently generated first speech signal block, obtaining at least one second speech signal block generated before the first speech signal block, and using the first speech signal block and the at least one second speech signal block as the multiple speech signal blocks; Acquiring acoustic features of the first speech signal block and acoustic features of the at least one second speech signal block; The first encoding network encodes a concatenation result of the acoustic features of the first speech signal block and the acoustic features of the at least one second speech signal block to obtain a first semantic vector corresponding to the first speech signal block.
11. The device according to claim 9, It is characterized in that The offline recognition module is specifically used to: determine a text sequence consisting of first characters output in sequence by the first decoding network; input the concatenation results of the acoustic features and the first semantic vectors corresponding to each of the multiple speech signal blocks and the text sequence into an offline speech recognition model, so as to output second characters corresponding to the multiple speech signal blocks through the offline speech recognition model.
12. The device according to claim 11, It is characterized in that The offline speech recognition model includes a second encoding network, a third encoding network and a second decoding network; The offline recognition module is specifically used for: Inputting the concatenation result of the acoustic features corresponding to each of the plurality of speech signal blocks and the first semantic vector into the second encoding network, so as to output the second semantic vector corresponding to each of the plurality of speech signal blocks through the second encoding network; Inputting the text sequence into the third encoding network to output a third semantic vector corresponding to the text sequence through the third encoding network; The second semantic vectors corresponding to each of the multiple speech signal blocks and the third semantic vector are input into the second decoding network, so as to output second texts corresponding to the multiple speech signal blocks through the second decoding network.
13. An electronic device, It is characterized in that include: A memory and a processor; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor executes the speech recognition method as described in any one of claims 1 to 8.
14. A non-transitory machine-readable storage medium, It is characterized in that The non-transitory machine-readable storage medium stores executable code, and when the executable code is executed by a processor of an electronic device, the processor is caused to execute the speech recognition method according to any one of claims 1 to 8.
15. A voice interaction method, It is characterized in that The streaming speech recognition model includes a first encoding network, a prediction network, an attention network and a first decoding network, and the method includes: Collecting speech signal blocks in a speech signal stream, each speech signal block having a preset duration; The collected speech signal blocks are uploaded to the server, so that the server encodes the acoustic features of the currently generated speech signal blocks through the first encoding network, so as to obtain first semantic vectors corresponding to each of the multiple speech signal blocks in sequence, and the first semantic vector corresponding to the currently generated first speech signal block is predicted through the prediction network to obtain the number of characters contained in the first speech signal block, the number of characters is used to constrain the number of calculations of the weight of the first semantic vector corresponding to the first speech signal block determined by the attention network, and according to the weight of the first semantic vector corresponding to the first speech signal block, the text corresponding to the first speech signal block is output through the first decoding network to obtain the first text corresponding to the multiple speech signal blocks outputted in sequence by the first decoding network, and the splicing results of the acoustic features and the first semantic vectors corresponding to each of the multiple speech signal blocks are input into the offline speech recognition model, so as to output second text corresponding to the multiple speech signal blocks through the offline speech recognition model; the multiple speech signal blocks correspond to a continuous speech; Displaying the first text received from the server; Update the first text according to the second text received from the server.
16. The method according to claim 15, It is characterized in that The server obtains the second character according to the following process: Determine a character sequence consisting of first characters sequentially output by the first decoding network; The concatenation results of the acoustic features and the first semantic vector corresponding to each of the multiple speech signal blocks and the text sequence are input into an offline speech recognition model, so as to output second text corresponding to the multiple speech signal blocks through the offline speech recognition model.
17. The method according to claim 16, It is characterized in that The offline speech recognition model includes a second encoding network, a third encoding network and a second decoding network; The server obtains the second character according to the following process: Inputting the concatenation result of the acoustic features corresponding to each of the plurality of speech signal blocks and the first semantic vector into the second encoding network, so as to output the second semantic vector corresponding to each of the plurality of speech signal blocks through the second encoding network; Inputting the text sequence into the third encoding network to output a third semantic vector corresponding to the text sequence through the third encoding network; The second semantic vectors corresponding to each of the multiple speech signal blocks and the third semantic vector are input into the second decoding network, so as to output second texts corresponding to the multiple speech signal blocks through the second decoding network.
18. A voice interaction device, It is characterized in that include: A memory, a processor, and a display; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor executes the voice interaction method as described in any one of claims 15 to 17.
19. A non-transitory machine-readable storage medium, It is characterized in that The non-transitory machine-readable storage medium stores executable code, and when the executable code is executed by a processor of an electronic device, the processor executes the voice interaction method as described in any one of claims 15 to 17.
20. A voice interaction method, It is characterized in that The streaming speech recognition model includes a first encoding network, a prediction network, an attention network and a first decoding network, and the method includes: Acquire voice signal blocks in the conference voice signal stream, each voice signal block having a preset duration; Encoding the acoustic features of the currently generated speech signal block through the first encoding network to sequentially obtain first semantic vectors corresponding to each of the plurality of speech signal blocks, wherein the plurality of speech signal blocks correspond to a continuous speech segment; Predicting the first semantic vector corresponding to the currently generated first speech signal block through the prediction network to obtain the number of characters contained in the first speech signal block, wherein the number of characters is used to constrain the number of calculations for determining the weight of the first semantic vector corresponding to the first speech signal block through the attention network; Outputting the text corresponding to the first speech signal block through the first decoding network according to the weight of the first semantic vector corresponding to the first speech signal block, so as to obtain the first text corresponding to the plurality of speech signal blocks sequentially output by the first decoding network; Inputting the concatenation result of the acoustic features and the first semantic vector corresponding to each of the plurality of speech signal blocks into an offline speech recognition model, so as to output second texts corresponding to the plurality of speech signal blocks through the offline speech recognition model; Update the first text corresponding to the plurality of speech signal blocks output by the streaming speech recognition model according to the second text; Generate meeting minutes based on the second text.
21. The method according to claim 20, It is characterized in that The offline speech recognition model includes a second encoding network, a third encoding network and a second decoding network; The process of obtaining the second character includes: Determine a character sequence consisting of first characters sequentially output by the first decoding network; Inputting the concatenation result of the acoustic features corresponding to each of the plurality of speech signal blocks and the first semantic vector into the second encoding network, so as to output the second semantic vector corresponding to each of the plurality of speech signal blocks through the second encoding network; Inputting the text sequence into the third encoding network to output a third semantic vector corresponding to the text sequence through the third encoding network; The second semantic vectors corresponding to each of the multiple speech signal blocks and the third semantic vector are input into the second decoding network, so as to output second texts corresponding to the multiple speech signal blocks through the second decoding network.
22. A voice interaction device, It is characterized in that include: A memory, a processor, and a display; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor executes the voice interaction generation method as described in claim 20 or 21.
23. A voice interaction method, It is characterized in that The streaming speech recognition model includes a first encoding network, a prediction network, an attention network and a first decoding network, and the method includes: Acquire voice signal blocks in the host's voice signal stream, each voice signal block having a preset duration; Encoding the acoustic features of the currently generated speech signal block through the first encoding network to sequentially obtain first semantic vectors corresponding to each of the plurality of speech signal blocks, wherein the plurality of speech signal blocks correspond to a continuous speech segment; Predicting the first semantic vector corresponding to the currently generated first speech signal block through the prediction network to obtain the number of characters contained in the first speech signal block, wherein the number of characters is used to constrain the number of calculations for determining the weight of the first semantic vector corresponding to the first speech signal block through the attention network; Outputting the text corresponding to the first speech signal block through the first decoding network according to the weight of the first semantic vector corresponding to the first speech signal block, so as to obtain the first text corresponding to the plurality of speech signal blocks sequentially output by the first decoding network; Displaying the first text in the live broadcast interface; Inputting the concatenation result of the acoustic features and the first semantic vector corresponding to each of the plurality of speech signal blocks into an offline speech recognition model, so as to output second texts corresponding to the plurality of speech signal blocks through the offline speech recognition model; The first text is replaced with the second text in the live broadcast interface.
24. The method according to claim 23, It is characterized in that The offline speech recognition model includes a second encoding network, a third encoding network and a second decoding network; The process of obtaining the second character includes: Determine a character sequence consisting of first characters sequentially output by the first decoding network; Inputting the concatenation result of the acoustic features corresponding to each of the plurality of speech signal blocks and the first semantic vector into the second encoding network, so as to output the second semantic vector corresponding to each of the plurality of speech signal blocks through the second encoding network; Inputting the text sequence into the third encoding network to output a third semantic vector corresponding to the text sequence through the third encoding network; The second semantic vectors corresponding to each of the multiple speech signal blocks and the third semantic vector are input into the second decoding network, so as to output second texts corresponding to the multiple speech signal blocks through the second decoding network.
25. A voice interaction device, It is characterized in that include: A memory, a processor, and a display; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor executes the voice interaction method as described in claim 23 or 24.
Citation Information
Patent Citations
Speech recognition method and device thereof
CN111261166A
Training method and decoding method of streaming end-to-end speech recognition model
CN111415667A