Voice interaction and voice recognition methods, devices, equipment, and storage media
Through the end-to-end speech recognition model, using preset blocking and grouping time division, combined with encoding and decoding networks, the real-time and accuracy problems in long speech recognition tasks are solved, and efficient recognition and word conversion processing of streaming speech signals are realized.
Patent Information
- Application Number
- CN202011134763.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-21
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-10-21
AI Technical Summary
Existing end-to-end speech recognition models are poor in long speech recognition tasks, especially in streaming speech recognition tasks, which are difficult to achieve real-time and accurate speech recognition.
An end-to-end speech recognition model is adopted to encode and decode speech signals by presetting the blocking time and grouping time of speech signals, and combining memory modules and attention networks to realize real-time speech recognition of long speech signals.
It improves the recognition accuracy and real-time nature of long voice signals, and can display voice content in real-time in streaming voice input scenarios, which is suitable for real-time voice-to-text conversion in conference and live broadcast scenarios.
Smart Images

Figure CN114464170B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and in particular, to a voice interaction and voice recognition method, apparatus, device, and storage medium. Background Art
[0002] In many application scenarios of artificial intelligence, voice recognition tasks are encountered. At present, a mainstream voice recognition solution is to use an end-to-end voice recognition model to complete the voice recognition task. The end-to-end voice recognition model jointly optimizes the acoustic model and the language model through one model, and can also not require the use of a pronunciation dictionary, which can not only greatly reduce the model training complexity, but also obtain good performance.
[0003] In the voice recognition task, there is a task type - the streaming voice recognition task, that is, the task of real-time voice recognition for the streaming generated voice signal. For example, in the real-time voice input scenario, it is necessary to display the voice content spoken by the user on the screen in real time; for another example, in the meeting scenario, it is necessary to record the voice spoken by the speaker in text form in real time.
[0004] At present, many end-to-end voice recognition models have good performance in the short voice recognition task. However, for the long voice generated by streaming, the performance is often poor. Among them, short voice and long voice can be divided by setting a voice duration threshold. For example, if the duration of a voice is less than the threshold, then this voice is considered a short voice, otherwise, it is a long voice. Summary of the Invention
[0005] Embodiments of the present invention provide a voice interaction and voice recognition method, apparatus, device, and storage medium, which can accurately complete the voice recognition processing of the long voice signal generated by streaming.
[0006] In a first aspect, an embodiment of the present invention provides a voice recognition method, which includes:
[0007] Obtain a first voice signal block in the voice signal stream based on a preset voice signal block duration;
[0008] Obtain at least one voice signal block included in a target group corresponding to the first voice signal block, where the target group includes the first voice signal block, and a group includes the voice signal blocks sequentially generated within a preset group duration;
[0009] Encode the at least one voice signal block through an encoding network in the voice recognition model to obtain a semantic vector corresponding to the first voice signal block;
[0010] Decode the semantic vectors corresponding to each of the at least one speech signal chunk through a decoding network in the speech recognition model to output the text corresponding to the first speech signal chunk.
[0011] In a second aspect, an embodiment of the present invention provides a speech recognition device, which includes:
[0012] An acquisition module, configured to acquire a first speech signal chunk in a speech signal stream based on a preset speech signal chunk duration; and acquire at least one speech signal chunk included in a target group corresponding to the first speech signal chunk, where the target group includes the first speech signal chunk, and a group includes speech signal chunks generated sequentially within a preset group duration;
[0013] An encoding module, configured to encode the at least one speech signal chunk through an encoding network in a speech recognition model to obtain a semantic vector corresponding to the first speech signal chunk;
[0014] A decoding module, configured to decode the semantic vectors corresponding to each of the at least one speech signal chunk through a decoding network in the speech recognition model to output the text corresponding to the first speech signal chunk.
[0015] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor; wherein, an executable code is stored on the memory, and when the executable code is executed by the processor, the processor can at least implement the speech recognition method as described in the first aspect.
[0016] In a fourth aspect, an embodiment of the present invention provides a non-transitory machine-readable storage medium, on which an executable code is stored, and when the executable code is executed by a processor of an electronic device, the processor can at least implement the speech recognition method as described in the first aspect.
[0017] In a fifth aspect, an embodiment of the present invention provides a speech interaction method, which includes:
[0018] Collect speech signal chunks in a speech signal stream, each speech signal chunk having a preset duration;
[0019] Chunk and upload the collected voice signals to the server, so that the server obtains the first voice signal chunk and at least one voice signal chunk included in the target group corresponding to the first voice signal chunk, and encode the at least one voice signal chunk through the encoding network in the voice recognition model to obtain a semantic vector corresponding to the first voice signal chunk; decode the semantic vectors corresponding to the at least one voice signal chunk respectively through the decoding network in the voice recognition model to output the text corresponding to the first voice signal chunk; wherein, the target group includes the first voice signal chunk, and a group includes the voice signal chunks generated in sequence within a preset group duration;
[0020] Display the text received from the server.
[0021] In a sixth aspect, an embodiment of the present invention provides a voice interaction device, which includes:
[0022] An acquisition module, configured to acquire voice signal chunks in a voice signal stream, and each voice signal chunk has a preset duration;
[0023] A sending module, configured to chunk and upload the acquired voice signals to the server, so that the server obtains the first voice signal chunk and at least one voice signal chunk included in the target group corresponding to the first voice signal chunk, and encode the at least one voice signal chunk through the encoding network in the voice recognition model to obtain a semantic vector corresponding to the first voice signal chunk; decode the semantic vectors corresponding to the at least one voice signal chunk respectively through the decoding network in the voice recognition model to output the text corresponding to the first voice signal chunk; wherein, the target group includes the first voice signal chunk, and a group includes the voice signal chunks generated in sequence within a preset group duration;
[0024] A display module, configured to display the text received from the server.
[0025] In a seventh aspect, an embodiment of the present invention provides a voice interaction device, including: a memory, a processor, and a display; wherein, an executable code is stored on the memory, and when the executable code is executed by the processor, the processor can at least implement the voice interaction method as described in the fifth aspect.
[0026] In an eighth aspect, an embodiment of the present invention provides a non-transitory machine-readable storage medium, on which an executable code is stored, and when the executable code is executed by a processor of a voice interaction device, the processor can at least implement the voice interaction method as described in the fifth aspect.
[0027] In a ninth aspect, an embodiment of the present invention provides a voice interaction method, which includes:
[0028] Obtain a first voice signal chunk from the conference voice signal stream based on a preset voice signal chunk duration;
[0029] Obtain at least one voice signal chunk included in a target group corresponding to the first voice signal chunk, where the target group includes the first voice signal chunk, and one group includes voice signal chunks generated sequentially within a preset group duration;
[0030] Encode the at least one voice signal chunk through an encoding network in a voice recognition model to obtain a semantic vector corresponding to the first voice signal chunk;
[0031] Decode the semantic vectors corresponding to the at least one voice signal chunk respectively through a decoding network in the voice recognition model to output the text corresponding to the first voice signal chunk;
[0032] Display the text.
[0033] In a tenth aspect, an embodiment of the present invention provides a voice interaction device, which includes:
[0034] An acquisition module, configured to obtain a first voice signal chunk from the conference voice signal stream based on a preset voice signal chunk duration; and obtain at least one voice signal chunk included in a target group corresponding to the first voice signal chunk, where the target group includes the first voice signal chunk, and one group includes voice signal chunks generated sequentially within a preset group duration;
[0035] An encoding module, configured to encode the at least one voice signal chunk through an encoding network in a voice recognition model to obtain a semantic vector corresponding to the first voice signal chunk;
[0036] A decoding module, configured to decode the semantic vectors corresponding to the at least one voice signal chunk respectively through a decoding network in the voice recognition model to output the text corresponding to the first voice signal chunk;
[0037] A display module, configured to display the text.
[0038] In an eleventh aspect, an embodiment of the present invention provides a voice interaction device, including: a memory, a processor, and a display; wherein, an executable code is stored on the memory, and when the executable code is executed by the processor, the processor can at least implement the voice interaction method as described in the ninth aspect.
[0039] In a twelfth aspect, an embodiment of the present invention provides a non-transitory machine-readable storage medium, on which executable code is stored. When the executable code is executed by a processor of a voice interaction device, the processor can at least implement the voice interaction method as described in the ninth aspect.
[0040] In a thirteenth aspect, an embodiment of the present invention provides a voice interaction method, which includes:
[0041] Based on a preset voice signal chunk duration, obtain a first voice signal chunk in the voice signal stream of the host;
[0042] Obtain at least one voice signal chunk included in a target group corresponding to the first voice signal chunk. The target group includes the first voice signal chunk, and a group includes voice signal chunks generated sequentially within a preset group duration;
[0043] Encode the at least one voice signal chunk through an encoding network in a voice recognition model to obtain a semantic vector corresponding to the first voice signal chunk;
[0044] Decode the semantic vectors corresponding to the at least one voice signal chunk respectively through a decoding network in the voice recognition model to output the text corresponding to the first voice signal chunk;
[0045] Display the text in a live broadcast interface.
[0046] In a fourteenth aspect, an embodiment of the present invention provides a voice interaction device, which includes:
[0047] An acquisition module, configured to obtain a first voice signal chunk in the voice signal stream of the host based on a preset voice signal chunk duration; and obtain at least one voice signal chunk included in a target group corresponding to the first voice signal chunk. The target group includes the first voice signal chunk, and a group includes voice signal chunks generated sequentially within a preset group duration;
[0048] An encoding module, configured to encode the at least one voice signal chunk through an encoding network in a voice recognition model to obtain a semantic vector corresponding to the first voice signal chunk;
[0049] A decoding module, configured to decode the semantic vectors corresponding to the at least one voice signal chunk respectively through a decoding network in the voice recognition model to output the text corresponding to the first voice signal chunk;
[0050] A display module, configured to display the text in a live broadcast interface.
[0051] Fifteenth aspect, an embodiment of the present invention provides a voice interaction device, including: a memory, a processor, and a display; wherein, an executable code is stored on the memory, and when the executable code is executed by the processor, the processor can at least implement the voice interaction method as described in the thirteenth aspect.
[0052] Sixteenth aspect, an embodiment of the present invention provides a non-transitory machine-readable storage medium, on which an executable code is stored. When the executable code is executed by a processor of a voice interaction device, the processor can at least implement the voice interaction method as described in the thirteenth aspect.
[0053] The voice recognition solution provided by the embodiments of the present invention can be applied to the voice recognition processing of streaming long voice signals.
[0054] Embodiments of the present invention provide an end-to-end voice recognition model for realizing real-time voice recognition processing of streaming long voice signals. The voice recognition model at least includes an encoding network and a decoding network. The voice recognition process based on this voice recognition model is as follows: Assume that a voice signal block (referred to as the first voice signal block) in the currently received voice signal stream is received. In order to recognize the text contained in the first voice signal block, first, at least one voice signal block included in the target group corresponding to the first voice signal block is obtained, and the at least one voice signal block includes the first voice signal block. Simply put, the target group includes each voice signal block that has been generated historically and belongs to the same group as the first voice signal block. After that, the at least one voice signal block included in the target group is encoded by the encoding network to obtain a semantic vector corresponding to the first voice signal block, so that the semantic vector includes historical information (each voice signal block generated before the first voice signal block in the target group) within a wider view corresponding to the first voice signal block. After that, the decoding network decodes the semantic vectors corresponding to each of the at least one voice signal block to obtain the text contained in the first voice signal block. It can be understood that during the decoding process, historical information within a wider view corresponding to the first voice signal block is also used to assist in decoding the first voice signal block.
[0055] There are two concepts involved in the above speech recognition processing: speech signal chunking and grouping composed of multiple speech signal chunks. Among them, the duration corresponding to the speech signal chunk can reflect the real-time degree of speech recognition. For example, 100 ms means a speech recognition delay of the order of 100 ms. In this way, whenever a speech signal chunk is generated, speech recognition processing is performed on the speech signal chunk in real time, which can ensure the real-time nature of speech recognition. The grouping duration restricts the maximum historical recall duration and is used to simulate the truncation of long speech, so as to complete the speech recognition of long speech by truncating the long speech. In this way, for a currently generated speech signal chunk, the relevant information of the speech signal chunks belonging to the same group can be used to assist in more accurately completing the speech recognition processing of the currently generated speech signal chunk and ensuring the accuracy of the speech recognition result. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0057] Figure 1 It is a schematic diagram of the composition of a speech recognition model provided by an embodiment of the present invention;
[0058] Figure 2 It is a schematic diagram of a speech signal chunk and grouping provided by an embodiment of the present invention;
[0059] Figure 3 It is a flowchart of a speech recognition method provided by an embodiment of the present invention;
[0060] Figure 4 It is a schematic diagram of the composition of another speech recognition model provided by an embodiment of the present invention;
[0061] Figure 5a It is a flowchart of a speech recognition method provided by an embodiment of the present invention;
[0062] Figure 5b It is a schematic diagram of the working process of a speech recognition model provided by an embodiment of the present invention;
[0063] Figure 6 It is a schematic diagram of the scenario of a speech interaction method provided by an embodiment of the present invention;
[0064] Figure 7a It is a flowchart of a speech interaction method provided by an embodiment of the present invention;
[0065] Figure 7bA schematic diagram of the scenario of a voice interaction method provided by an embodiment of the present invention;
[0066] Figure 8a A flowchart of a voice interaction method provided by an embodiment of the present invention;
[0067] Figure 8b A schematic diagram of the scenario of a voice interaction method provided by an embodiment of the present invention;
[0068] Figure 9 A schematic structural diagram of a voice recognition device provided by an embodiment of the present invention;
[0069] Figure 10 For Figure 9 A schematic structural diagram of an electronic device corresponding to the voice recognition device provided by the embodiment shown;
[0070] Figure 11 A schematic structural diagram of a voice interaction device provided by an embodiment of the present invention;
[0071] Figure 12 For Figure 11 A schematic structural diagram of a voice interaction device corresponding to the voice interaction device provided by the embodiment shown;
[0072] Figure 13 A schematic structural diagram of a voice interaction device provided by an embodiment of the present invention;
[0073] Figure 14 For Figure 13 A schematic structural diagram of a voice interaction device corresponding to the voice interaction device provided by the embodiment shown;
[0074] Figure 15 A schematic structural diagram of a voice interaction device provided by an embodiment of the present invention;
[0075] Figure 16 For Figure 15 A schematic structural diagram of a voice interaction device corresponding to the voice interaction device provided by the embodiment shown. Detailed implementation manners
[0076] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0077] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "said", and "the" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. "Plural" generally includes at least two.
[0078] Depending on the context, the words "if" and "when" as used herein may be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrases "if determined" or "if detected (stated condition or event)" may be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".
[0079] In addition, the step timings in the following method embodiments are only examples and not strictly limited.
[0080] The speech recognition method provided by the embodiments of the present invention can be executed by an electronic device, which can be a terminal device such as a PC, a laptop, a smart phone, a smart robot, etc., or a server in the cloud. The server can be a physical server including an independent host, or can also be a virtual server, or can also be a cloud server.
[0081] The speech recognition method provided by the embodiments of the present invention is applicable to the speech recognition processing of streaming long speech signals. Simply put, it is applicable to the speech recognition processing of real-time generated long speech signal streams.
[0082] It should be noted that the applicability to the speech recognition processing of long speech signals does not limit that the speech recognition method is only applicable to long speech signals, but can be generally applicable to any speech signal stream. However, compared with other traditional speech recognition schemes, this scheme will have better recognition effects in the recognition of long speech.
[0083] In order to achieve the real-time speech recognition processing of streaming long speech signals, the embodiments of the present invention provide an end-to-end speech recognition model, such as Figure 1 shown. The speech recognition model at least includes an encoding network and a decoding network. The model will be introduced in detail below.
[0084] In addition, to achieve the real-time speech recognition processing of streaming long speech signals, the embodiments of the present invention also define the following two concepts: speech signal chunking and a group composed of multiple speech signal chunks.
[0085] Among them, the voice signal block is defined by the preset voice signal block duration, and the group is defined by the preset group duration. For example, if the preset voice signal block duration is 100 milliseconds (ms), and the preset group duration is 9 seconds (s), then from the moment the voice signal starts to be output, a voice signal block is intercepted every 100ms, and the 90 voice signal blocks generated in sequence within 9s constitute a group. In other words, the maximum number of voice signal blocks that can be included in a group is determined by the preset group duration and the voice signal block duration: 9s / 100ms=90.
[0086] The duration of voice signal segmentation and grouping can be preset according to the actual application scenario. For example, in a scenario that is particularly sensitive to delay, the duration of voice signal segmentation can be set to be slightly shorter, while in a scenario that is not very sensitive to delay, the duration of voice signal segmentation can be set to be slightly longer.
[0087] In order to more intuitively understand the meaning of the above speech signal block and group, combined with Figure 2 To illustrate by way of example.
[0088] exist Figure 2 In the example, it is assumed that the user starts to output the voice signal stream from time T0=0s, and it is assumed that the preset voice signal block length is 100ms and the preset group length is 9s. Based on this assumption, it can be known that a group will contain 90 voice signal blocks. Figure 2 As shown in , starting from the moment T0=0s, when the timing reaches 100ms, the voice signal generated within this 100ms is intercepted as the first voice signal block, and the first group is created, and the first voice signal block is marked as belonging to the first group. After that, every time the timing reaches 100ms, the voice signal generated within the current 100ms can be intercepted as a voice signal block, and it is marked as belonging to the first group, until the 90th voice signal block is obtained, the 90th voice signal block is marked to the first group, and the first group reaches the capacity limit, at this time, the first group is formed. Figure 2 After that, when the timing continues for 100ms and the 91st speech signal block is obtained, a second group is created and the 91st block is marked as belonging to the second group, that is, the 91st speech signal block is the first speech signal block in the second group, and so on. Figure 2 When the block 180 shown in FIG. 1 is reached, the 90 speech signal blocks consisting of blocks 91 to 180 will reach the upper limit of the capacity of the second group.
[0089] From the above-described generation method of voice signal chunking and grouping, it can be known that: First, there are no duplicate voice signal chunks between different groups; Second, for example, when the user stops outputting voice signals at a certain moment, at this time, the number of voice signal chunks contained in the last generated group is less than or equal to 90.
[0090] In the embodiments of the present invention, the functions of defining the above two concepts are as follows: The duration of the voice signal chunk can reflect the real-time degree of speech recognition. For example, 100 ms means that there is a speech recognition delay of the order of 100 ms. In this way, every time a voice signal chunk is generated, speech recognition processing is performed on the voice signal chunk in real time, which can ensure the real-time nature of speech recognition. The grouping duration restricts the maximum historical recall duration and is used to simulate the truncation of long speech so as to complete the speech recognition of long speech by truncating the long speech. In this way, for a currently generated voice signal chunk, the relevant information of the voice signal chunks belonging to the same group as it can be used to assist in more accurately completing the speech recognition processing of the currently generated voice signal chunk and ensuring the accuracy of the speech recognition result.
[0091] That is to say, when the voice signal stream to be subjected to speech recognition is long, based on the above setting of the grouping duration, the speech recognition model can simulate the effect of truncating the long speech into short speech restricted by the grouping duration for speech recognition processing of the short speech, so that the speech recognition model can still obtain a good recognition effect similar to that of short speech when facing the recognition task of long speech.
[0092] The following describes the specific speech recognition process in conjunction with some embodiments.
[0093] Figure 3 It is a flowchart of a speech recognition method provided by an embodiment of the present invention. As Figure 3 shown, the method includes the following steps:
[0094] 301. Based on a preset voice signal chunk duration, obtain a first voice signal chunk in the voice signal stream.
[0095] 302. Obtain at least one voice signal chunk included in the target group corresponding to the first voice signal chunk. The target group includes the first voice signal chunk, and a group includes the voice signal chunks generated in sequence within the preset grouping duration.
[0096] For the speech recognition method provided in this embodiment, optionally, it can be executed by a server in the cloud. At this time, the terminal device on the user side can collect the voice signal stream output by the user and intercept voice signal chunks one by one based on the above voice signal chunk duration (such as 100 ms), and upload the intercepted voice signal chunks to the server in real time.
[0097] Combined with Figure 2 Taking the example in Figure 2 , assume that the user starts to output a voice signal stream from time T0. Then, whenever a voice signal chunk is obtained, voice recognition processing will be immediately performed on this voice signal chunk. Therefore, the above first voice signal chunk can be a certain voice signal chunk currently obtained. Since the voice recognition processes for each voice signal chunk are similar, in the embodiments of the present invention, only the voice recognition process for one voice signal chunk is taken as an example for illustration, and this voice signal chunk is the first voice signal chunk.
[0098] During the process of performing voice recognition processing on the first voice signal chunk, first, it is necessary to obtain each voice signal chunk generated historically that belongs to the same group as this first voice signal chunk, that is, at least one voice signal chunk included in the target group corresponding to the first voice signal chunk.
[0099] For example, assume that the first voice signal chunk is Figure 2 the voice signal chunk 1 shown in Figure 2 . Since this voice signal chunk 1 is the first voice signal chunk in the first group and there is no voice signal chunk generated before the voice signal chunk 1 in the first group, therefore, at this time, the acquisition result of at least one voice signal chunk included in the target group (i.e., the first group) corresponding to the first voice signal chunk is only the first voice signal chunk, and at this time, it is equivalent to only needing to perform the subsequent steps of processing on the first voice signal chunk.
[0100] Assume that the first voice signal chunk is Figure 2 the voice signal chunk 10 shown in Figure 2 . Since this voice signal chunk 10 belongs to the first group and voice signal chunks 1 to 9 have been generated before the voice signal chunk 10 in this first group, therefore, at this time, the acquisition result of at least one voice signal chunk included in the target group (i.e., the first group) corresponding to the first voice signal chunk includes the voice signal chunk 10 as the first voice signal chunk and voice signal chunks 1 to 9 that have been generated before the first voice signal chunk in the target group. For the convenience of description, hereinafter, the voice signal chunks in the target group other than the first voice signal chunk are called second voice signal chunks. Thus, in addition to the first voice signal chunk, the target group may also include at least one second voice signal chunk.
[0101] Assume again that the first voice signal chunk is Figure 2The voice signal block 91 shown in the figure. Since the voice signal block 91 belongs to the second group and is the first voice signal block within the second group, there is no voice signal block generated before the voice signal block 91 within the second group. Therefore, at this time, the acquisition result of at least one voice signal block included in the target group (i.e., the second group) corresponding to the first voice signal block is only the first voice signal block.
[0102] As can be seen from the above examples: First, after a voice signal block is generated, the relevant information of the voice signal block can be stored (the relevant information can be the voice signal block itself, or can include the acoustic features and semantic vectors of the voice signal block mentioned below), so as to query and obtain the above-mentioned second voice signal block based on the storage result. Second, how long the relevant information of the above voice signal block needs to be stored is restricted by the remaining capacity within its target group. When the remaining capacity of the target group is 0, the relevant information of each voice signal block belonging to the target group that has been stored is deleted.
[0103] For example, assume that the first voice signal block received currently is the voice signal block 10 within the first group. This means that the relevant information of voice signal blocks 1 to 9 has been stored previously. At least one second voice signal block that can be queried from the storage record is: voice signal blocks 1 to 9, and moreover, the relevant information of voice signal block 10 can also be stored. Another example, assume that the first voice signal block received currently is the voice signal block 90 within the first group. This means that the relevant information of voice signal blocks 1 to 89 has been stored previously. At least one second voice signal block that can be queried from the storage record is: voice signal blocks 1 to 89. At this time, after completing the speech recognition processing of voice signal block 90 based on the above steps, since the first group has reached the capacity limit of 90, that is, there is no remaining capacity, the relevant information of each voice signal block included in the first group that has been stored can be deleted.
[0104] 303. Encode the at least one voice signal block through an encoding network in the speech recognition model to obtain a semantic vector corresponding to the first voice signal block.
[0105] Optionally, the encoding network can be implemented using a multi-layer neural network, which can be any of the following: Feedforward Sequential Memory Networks (FSMN), Deep Feedforward Sequential Memory Networks (DFSMN), Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM), Bi-directional Long Short-Term Memory (BiLSTM), and so on.
[0106] As described above, the at least one speech signal chunk may only contain the first speech signal chunk (corresponding to the case where the first speech signal chunk is the first speech signal chunk within the target group), or may include the first speech signal chunk and at least one second speech signal chunk (corresponding to the case where the first speech signal chunk is not the first speech signal chunk within the target group).
[0107] For the above two cases, the encoding methods for the at least one speech signal chunk are different.
[0108] Specifically, when the target group includes at least one second speech signal chunk, encoding the at least one speech signal chunk through the encoding network in the speech recognition model to obtain a semantic vector corresponding to the first speech signal chunk can be achieved as follows:
[0109] Obtain the acoustic features of the first speech signal chunk and the acoustic features of the at least one second speech signal chunk;
[0110] Encode the concatenation result of the acoustic features of the first speech signal chunk and the acoustic features of the at least one second speech signal chunk through the encoding network to obtain a semantic vector corresponding to the first speech signal chunk.
[0111] Among them, the acoustic features can be any of the following: Mel-Frequency Cepstral Coefficients (MFCC), Linear Predictive Cepstral Coefficient (LPCC), short-term average energy, amplitude average change rate, Fbank features, and so on.
[0112] Taking the first speech signal block as an example, the process of obtaining the acoustic features of any speech signal block can be implemented as follows: performing frame segmentation on the first speech signal block to obtain multiple frames of speech signals; extracting the acoustic features corresponding to each of the multiple frames of speech signals, and determining the acoustic features of the first speech signal block according to the acoustic features corresponding to each of the multiple frames of speech signals.
[0113] Among them, the duration of each frame of speech signal can be preset, for example, preset to 10 ms. Assuming that the duration of the speech signal block is 100 ms, this means that a speech signal block can be divided into 10 frames of speech signals. After obtaining the acoustic features corresponding to each of the multiple frames of speech signals respectively, the obtained acoustic features are concatenated together as the acoustic features of the first speech signal block.
[0114] By sending the acoustic features of the first speech signal block and at least one second speech signal block belonging to the same group as the first speech signal block into the encoding network for encoding, the semantic vector corresponding to the first speech signal block output by the encoding (for ease of description, referred to as the first semantic vector) can contain historical semantics within a larger field of view, which can help with the accurate recognition of the text contained in the first speech signal block. Among them, the historical semantics within the larger field of view refers to being restricted by the maximum historical recall duration defined by the group duration, that is, the semantic information of each speech signal block that has been generated within the same group can be used to assist in the speech recognition processing of the currently generated first speech signal block, which helps to obtain a more accurate recognition result.
[0115] In the above-described solution, since the grouping defined based on the preset group duration can truncate the long speech, so that the speech recognition model can only focus on a limited historical recall duration to perform speech recognition processing on the currently generated speech signal block. It can be considered that different groups are independent of each other, that is, the speech processing process of the speech signal block in the next group will not use the relevant information of the speech signal block in the previous group.
[0116] However, in practical applications, the following situation may be encountered: at least one frame of speech signal containing a certain word spoken by the user may just be divided into two groups, that is, a part is in the previous group and the other part is in the next group, which will result in discontinuous information between different groups. To ensure the continuity of information between different groups, the following optional solution is provided in the embodiments of the present invention:
[0117] The configured encoding network is an encoding network with a memory block for memorizing the acoustic features of at least one frame of speech signal generated last within the previous packet, where the at least one frame of speech signal is located within at least one speech signal block generated last within the previous packet. The principle and structural composition of the memory block can be implemented with reference to existing related technologies and will not be elaborated here.
[0118] For example, assuming that a speech signal block includes 10 frames of speech signals and a packet can include 100 speech signal blocks, the above-mentioned memory block can be set to memorize, for example, 12 frames of speech signals generated last within the previous packet. Based on this, assuming that the speech processing of the first speech signal block in packet i is currently in progress, these 12 frames of speech signals are the 10 frames of speech signals included in the 100th speech signal block in packet (i - 1) and the last two frames of speech signals included in the 99th speech signal block.
[0119] Continuing with the above example, at this time, assuming that the first speech signal block in the next packet, i.e., speech signal block 101, is generated, the encoding network can obtain the semantic vector corresponding to speech signal block 101 through the following process:
[0120] The encoding network encodes the concatenation result of the acoustic features of the at least one frame of speech signal currently memorized (such as the 12 frames of speech signals in the above example) and the acoustic features of speech signal block 101 (i.e., the acoustic features of the 10 frames of speech signals included in speech signal block 101) to obtain the semantic vector corresponding to speech signal block 101, thereby ensuring the continuity of information between adjacent packets.
[0121] Briefly summarized, if the current first speech signal block is the first speech signal block within its target packet, then when encoding it, the acoustic features of a set number of speech signal frames generated last within the previous packet can be incorporated. The set number of speech signal frames may correspond to one speech signal block or several speech signal blocks. 304. Decode the semantic vectors corresponding to the at least one speech signal block through the decoding network in the speech recognition model to output the text corresponding to the first speech signal block.
[0122] It should be noted that in this embodiment, after obtaining the first semantic vector corresponding to the first speech signal block by encoding the acoustic features corresponding to each of the at least one speech signal block, during the decoding process, not only the first semantic vector is decoded, but in the case where the at least one speech signal block is composed of the first speech signal block and at least one second speech signal block, the first semantic vector corresponding to the first speech signal block and the semantic vectors (for ease of description, referred to as second semantic vectors) corresponding to each of the at least one second speech signal block are decoded to obtain the text included in the first speech signal block.
[0123] Among them, the process of obtaining the second semantic vector corresponding to each second speech signal block is similar to the process of obtaining the first semantic vector corresponding to the first speech signal block, and will not be elaborated.
[0124] It can be seen that during the decoding process, the historical information within a wider view corresponding to the first speech signal block is also utilized to assist in decoding the first speech signal block, so as to obtain a more accurate recognition result.
[0125] In another alternative embodiment, in order to implement real-time speech recognition processing for streaming long speech signals, the embodiment of the present invention further provides an end-to-end speech recognition model, as Figure 4 shown. The speech recognition model at least includes an encoding network, a prediction network, an attention network, and a decoding network.
[0126] Based on Figure 4 the speech recognition model shown, there is also provided a speech recognition solution as Figure 5a shown, which may include the following steps:
[0127] 501. Based on a preset speech signal block duration, obtain the first speech signal block in the speech signal stream.
[0128] 502. Obtain at least one speech signal block included in the target group corresponding to the first speech signal block. The target group includes the first speech signal block, and a group includes the speech signal blocks generated in sequence within a preset group duration.
[0129] 503. Encode the at least one speech signal block through the encoding network in the speech recognition model to obtain the semantic vector corresponding to the first speech signal block.
[0130] 504. Predict the first semantic vector through the prediction network in the speech recognition model to obtain the number of words included in the first speech signal block.
[0131] 505. Determine the weights of the semantic vectors corresponding to the at least one speech signal block and the weighted sum result in a weight calculation process through an attention network in a speech recognition model, and input the weighted sum result into a decoding network to output text corresponding to the weighted sum result through the decoding network, where the number of text is used to constrain the number of weight calculation times.
[0132] In the embodiments of the present invention, the prediction network, the attention network, and the decoding network can also be composed of multi-layer neural networks.
[0133] The prediction network is used to determine the number of text included in the speech signal block, so as to control the work of the attention network based on the guidance of the number of text.
[0134] Specifically, after obtaining the semantic vector corresponding to the first speech signal block, input the semantic vector into the prediction network, and the prediction network outputs a corresponding prediction sequence. The prediction sequence is used to indicate whether each of the multiple frames of speech signals included in the first speech signal block corresponds to text. Thus, according to the number of frames of speech signals corresponding to text in the prediction sequence, the number of text included in the first speech signal block can be determined.
[0135] For example, assume that the first speech signal block includes 10 frames of speech signals. Then the prediction sequence output by the prediction network is a sequence with a total length of 10 composed of 0 and 1, where 0 indicates that there is no text in the corresponding frame of speech signal, and 1 indicates that there is text in the corresponding frame of speech signal. For example, if the output prediction sequence is [0, 0, 0, 1, 1, 0, 0, 0, 0, 0], it means that there is text in the fourth and fifth frames of speech signals among these 10 frames of speech signals, that is, the first speech signal block includes two text.
[0136] Briefly speaking, the prediction network can be trained in the following way: obtain several speech signal block samples, label the number of text included in each speech signal block sample, and use this labeled information to train the prediction network. The trained prediction network thus has the function of predicting the number of text included in the speech signal block.
[0137] Generally speaking, the guiding effect of the number of text output by the above prediction network on the attention network is mainly reflected in: constraining the number of times the attention network calculates the weights of the semantic vectors corresponding to the at least one speech signal block.
[0138] Under the guidance of the number N (N is greater than or equal to 1) of text included in the first speech signal block predicted by the prediction network, the specific working process of the attention network is as follows:
[0139] Initialize the number of weight calculation times to N, and iteratively execute the following process until the number of weight calculation times is reduced to 0:
[0140] If the current weight calculation count is not 0, input the semantic vectors corresponding to each of the at least one speech signal block and the previous character output by the decoding network into the attention network, so that the attention network determines the weights of the semantic vectors corresponding to each of the at least one speech signal block based on the previous character, and determines the weighted sum result of the semantic vectors corresponding to each of the at least one speech signal block according to the weights, and input the weighted sum result into the decoding network;
[0141] Obtain the currently output character through the decoding network, and input the character into the attention network;
[0142] Decrease the weight calculation count by one.
[0143] In summary, in the embodiment of the present invention, the prediction network outputs the number of characters included in a certain speech signal block, which can be used to guide the attention network to determine which of the previous semantic vectors need to be noted during the process of decoding the current character. At the same time, the historical character (the previous character) finally output by the decoding network will also be used as the input of the attention network to predict the current character.
[0144] Specifically, assume that the current first speech signal block is Figure 2 the speech signal block 10 shown in, then the at least one speech signal block includes the previously generated speech signal blocks 1 to 9 and the speech signal block 10, and assume that the semantic vector corresponding to the speech signal block 10 is represented as C10, and the semantic vectors corresponding to the speech signal blocks 1 to 9 are represented as C1 to C9 respectively. Assume that the prediction network determines that there are two characters in the speech signal block 10. Since the speech signal blocks included in the group where the speech signal block 10 is located are the speech signal blocks 1 to 9, therefore, since the prediction result input by the prediction network to the attention network is the prediction result corresponding to the speech signal block 10, it makes the attention network need to pay attention to the semantic vectors corresponding to each of the speech signal blocks that have been generated within the group to which the speech signal block 10 belongs: C1 to C9. In addition, assume that when the prediction network performs prediction processing on C3, it determines that there are no characters in the corresponding speech signal block. Optionally, at this time, the attention network can only pay attention to the semantic vectors corresponding to the speech signal blocks 1, 2, and 4 to 9, and does not need to pay attention to the semantic vector corresponding to the speech signal block 3. That is to say, the attention network can only pay attention to the semantic vectors corresponding to the speech signal blocks that contain characters within the current group.
[0145] Assume that determining attention requires focusing on C1 to C9, and assume that the previous word output by the decoding network is "big". Then, under the guidance of the prediction network outputting a speech signal chunk containing two words in chunk 10, the attention network first performs a weight calculation on C1 to C9 based on the previous word ("big") input to the decoding network. Based on the weight calculation results, a weighted sum of C1 to C9 is performed, and the semantic vector obtained from the weighted sum is output to the decoding network. Assume that the current word decoded and output by the decoding network based on the current input is: "home". At this time, since the attention network has completed one weight calculation, the weight calculation count is updated to 1. Additionally, at this time, the output "home" of the decoding network is input to the attention network. Since the weight calculation count is not 0, the attention network continues to perform a weight calculation on C1 to C9 based on the previous word output by the decoding network (which is "home" at this time). Based on the weight calculation results, a weighted sum of C1 to C9 is performed, and the semantic vector obtained from the weighted sum is output to the decoding network. Assume that the current word decoded and output by the decoding network based on the current input is: "good". At this time, since the attention network has completed another weight calculation, the weight calculation count is updated to 0. Additionally, at this time, the output "good" of the decoding network is input to the attention network. Since the weight calculation count is 0, the attention network pauses the weight calculation process and waits for the arrival of the semantic vector and the predicted result of the number of words corresponding to the next speech signal chunk 11.
[0146] For ease of understanding, the following combines Figure 5b the speech recognition model shown in
[0147] In Figure 5b , assume that the currently obtained first speech signal chunk is speech signal chunk m. Within the target group to which the first speech signal chunk belongs, speech signal chunks 1 to m - 1 have been generated historically. Assume that the acoustic feature representations corresponding to speech signals 1 to speech signal chunk m are: X1 to X m . Assume that the semantic vector representations corresponding to speech signals 1 to speech signal chunk m are: C1 to C m . Additionally, assume that the previous word output by the decoding network is represented as Y t-1 , and the currently output word is represented as Y t .
[0148] Based on the above assumptions, when speech signal chunk m is generated, the acoustic feature X m corresponding to speech signal chunk m is extracted, and the acoustic features X1 to X m-1 corresponding to the historical speech signal chunks - speech signal chunks 1 to m - 1 belonging to the same group as it are obtained. After that, the acoustic features X1 to X mThe input encoding network outputs a semantic vector C corresponding to the m-th block of the speech signal m 。
[0149] The semantic vector C m is input into the prediction network, and the prediction network outputs a corresponding prediction sequence, assumed to be P m =[0, 0, 0, 1, 1, 0, 0, 0, 0, 0], indicating that there are words in the fourth and fifth frames of the speech signal in the m-th block of the speech signal, that is, the m-th block of the speech signal includes two words.
[0150] This prediction sequence, the previous word Y output by the decoding network t-1 and the semantic vectors C1 to C m are all input into the attention network. Under the guidance of this prediction sequence, the attention network calculates weights for the semantic vectors C1 to C m based on the previous word output by the decoding network, and inputs the weighted sum result of the semantic vectors obtained based on the calculated weights into the decoding network. Among them, the guiding role of the prediction sequence and the calculation process of the attention network refer to the relevant descriptions in the above text and will not be elaborated here.
[0151] From the working process of the speech recognition model introduced above, it can be seen that during the training of this speech recognition model, two loss functions are required: one is the loss function corresponding to the prediction network, and the other is the loss function corresponding to the final output result of the entire model, that is, the loss function corresponding to the output of the decoding network.
[0152] The prediction network can be trained simultaneously while training the speech recognition model. The function of the prediction network is to predict how many words are included in each block of the speech signal. Therefore, the training data during the training process can include a large number of speech signal samples and word quantity annotation information. Among them, the speech signal samples can include samples with a duration greater than the duration threshold corresponding to the long speech, and can also include samples with a duration less than the duration threshold corresponding to the long speech. Among them, the acquisition of the word quantity annotation information can be: performing block cutting on the speech signal samples to obtain multiple speech signal blocks, and annotating the number of words included in each speech signal block, and this annotation can be implemented manually.
[0153] In addition, in an optional embodiment, if it is predicted that there are no words in a continuous set number of speech signal blocks within a group, then this group can be terminated at this time, and the historical information can be cleared. Specifically, the starting speech signal block of the new group is re-determined as the next speech signal block containing words, and the relevant information of each speech signal block within each previous group generated by this new group is cleared, such as acoustic features and semantic vectors.
[0154] Among them, the set quantity is, for example, one or more. Taking the first voice signal chunk in the above text as an example, when it is determined based on the prediction result of the prediction network that the first voice signal chunk does not contain text, or if it is determined that none of the consecutive preset quantity of voice signal chunks starting from the first voice signal chunk contain text, a new group is generated, the starting voice signal chunk of the new group is determined to be the next voice signal chunk containing text, and the historical storage information corresponding to each previous group is deleted. Among them, the historical storage information includes the acoustic features and semantic vectors corresponding to the voice signal chunks, and among them, the above consecutive preset quantity of voice signal chunks all belong to this group.
[0155] For example, assume that the first voice signal chunk is voice signal chunk 5, which belongs to the first group, and assume that the preset setting is that if two consecutive voice signal chunks do not contain text, the group needs to be terminated, and assume that a group includes 100 voice signal chunks. Assume that the prediction network predicts that voice signal chunk 5 does not contain text, then at this time the attention network will not perform the weight calculation work. Assume that the next arriving voice signal chunk 6 also does not contain text, then it is determined that a second group needs to be created. Specifically, assume that the next arriving voice signal chunk 7 also does not contain text, then continue to wait for the next voice signal chunk. If the next arriving voice signal chunk 8 contains text, then it is determined that the starting voice signal chunk of the second group is voice signal chunk 8. In this way, voice signal chunks 8 to 108 may form a group.
[0156] Since none of the consecutive set quantity of voice signal chunks contain text, it often means that it is in a silent segment at this time, that is, the user may have finished speaking a paragraph before and there is a long pause in the middle. To reduce the computational load and avoid the accumulation of errors in the previous speech recognition model, the above processing of terminating the current group can be performed.
[0157] The speech recognition method provided in the above embodiments of the present invention can be applied to an architecture composed of a client and a server. Among them, the client runs in the terminal device on the user side, and the server can be composed of a server or a server cluster in the cloud, and the server provides a speech recognition service. In practical applications, according to different application scenarios, the client will change accordingly. For example, in a conference application scenario, the client can be a client providing conference-related functions; in a live broadcast scenario, the client can be a host client.
[0158] Under the architecture of the client and the server, the embodiments of the present invention provide a voice interaction method, which can be executed by the client. The voice interaction method may include the following steps:
[0159] Collect voice signal chunks in the voice signal stream, and each voice signal chunk has a preset duration;
[0160] Chunk the collected voice signals and upload them to the server, so that the server can obtain the first voice signal chunk and at least one voice signal chunk included in the target group corresponding to the first voice signal chunk, and encode the at least one voice signal chunk through the encoding network in the voice recognition model to obtain a semantic vector corresponding to the first voice signal chunk; decode the semantic vectors corresponding to the at least one voice signal chunk respectively through the decoding network in the voice recognition model to output the text corresponding to the first voice signal chunk; wherein, the target group includes the first voice signal chunk, and one group includes the voice signal chunks generated sequentially within a preset group duration.
[0161] Display the text received from the server.
[0162] As described above, when the target group includes at least one second voice signal chunk, the process of obtaining the semantic vector corresponding to the first voice signal chunk includes:
[0163] Obtain the acoustic features of the first voice signal chunk and the acoustic features of at least one second voice signal chunk.
[0164] Encode the concatenation result of the acoustic features of the first voice signal chunk and the acoustic features of at least one second voice signal chunk through the encoding network to obtain a semantic vector corresponding to the first voice signal chunk.
[0165] When the target group only includes the first voice signal chunk, the process of obtaining the semantic vector corresponding to the first voice signal chunk includes:
[0166] Encode the concatenation result of the acoustic features of at least one frame of voice signal currently memorized and the acoustic features of the first voice signal chunk through the encoding network to obtain a semantic vector corresponding to the first voice signal chunk.
[0167] Wherein, the encoding network is an encoding network with a memory module, and the memory module is used to memorize the acoustic features of at least one frame of voice signal generated last in the previous group, and the at least one frame of voice signal is within at least one voice signal chunk generated last in the previous group.
[0168] The following combines Figure 6 to exemplarily illustrate the execution process of the voice interaction method provided by the embodiments of the present invention under the above-mentioned client-server architecture.
[0169] In Figure 6Among them, it is assumed that user A is using a certain client. During the use of the client, user A will output a voice signal stream. The client collects the voice signal chunks in the voice signal stream output by user A and uploads each sequentially collected voice signal chunk to the server in real time. Among them, each voice signal chunk has a preset duration. For the convenience of description, it is assumed that the output voice of user A is intercepted into voice signal chunks, and V1 to V m These m voice signal chunks are intercepted, and it is assumed that these m voice signal chunks belong to the same group.
[0170] Based on the above assumed situation, as Figure 6 shown, whenever the client collects a voice signal chunk, it uploads it to the server. Let V i represent any voice signal chunk collected, and V i is one of the above m voice signal chunks. There is a pre-trained speech recognition model maintained in the server. Among them, the speech recognition model includes an encoding network and a decoding network, which are respectively represented as: a streaming encoder and a streaming decoder in Figure 6 .
[0171] When the server receives the voice signal chunk V i , it extracts its corresponding acoustic feature X i , and obtains the acoustic features corresponding to the other already generated voice signal chunks within the same group as the voice signal chunk V i . In Figure 6 , it is assumed that the other already generated voice signal chunks within the same group as the voice signal chunk V i are: voice signal chunk V i-1 , voice signal chunk V i-2 . And it is assumed that the acoustic features corresponding to the voice signal chunk V i-1 , voice signal chunk V i-2 are respectively represented as X i-1 , X i-2 .
[0172] All the acoustic features X i , X i-1 , X i-2 are sent to the streaming encoder for encoding to obtain the semantic vector corresponding to the voice signal chunk V i , represented as C i .
[0173] After that, the server obtains the voice signal chunks V i-1 , V i-2 that have been obtained during the speech recognition process of the voice signal chunks V i-1 , V i-2Their corresponding semantic vectors, denoted as: C i-1 and C i-2 .
[0174] The streaming decoder obtains the text contained in the speech signal block V i and C i-1 and C i-2 through decoding processing. i The text contained in the speech signal block V
[0175] In Figure 6 , assume that the texts contained in the m speech signal blocks V1 to V m are respectively: "It's an honor", "to", "be invited", "to", "explain", "blockchain" for each of these m speech signal blocks. In this assumption, the text corresponding to the above speech signal block V i-2 can be "It's an honor", the text corresponding to the speech signal block V i-1 can be "to", the text corresponding to the speech signal block V i can be "be invited", and so on.
[0176] After obtaining each speech signal block, the server performs real-time speech recognition processing on it to obtain the corresponding text and feeds it back to the client, and the client displays the received text in real time.
[0177] For other component structures of the speech recognition model and the detailed execution process of the server, reference can be made to the relevant descriptions in the foregoing other embodiments.
[0178] It should be noted that the steps performed by the above server can also be transferred to the local client when the client has sufficient computing resources.
[0179] The speech interaction method provided by the embodiments of the present invention above can be applied to any scenario that requires speech recognition of a speech signal stream. For example, in a meeting scenario or a live broadcast scenario.
[0180] Taking the meeting scenario as an example, the embodiments of the present invention can provide a speech interaction method applicable to the meeting scenario, as Figure 7a shown, this speech interaction method may include the following steps:
[0181] 701. Obtain a first speech signal block in the meeting speech signal stream based on a preset speech signal block duration.
[0182] 702. Obtain at least one speech signal block contained in the target group corresponding to the first speech signal block, where the target group includes the first speech signal block, and a group includes speech signal blocks generated in sequence within a preset group duration.
[0183] 703. Encode the at least one voice signal in chunks through an encoding network in the speech recognition model to obtain a semantic vector corresponding to the first voice signal chunk.
[0184] 704. Decode the semantic vectors corresponding to the at least one voice signal chunk respectively through a decoding network in the speech recognition model to output the text corresponding to the first voice signal chunk.
[0185] 705. Display the text.
[0186] Optionally, the method may further include: generating a meeting record according to the text.
[0187] Optionally, when the target group includes at least one second voice signal chunk, the process of obtaining the semantic vector corresponding to the first voice signal chunk includes:
[0188] Obtain the acoustic features of the first voice signal chunk and the acoustic features of the at least one second voice signal chunk;
[0189] Encode the concatenation result of the acoustic features of the first voice signal chunk and the acoustic features of the at least one second voice signal chunk through the encoding network to obtain a semantic vector corresponding to the first voice signal chunk.
[0190] When the target group only includes the first voice signal chunk, the process of obtaining the semantic vector corresponding to the first voice signal chunk includes:
[0191] Encode the concatenation result of the acoustic features of at least one frame of voice signal currently memorized and the acoustic features of the first voice signal chunk through the encoding network to obtain a semantic vector corresponding to the first voice signal chunk. Wherein, the encoding network is an encoding network with a memory module, and the memory module is used to memorize the acoustic features of at least one frame of voice signal generated last in the previous group, and the at least one frame of voice signal is within at least one voice signal chunk generated last in the previous group.
[0192] Optionally, in addition to the encoding network and the decoding network, the speech recognition model may further include a prediction network and an attention network. Based on this, decoding the semantic vectors corresponding to the at least one voice signal chunk respectively through the decoding network in the speech recognition model to output the text corresponding to the first voice signal chunk can be implemented as:
[0193] Predict the semantic vector corresponding to the first voice signal chunk through the prediction network in the speech recognition model to obtain the number of words contained in the first voice signal chunk;
[0194] Determine the weights of the semantic vectors corresponding to the at least one speech signal block and the weighted sum result in a single weight calculation process through the attention network in the speech recognition model, and input the weighted sum result into the decoding network to output the text corresponding to the weighted sum result through the decoding network, where the number of characters is used to constrain the number of weight calculation times.
[0195] Among them, the specific working processes of the prediction network, the attention network, and the decoding network can refer to the relevant descriptions in the foregoing other embodiments and will not be elaborated here.
[0196] As Figure 7b shown, in a conference scenario, the above speech interaction method can be executed by a speech interaction device, which can be a type of terminal device in a conference scenario, simply referred to as a conference terminal. When the conference terminal has sufficient computing power, the above solution can be fully executed locally on the conference terminal. When the computing power of the conference terminal is insufficient, the core steps can be executed by a server in the cloud. At this time, the conference terminal can only complete the functions of collecting and uploading the speech signal blocks in the speech signal stream and displaying and recording the text recognition results.
[0197] Suppose in a conference scenario, it is required to display the speech spoken by the speaker on the screen in real time, and in addition, text recording can also be performed. At this time, based on the solution provided in the embodiments of the present invention, the speech content of the speaker can be displayed on the screen in real time in an accompanying manner and the speech content can be recorded. For example, suppose that following the speech output of the speaker, based on the above speech recognition process, the words "It's an honor", "to", "be invited", "to", "explain", "blockchain" are successively displayed on the screen. And the sentence "It's an honor to be invited to explain blockchain" is written in the conference record file.
[0198] Taking the live broadcast scenario as an example, the embodiments of the present invention can provide a speech interaction method applicable to the live broadcast scenario, as Figure 8a shown, the speech interaction method may include the following steps:
[0199] 801. Based on a preset speech signal block duration, obtain a first speech signal block in the speech signal stream of the anchor.
[0200] 802. Obtain at least one speech signal block included in a target group corresponding to the first speech signal block, where the target group includes the first speech signal block, and a group includes the speech signal blocks generated successively within a preset group duration.
[0201] 803. Encode the at least one speech signal block through an encoding network in the speech recognition model to obtain a semantic vector corresponding to the first speech signal block.
[0202] 804. Decode the semantic vectors corresponding to each block of the at least one speech signal through a decoding network in the speech recognition model to output the text corresponding to the first speech signal block.
[0203] 805. Display the text in the live broadcast interface.
[0204] Optionally, when the target group includes at least one second speech signal block, the process of obtaining the semantic vector corresponding to the first speech signal block includes:
[0205] Obtain the acoustic features of the first speech signal block and the acoustic features of the at least one second speech signal block.
[0206] Encode the concatenation result of the acoustic features of the first speech signal block and the acoustic features of the at least one second speech signal block through an encoding network to obtain the semantic vector corresponding to the first speech signal block.
[0207] When the target group only includes the first speech signal block, the process of obtaining the semantic vector corresponding to the first speech signal block includes:
[0208] Encode the concatenation result of the acoustic features of at least one frame of the speech signal currently memorized and the acoustic features of the first speech signal block through an encoding network to obtain the semantic vector corresponding to the first speech signal block. Herein, the encoding network is an encoding network with a memory module, and the memory module is used to memorize the acoustic features of at least one frame of the speech signal generated last in the previous group, and the at least one frame of the speech signal is within at least one speech signal block generated last in the previous group.
[0209] Optionally, in addition to the encoding network and the decoding network, the speech recognition model may further include a prediction network and an attention network. Based on this, decoding the semantic vectors corresponding to each block of the at least one speech signal through the decoding network in the speech recognition model to output the text corresponding to the first speech signal block can be implemented as:
[0210] Predict the semantic vector corresponding to the first speech signal block through the prediction network in the speech recognition model to obtain the number of words included in the first speech signal block.
[0211] Determine the weights and the weighted sum result of the semantic vectors corresponding to each block of the at least one speech signal in one weight calculation process through the attention network in the speech recognition model, and input the weighted sum result into the decoding network to output the text corresponding to the weighted sum result through the decoding network, and the number of words is used to constrain the number of weight calculation times.
[0212] Among them, the specific working processes of the prediction network, the attention network, and the decoding network can refer to the relevant descriptions in the foregoing other embodiments, and will not be elaborated here.
[0213] As Figure 8b shown, the terminal device of the host can segment the speech signal stream spoken by the host into speech signal chunks, and transmit the obtained speech signal chunks to the server in the cloud in real time. The server displays the speech content of the host in the form of subtitles in the live broadcast interface. In this way, the audience who pulls the live video stream can see the speech content of the host through the live broadcast interface. Of course, the terminal device of the host can also directly upload the collected speech signal stream to the server in the cloud, and the server completes the segmentation and subsequent processing of the speech signal chunks. The detailed processing process of the server can refer to the relevant descriptions in the foregoing embodiments, and will not be elaborated here.
[0214] The speech recognition device and the speech interaction device of one or more embodiments of the present invention will be described in detail below. Those skilled in the art can understand that these speech recognition devices and speech interaction devices can be configured by using commercially available hardware components through the steps taught by this solution.
[0215] Figure 9 is a schematic structural diagram of a speech recognition device provided by an embodiment of the present invention. As Figure 9 shown, the device includes: an acquisition module 11, an encoding module 12, and a decoding module 13.
[0216] The acquisition module 11 is configured to obtain a first speech signal chunk in the speech signal stream based on a preset speech signal chunk duration; and obtain at least one speech signal chunk included in a target group corresponding to the first speech signal chunk. The target group includes the first speech signal chunk, and a group includes speech signal chunks generated sequentially within a preset group duration.
[0217] The encoding module 12 is configured to encode the at least one speech signal chunk through an encoding network in the speech recognition model to obtain a semantic vector corresponding to the first speech signal chunk.
[0218] The decoding module 13 is configured to decode the semantic vectors corresponding to the at least one speech signal chunk through a decoding network in the speech recognition model to output the text corresponding to the first speech signal chunk.
[0219] Optionally, at least one second speech signal chunk is included in the target group. At this time, the encoding module 12 is specifically configured to: obtain the acoustic features of the first speech signal chunk and the acoustic features of the at least one second speech signal chunk; encode the concatenation result of the acoustic features of the first speech signal chunk and the acoustic features of the at least one second speech signal chunk through the encoding network, so as to obtain a semantic vector corresponding to the first speech signal chunk.
[0220] Optionally, the apparatus further includes: a feature extraction module, configured to obtain the acoustic features of any speech signal chunk in the following manner: perform frame segmentation on the any speech signal chunk to obtain multiple frames of speech signals; extract the acoustic features corresponding to each of the multiple frames of speech signals; and determine the acoustic features of the any speech signal chunk according to the acoustic features corresponding to each of the multiple frames of speech signals.
[0221] Optionally, the encoding network is an encoding network with a memory module, and the memory module is configured to memorize the acoustic features of at least one frame of speech signal generated last in the previous group, and the at least one frame of speech signal is located in at least one speech signal chunk generated last in the previous group. Based on this, if the first speech signal chunk is the first speech signal chunk in the target group, the encoding module 12 may specifically be configured to: encode the concatenation result of the acoustic features of the at least one frame of speech signal currently memorized and the acoustic features of the first speech signal chunk through the encoding network, so as to obtain a semantic vector corresponding to the first speech signal chunk.
[0222] Optionally, the decoding module 13 may specifically be configured to: predict the number of characters included in the first speech signal chunk through the prediction network in the speech recognition model; determine the weights and the weighted sum result of the semantic vectors corresponding to the at least one speech signal chunk respectively in one weight calculation process through the attention network in the speech recognition model, and input the weighted sum result into the decoding network, so as to output, through the decoding network, the characters corresponding to the weighted sum result, and the number of characters is used to constrain the number of weight calculation times.
[0223] Optionally, the decoding module 13 may specifically be configured to: output, through the prediction network, a prediction sequence corresponding to the semantic vector corresponding to the first speech signal chunk, where the prediction sequence is used to indicate whether each of the multiple frames of speech signals included in the first speech signal chunk corresponds to a character; and the number of characters included in the first speech signal chunk is determined by the number of frames of speech signals corresponding to characters in the prediction sequence.
[0224] Among them, optionally, the first speech signal block contains N characters, where N is greater than or equal to 1. Specifically, the decoding module 13 may be configured to:
[0225] Initialize the weight calculation times to N, and iteratively execute the following process until the weight calculation times are reduced to 0:
[0226] If the current weight calculation times are not 0, input the semantic vectors corresponding to the at least one speech signal block and the previous character output by the decoding network into the attention network, so that the attention network determines the weights of the semantic vectors corresponding to the at least one speech signal block based on the previous character, and determines the weighted sum result of the semantic vectors corresponding to the at least one speech signal block according to the weights, and input the weighted sum result into the decoding network;
[0227] Obtain the currently output character through the decoding network, and input the character into the attention network;
[0228] Decrement the weight calculation times by one.
[0229] Optionally, the device may further include: a grouping processing module, configured to, if the first speech signal block does not contain characters, or if none of the consecutive preset number of speech signal blocks starting from the first speech signal block contain characters, use the next speech signal block containing characters as the starting point of a new group; delete the historical storage information corresponding to the new group, where the historical storage information includes the acoustic features and semantic vectors corresponding to each speech signal block in each group before the new group is generated.
[0230] Figure 9 The shown device may execute the speech recognition method provided in the foregoing Figures 1 to 5a For the detailed execution process and technical effects, refer to the description in the foregoing embodiments, which will not be elaborated here.
[0231] In a possible design, the structure of the foregoing Figure 9 shown speech recognition device may be implemented as an electronic device. As Figure 10 shown, the electronic device may include: a processor 21 and a memory 22. Among them, an executable code is stored on the memory 22. When the executable code is executed by the processor 21, the processor 21 can at least implement the speech recognition method provided in the foregoing Figures 1 to 5a shown embodiments.
[0232] Optionally, the electronic device may further include a communication interface 23 for communicating with other devices.
[0233] In addition, an embodiment of the present invention provides a non-transitory machine-readable storage medium, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor can at least implement the speech recognition method provided in the foregoing Figures 1 to 5a as shown in the embodiment.
[0234] Figure 11 FIG. is a schematic structural diagram of a voice interaction device provided by an embodiment of the present invention. As Figure 11 shown, the device includes: an acquisition module 31, a sending module 32, and a display module 33.
[0235] The acquisition module 31 is configured to acquire speech signal blocks in a speech signal stream, and each speech signal block has a preset duration.
[0236] The sending module 32 is configured to upload the acquired speech signal blocks to a server, so that the server obtains a first speech signal block and at least one speech signal block included in a target group corresponding to the first speech signal block, and encodes the at least one speech signal block through an encoding network in a speech recognition model to obtain a semantic vector corresponding to the first speech signal block; decodes the semantic vectors corresponding to the at least one speech signal block through a decoding network in the speech recognition model to output the text corresponding to the first speech signal block; wherein, the target group includes the first speech signal block, and a group includes speech signal blocks generated in sequence within a preset group duration.
[0237] The display module 33 is configured to display the text received from the server.
[0238] Optionally, the target group includes at least one second speech signal block; the process of obtaining the semantic vector corresponding to the first speech signal block includes:
[0239] Obtaining the acoustic features of the first speech signal block and the acoustic features of the at least one second speech signal block;
[0240] Encoding the concatenation result of the acoustic features of the first speech signal block and the acoustic features of the at least one second speech signal block through the encoding network to obtain a semantic vector corresponding to the first speech signal block.
[0241] In a possible design, the structure of the foregoing Figure 11 as shown in the voice interaction device can be implemented as a voice interaction device, such as Figure 12As shown in the figure, the voice interaction device may include: a processor 41, a memory 42, and a display 43; wherein, an executable code is stored on the memory 42, and when the executable code is executed by the processor 41, the processor 41 is caused to execute the following steps:
[0242] Collect voice signal blocks in the voice signal stream, each voice signal block having a preset duration;
[0243] Upload the collected voice signal blocks to the server, so that the server obtains the first voice signal block and at least one voice signal block included in the target group corresponding to the first voice signal block, and encodes the at least one voice signal block through an encoding network in the voice recognition model to obtain a semantic vector corresponding to the first voice signal block; decode the semantic vectors corresponding to the at least one voice signal block respectively through a decoding network in the voice recognition model to output the text corresponding to the first voice signal block; wherein, the target group includes the first voice signal block, and a group includes voice signal blocks generated in sequence within a preset group duration;
[0244] Display the text received from the server through the display 43.
[0245] The processor 41 may also execute other related steps in the voice interaction method provided in the foregoing related embodiments, which will not be elaborated herein.
[0246] Optionally, the voice interaction device may further include a communication interface 44 for communicating with other devices.
[0247] In addition, an embodiment of the present invention further provides a non-transitory machine-readable storage medium, on which an executable code is stored, and when the executable code is Figure 12 executed by the processor of the voice interaction device shown, the processor is caused to execute the corresponding voice interaction method.
[0248] Figure 13 The structure diagram of a voice interaction device provided by an embodiment of the present invention is as Figure 13 shown, the device includes: an acquisition module 51, an encoding module 52, a decoding module 53, and a display module 54.
[0249] The acquisition module 51 is configured to obtain a first voice signal block in the conference voice signal stream based on a preset voice signal block duration; and obtain at least one voice signal block included in the target group corresponding to the first voice signal block, the target group includes the first voice signal block, and a group includes voice signal blocks generated in sequence within a preset group duration.
[0250] An encoding module 52, configured to encode the at least one voice signal in chunks through an encoding network in a voice recognition model, so as to obtain a semantic vector corresponding to the first voice signal chunk.
[0251] A decoding module 53, configured to decode the semantic vectors corresponding to the at least one voice signal chunk respectively through a decoding network in the voice recognition model, so as to output the text corresponding to the first voice signal chunk.
[0252] A display module 54, configured to display the text.
[0253] Optionally, the device further includes: a recording module, configured to generate a meeting record according to the text.
[0254] Optionally, the target group includes at least one second voice signal chunk. Based on this, the encoding module 52 is specifically configured to: obtain the acoustic features of the first voice signal chunk and the acoustic features of the at least one second voice signal chunk; encode the splicing result of the acoustic features of the first voice signal chunk and the acoustic features of the at least one second voice signal chunk through the encoding network, so as to obtain a semantic vector corresponding to the first voice signal chunk.
[0255] Optionally, the decoding module 53 may specifically be configured to: predict the semantic vector corresponding to the first voice signal chunk through a prediction network in the voice recognition model, so as to obtain the number of words included in the first voice signal chunk; determine the weights and the weighted sum result of the semantic vectors corresponding to the at least one voice signal chunk respectively in one weight calculation process through an attention network in the voice recognition model, and input the weighted sum result into the decoding network, so as to output, through the decoding network, the text corresponding to the weighted sum result, where the number of words is used to constrain the number of weight calculation times.
[0256] In a possible design, the structure of the above Figure 13 shown voice interaction device may be implemented as a voice interaction device, as Figure 14 shown, the voice interaction device may include: a processor 61, a memory 62, and a display 63; wherein, an executable code is stored on the memory 62, and when the executable code is executed by the processor 61, the processor 61 is caused to execute the voice interaction method as Figure 7a shown.
[0257] The processor 61 may also execute other related steps in the voice interaction method provided in the foregoing related embodiments, which will not be elaborated herein.
[0258] Optionally, the voice interaction device may further include a communication interface 64 for communicating with other devices.
[0259] In addition, an embodiment of the present invention further provides a non-transitory machine-readable storage medium, on which executable code is stored. When the executable code is Figure 14 executed by a processor of the voice interaction device as shown, the processor is caused to execute the voice interaction method as Figure 7a shown.
[0260] Figure 15 FIG. is a schematic structural diagram of a voice interaction device provided by an embodiment of the present invention. As Figure 15 shown, the device includes: an acquisition module 71, an encoding module 72, a decoding module 73, and a display module 74.
[0261] The acquisition module 71 is configured to obtain a first voice signal block in the voice signal stream of the host based on a preset voice signal block duration; and obtain at least one voice signal block included in a target group corresponding to the first voice signal block, where the target group includes the first voice signal block, and a group includes voice signal blocks sequentially generated within a preset group duration.
[0262] The encoding module 72 is configured to encode the at least one voice signal block through an encoding network in a voice recognition model to obtain a semantic vector corresponding to the first voice signal block.
[0263] The decoding module 73 is configured to decode the semantic vectors corresponding to the at least one voice signal block through a decoding network in the voice recognition model to output the text corresponding to the first voice signal block.
[0264] The display module 74 is configured to display the text.
[0265] Optionally, the target group includes at least one second voice signal block. Based on this, the encoding module 72 is specifically configured to: obtain the acoustic features of the first voice signal block and the acoustic features of the at least one second voice signal block; encode the concatenation result of the acoustic features of the first voice signal block and the acoustic features of the at least one second voice signal block through the encoding network to obtain a semantic vector corresponding to the first voice signal block.
[0266] Optionally, the decoding module 73 may specifically be configured to: predict the semantic vectors corresponding to the segmented first voice signal through a prediction network in the voice recognition model to obtain the number of characters included in the segmented first voice signal; determine the weights and weighted sum results of the semantic vectors corresponding to each of the at least one voice signal segment in one weight calculation process through an attention network in the voice recognition model, and input the weighted sum result into the decoding network to output, through the decoding network, characters corresponding to the weighted sum result, where the number of characters is used to constrain the number of weight calculation times.
[0267] In a possible design, the structure of the above Figure 15 shown voice interaction device may be implemented as a voice interaction device, such as Figure 16 shown. The voice interaction device may include: a processor 81, a memory 82, and a display 83; wherein, an executable code is stored on the memory 82, and when the executable code is executed by the processor 81, the processor 81 is caused to execute the Figure 8a shown voice interaction method.
[0268] The processor 81 may also execute other related steps in the voice interaction method provided in the foregoing related embodiments, which will not be elaborated herein.
[0269] Optionally, the voice interaction device may further include a communication interface 84 for communicating with other devices.
[0270] In addition, an embodiment of the present invention further provides a non-transitory machine-readable storage medium, on which an executable code is stored. When the executable code is executed by Figure 16 the processor of the shown voice interaction device, the processor is caused to execute the Figure 8a shown voice interaction method.
[0271] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0272] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform. Of course, it can also be implemented by a combination of hardware and software. Based on such an understanding, the essence of the above technical solution or the part that contributes to the prior art can be embodied in the form of a computer product. The present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0273] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A speech recognition method, characterized in that, Including: Obtaining a first speech signal chunk in a speech signal stream based on a preset speech signal chunk duration; Obtaining at least one speech signal chunk included in a target group corresponding to the first speech signal chunk, where the at least one speech signal chunk includes the first speech signal chunk and speech signal chunks that are generated before the first speech signal and belong to the same target group as the first speech signal, and the duration of one group is a preset group duration; Encoding the at least one speech signal chunk through an encoding network in a speech recognition model to obtain a semantic vector corresponding to the first speech signal chunk; Decoding the semantic vectors corresponding to the at least one speech signal chunk through a decoding network in the speech recognition model to output the text corresponding to the first speech signal chunk.
2. The method according to claim 1, wherein The target group includes at least one second speech signal chunk; The encoding the at least one speech signal chunk through an encoding network in a speech recognition model to obtain a semantic vector corresponding to the first speech signal chunk includes: Obtaining the acoustic features of the first speech signal chunk and the acoustic features of the at least one second speech signal chunk; Encoding the concatenation result of the acoustic features of the first speech signal chunk and the acoustic features of the at least one second speech signal chunk through the encoding network to obtain a semantic vector corresponding to the first speech signal chunk.
3. The method according to claim 2, wherein Obtaining the acoustic features of any speech signal chunk in the following manner: Performing frame division processing on the any speech signal chunk to obtain multiple frames of speech signals; Extracting the acoustic features corresponding to the multiple frames of speech signals; Determining the acoustic features of the any speech signal chunk according to the acoustic features corresponding to the multiple frames of speech signals.
4. The method according to claim 1, wherein The encoding network is an encoding network with a memory module, and the memory module is used to memorize the acoustic features of at least one frame of speech signal generated last in the previous group, and the at least one frame of speech signal is within at least one speech signal chunk generated last in the previous group.
5. The method according to claim 4, characterized in that, The first speech signal chunk is the first speech signal chunk in the target group; The encoding the at least one speech signal chunk through an encoding network in a speech recognition model to obtain a semantic vector corresponding to the first speech signal chunk includes: Encoding the concatenation result of the acoustic features of at least one frame of speech signal currently memorized and the acoustic features of the first speech signal chunk through the encoding network to obtain a semantic vector corresponding to the first speech signal chunk.
6. The method according to claim 1, characterized in that, The decoding the semantic vectors corresponding to the at least one speech signal chunk through a decoding network in the speech recognition model to output the text corresponding to the first speech signal chunk includes: Predicting the semantic vector corresponding to the first speech signal chunk through a prediction network in the speech recognition model to obtain the number of words included in the first speech signal chunk; Determine the weights of the semantic vectors corresponding to each of the at least one speech signal chunk and the weighted sum result in one weight calculation process through the attention network in the speech recognition model, and input the weighted sum result into the decoding network to output, through the decoding network, the text corresponding to the weighted sum result, where the number of texts is used to constrain the number of weight calculation times.
7. The method according to claim 6, characterized in that The predicting, by the prediction network in the speech recognition model, the semantic vector corresponding to the first speech signal chunk to obtain the number of texts included in the first speech signal chunk includes: Output, through the prediction network, a prediction sequence corresponding to the semantic vector corresponding to the first speech signal chunk, where the prediction sequence is used to indicate whether each of the multiple frames of speech signals included in the first speech signal chunk corresponds to a text; The number of texts included in the first speech signal chunk is determined by the number of frames of the speech signals corresponding to the texts in the prediction sequence.
8. The method according to claim 6, wherein The first speech signal chunk includes N texts, where N is greater than or equal to 1; The determining, by the attention network in the speech recognition model, the weights of the semantic vectors corresponding to each of the at least one speech signal chunk and the weighted sum result in one weight calculation process, and inputting the weighted sum result into the decoding network to output, through the decoding network, the text corresponding to the weighted sum result includes: Initialize the number of weight calculation times to N, and iteratively execute the following process until the number of weight calculation times is reduced to 0: If the current number of weight calculation times is not 0, input the semantic vectors corresponding to each of the at least one speech signal chunk and the previous text output by the decoding network into the attention network, so that the attention network determines the weights of the semantic vectors corresponding to each of the at least one speech signal chunk based on the previous text, and determines the weighted sum result of the semantic vectors corresponding to each of the at least one speech signal chunk according to the weights, and input the weighted sum result into the decoding network; Obtain the currently output text through the decoding network, and input the text into the attention network; Decrease the number of weight calculation times by one.
9. The method according to claim 6, wherein The method further includes: If the first speech signal chunk does not include a text, or if none of the continuously preset number of speech signal chunks starting from the first speech signal chunk includes a text, use the next speech signal chunk including a text as the starting point of a new group; Delete the historical storage information corresponding to the new group, where the historical storage information includes the acoustic features and semantic vectors corresponding to each speech signal chunk in each group before the generation of the new group.
10. A voice recognition device, characterized in that, including: An obtaining module, configured to obtain a first speech signal chunk in a speech signal stream based on a preset speech signal chunk duration; And, obtaining at least one speech signal chunk included in a target group corresponding to the first speech signal chunk, where the at least one speech signal chunk includes the first speech signal chunk and speech signal chunks that are generated before the first speech signal and belong to the same target group as the first speech signal, and the duration of one group is a preset group duration; An encoding module, configured to encode the at least one speech signal chunk through an encoding network in a speech recognition model to obtain a semantic vector corresponding to the first speech signal chunk; A decoding module, configured to decode the semantic vectors corresponding to the at least one speech signal chunk respectively through a decoding network in the speech recognition model to output the text corresponding to the first speech signal chunk.
11. The device according to claim 10, characterized in that, The target group includes at least one second speech signal chunk; Specifically, the encoding module is configured to: Obtain the acoustic features of the first speech signal chunk and the acoustic features of the at least one second speech signal chunk; Encode the concatenation result of the acoustic features of the first speech signal chunk and the acoustic features of the at least one second speech signal chunk through the encoding network to obtain a semantic vector corresponding to the first speech signal chunk.
12. An electronic device, characterized in that, Includes: A memory and a processor; wherein, executable code is stored on the memory, and when the executable code is executed by the processor, the processor is caused to execute the speech recognition method according to any one of claims 1 to 9.
13. A non-transitory machine-readable storage medium, characterized in that, Executable code is stored on the non-transitory machine-readable storage medium, and when the executable code is executed by a processor of an electronic device, the processor is caused to execute the speech recognition method according to any one of claims 1 to 9.
14. A voice interaction method, characterized in that, Includes: Collecting speech signal chunks in a speech signal stream, each speech signal chunk having a preset duration; Uploading the collected speech signal chunks to a server, so that the server obtains a first speech signal chunk and at least one speech signal chunk included in a target group corresponding to the first speech signal chunk, encodes the at least one speech signal chunk through an encoding network in a speech recognition model to obtain a semantic vector corresponding to the first speech signal chunk; decodes the semantic vectors corresponding to the at least one speech signal chunk respectively through a decoding network in the speech recognition model to output the text corresponding to the first speech signal chunk; wherein, the at least one speech signal chunk includes the first speech signal chunk and speech signal chunks that are generated before the first speech signal and belong to the same target group as the first speech signal, and the duration of one group is a preset group duration; Displaying the text received from the server.
15. The method according to claim 14, wherein The target group includes at least one second speech signal chunk; the process of obtaining the semantic vector corresponding to the first speech signal chunk includes: Obtaining the acoustic features of the first speech signal chunk and the acoustic features of the at least one second speech signal chunk; Encoding the concatenation result of the acoustic features of the first speech signal block and the acoustic features of the at least one second speech signal block through the encoding network to obtain a semantic vector corresponding to the first speech signal block.
16. A voice interaction device, characterized in that, Comprising: A memory, a processor, and a display; wherein, executable code is stored on the memory, and when the executable code is executed by the processor, the processor executes the speech interaction method as claimed in claim 14 or 15.
17. A non-transitory machine-readable storage medium, characterized in that, Executable code is stored on the non-transitory machine-readable storage medium, and when the executable code is executed by the processor of the electronic device, the processor executes the speech interaction method as claimed in claim 14 or 15.
18. A voice interaction method, characterized in that, Comprising: Obtaining a first speech signal block in the conference speech signal stream based on a preset speech signal block duration. Obtaining at least one speech signal block included in a target group corresponding to the first speech signal block, the at least one speech signal block including the first speech signal block and the speech signal blocks that are generated before the first speech signal and belong to the same target group as the first speech signal, and the duration of one group is a preset group duration. Encoding the at least one speech signal block through the encoding network in the speech recognition model to obtain a semantic vector corresponding to the first speech signal block. Decoding the semantic vectors corresponding to the at least one speech signal block respectively through the decoding network in the speech recognition model to output the text corresponding to the first speech signal block. Displaying the text.
19. The method according to claim 18, wherein The method further comprises: Generating a meeting record according to the text.
20. The method according to claim 18, wherein The target group includes at least one second speech signal block. The encoding the at least one speech signal block through the encoding network in the speech recognition model to obtain a semantic vector corresponding to the first speech signal block includes: Obtaining the acoustic features of the first speech signal block and the acoustic features of the at least one second speech signal block. Encoding the concatenation result of the acoustic features of the first speech signal block and the acoustic features of the at least one second speech signal block through the encoding network to obtain a semantic vector corresponding to the first speech signal block.
21. The method according to claim 18, characterized in that, The decoding the semantic vectors corresponding to the at least one speech signal block respectively through the decoding network in the speech recognition model to output the text corresponding to the first speech signal block includes: Predicting the semantic vector corresponding to the first speech signal block through the prediction network in the speech recognition model to obtain the number of words included in the first speech signal block. Determining the weights and the weighted sum result of the semantic vectors corresponding to the at least one speech signal block respectively in one weight calculation process through the attention network in the speech recognition model, and inputting the weighted sum result into the decoding network to output the text corresponding to the weighted sum result through the decoding network, and the number of words is used to constrain the number of weight calculation times.
22. A voice interaction device, characterized in that, Comprising: A memory and a processor; wherein, executable code is stored on the memory, and when the executable code is executed by the processor, the processor executes the voice interaction generation method according to any one of claims 18 to 21.
23. A voice interaction method, characterized in that, Including: Based on a preset voice signal block duration, obtaining a first voice signal block in the voice signal stream of the anchor; Obtaining at least one voice signal block included in a target group corresponding to the first voice signal block, the at least one voice signal block including the first voice signal block and voice signal blocks belonging to the same target group and generated before the first voice signal, and the duration of one group is a preset group duration; Encoding the at least one voice signal block through an encoding network in a voice recognition model to obtain a semantic vector corresponding to the first voice signal block; Decoding the semantic vectors corresponding to the at least one voice signal block through a decoding network in the voice recognition model to output the text corresponding to the first voice signal block; Displaying the text in a live broadcast interface.
24. The method according to claim 23, characterized in that, The target group includes at least one second voice signal block; The encoding the at least one voice signal block through an encoding network in a voice recognition model to obtain a semantic vector corresponding to the first voice signal block includes: Obtaining the acoustic features of the first voice signal block and the acoustic features of the at least one second voice signal block; Encoding the concatenation result of the acoustic features of the first voice signal block and the acoustic features of the at least one second voice signal block through the encoding network to obtain a semantic vector corresponding to the first voice signal block.
25. The method according to claim 23, wherein The decoding the semantic vectors corresponding to the at least one voice signal block through a decoding network in the voice recognition model to output the text corresponding to the first voice signal block includes: Predicting the semantic vector corresponding to the first voice signal block through a prediction network in the voice recognition model to obtain the number of characters included in the first voice signal block; Determining the weights and weighted summation results of the semantic vectors corresponding to the at least one voice signal block in one weight calculation process through an attention network in the voice recognition model, and inputting the weighted summation result into the decoding network to output the text corresponding to the weighted summation result through the decoding network, and the number of characters is used to constrain the number of weight calculation times.
26. A voice interaction device, characterized in that, Including: A memory and a processor; wherein, executable code is stored on the memory, and when the executable code is executed by the processor, the processor executes the voice interaction generation method according to any one of claims 23 to 25.
Citation Information
Patent Citations
System and method for real-time transcription of an audio signal into texts
CN109417583A
Voice recognition model generation method and device and voice recognition method and device
CN111696526A