Voice processing method and electronic device

By setting different processing units in the speech recognition model to process the end positions and non-end positions of the voice audio stream, the problem of difficulty in balancing the recognition accuracy and speed of traditional technologies is solved, and the efficient recognition effect in streaming speech recognition is achieved.

CN114582350BActive Publication Date: 2025-05-27LENOVO (BEIJING) LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210282952.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-22
Publication Date
2025-05-27
Estimated Expiration
2042-03-22

AI Technical Summary

Technical Problem

Traditional end-to-end speech recognition technology is difficult to balance recognition accuracy and recognition speed simultaneously, especially in streaming speech recognition scenarios.

Method used

By setting different processing units in the speech recognition model, speech blocks at the end positions and non-end positions of the speech audio stream are processed respectively. The first processing unit processes the voice block at the end position, and the second processing unit processes the voice block at the non-end position, ensuring that the processing time of the first processing unit is lower than that of the second processing unit.

Benefits of technology

It realizes the effect of taking into account the recognition accuracy and speed in streaming speech recognition, reduces the processing delay of voice blocks at the end position, and ensures the overall recognition rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114582350B_ABST
    Figure CN114582350B_ABST
Patent Text Reader

Abstract

The present application discloses a voice processing method and an electronic device. For the voice to be processed, the present application processes the voice blocks at the end positions of the voice to be processed through a first processing unit of a voice recognition model, and processes the voice blocks other than the end positions (i.e., the voice blocks at non-end positions) through a second processing unit of the voice recognition model. Among them, the time required for the first processing unit in the voice recognition model to process the voice blocks to be processed is shorter than the time required for the second processing unit to process the voice blocks to be processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of information processing, and particularly relates to a voice processing method and an electronic device. Background Art

[0002] Streaming speech recognition is an essential skill for the application of speech recognition engines. It is more difficult than offline recognition because streaming speech recognition needs to balance the recognition rate and recognition speed at the same time. That is to say, it requires both high speech recognition accuracy and fast recognition speed. However, traditional end-to-end speech recognition technologies are difficult to balance the recognition accuracy and recognition speed at the same time. Summary of the Invention

[0003] Therefore, this application discloses the following technical solutions:

[0004] A voice processing method, comprising:

[0005] Obtaining a to-be-processed speech block of the to-be-processed speech currently;

[0006] In response to the to-be-processed speech block being a speech block at an end position of the to-be-processed speech, processing the to-be-processed speech block at least through a first processing unit of a speech recognition model to obtain a first processing result;

[0007] In response to the to-be-processed speech block being a speech block other than the end position of the to-be-processed speech, processing the to-be-processed speech block at least through a second processing unit of the speech recognition model to obtain a second processing result;

[0008] Wherein, the time required for the first processing unit to process the to-be-processed speech block is shorter than the time required for the second processing unit to process the to-be-processed speech block.

[0009] Optionally, the end position includes at least one of the following:

[0010] Positions corresponding to a first number of speech blocks at the head of the speech audio stream;

[0011] Positions corresponding to a second number of speech blocks at the tail of the speech audio stream;

[0012] Wherein, the first number of speech blocks at the head of the speech audio stream start from the first-word speech block of the speech audio stream, and the first-word speech block is the speech block corresponding to the speech audio stream that contains the first speech word of the speech audio stream.

[0013] Optionally, the process of identifying the first-word speech block of the speech audio stream includes:

[0014] At least through the first processing unit, start processing the speech audio stream from the first speech block of the speech audio stream to obtain the processing result of the currently processed speech block;

[0015] If the processing result of the currently processed speech block contains recognized text characters, determine that the currently processed speech block is the first-word speech block of the speech audio stream;

[0016] If the processing result of the currently processed speech block does not contain recognized text characters, at least through the first processing unit, process the next speech block of the currently processed speech block until the processing result of the speech block contains recognized text characters, and determine the speech block corresponding to the processing result containing text characters as the first-word speech block of the speech audio stream.

[0017] Optionally, the speech recognition model includes a first encoding unit, a second encoding unit, and a decoding unit, the first processing unit is the first encoding unit, and the second processing unit is the second encoding unit;

[0018] The at least through the first processing unit of the speech processing model processes the speech block to be processed to obtain a first processing result, including:

[0019] Perform codec-based speech recognition on the speech block to be processed through the first encoding method corresponding to the first encoding unit and the decoding method corresponding to the decoding unit of the speech recognition model to obtain a first recognition result;

[0020] The at least through the second processing unit of the speech recognition model processes the speech block to be processed to obtain a second processing result, including:

[0021] Perform codec-based speech recognition on the speech block to be processed through the second encoding method corresponding to the second encoding unit and the decoding method corresponding to the decoding unit of the speech recognition model to obtain a second recognition result.

[0022] Optionally, the speech recognition model includes an encoding unit, a first decoding unit, and a second decoding unit, the first processing unit is the first decoding unit, and the second processing unit is the second decoding unit;

[0023] The at least through the first processing unit of the speech processing model processes the speech block to be processed to obtain a first processing result, including:

[0024] Perform codec-based speech recognition on the speech block to be processed through the encoding method corresponding to the encoding unit and the first decoding method corresponding to the first decoding unit of the speech recognition model to obtain a first recognition result;

[0025] The second processing unit that processes the to-be-processed speech block at least through the speech recognition model obtains a second processing result, including:

[0026] Performing codec-based speech recognition on the to-be-processed speech block through the encoding method corresponding to the encoding unit of the speech recognition model and the second decoding method corresponding to the second decoding unit to obtain a second recognition result.

[0027] Optionally, the speech recognition model includes a first encoding unit, a second encoding unit, a first decoding unit, and a second decoding unit. The first processing unit includes the first encoding unit and the first decoding unit, and the second processing unit includes the second encoding unit and the second decoding unit;

[0028] The first processing unit that processes the to-be-processed speech block at least through the speech processing model obtains a first processing result, including:

[0029] Performing codec-based speech recognition on the to-be-processed speech block through the first encoding method corresponding to the first encoding unit of the speech recognition model and the first decoding method corresponding to the first decoding unit to obtain a first recognition result;

[0030] The second processing unit that processes the to-be-processed speech block at least through the speech recognition model obtains a second processing result, including:

[0031] Performing codec-based speech recognition on the to-be-processed speech block through the second encoding method corresponding to the second encoding unit of the speech recognition model and the second decoding method corresponding to the second decoding unit to obtain a second recognition result.

[0032] Optionally, the first encoding unit and the second encoding unit each include at least one encoding layer for performing speech encoding;

[0033] The first encoding unit and the second encoding unit satisfy at least one of the following conditions:

[0034] The number of encoding layers of the first encoding unit is lower than the number of encoding layers of the second encoding unit;

[0035] The compositional structure complexity of the encoding layers of the first encoding unit is lower than the compositional structure complexity of the encoding layers of the second encoding unit.

[0036] Optionally, the first decoding unit and the second decoding unit each include at least one decoding layer for performing speech decoding;

[0037] The first decoding unit and the second decoding unit satisfy at least one of the following conditions:

[0038] The number of decoding layers of the first decoding unit is lower than that of the second decoding unit;

[0039] The compositional structure complexity of the decoding layers of the first decoding unit is lower than that of the decoding layers of the second decoding unit.

[0040] Optionally, the first encoding unit and the second encoding unit share at least one encoding layer.

[0041] An electronic device, comprising:

[0042] A memory for storing at least a set of instruction sets;

[0043] A processor for calling and executing the instruction sets in the memory, and implementing the voice processing method as described in any one of the above by executing the instruction sets.

[0044] As can be seen from the above solutions, for the voice to be processed in the voice processing method and the electronic device disclosed in this application, the voice blocks at the end positions are processed by the first processing unit of the voice recognition model, and the voice blocks other than the end positions (i.e., the voice blocks at non-end positions) are processed by the second processing unit of the voice recognition model. Among them, the time required for the first processing unit in the voice recognition model to process the voice blocks to be processed is shorter than the time required for the second processing unit to process the voice blocks to be processed.

[0045] This application changes the model structure of the voice recognition model, sets different processing units in the voice recognition model, and accelerates the processing speed of the voice blocks at the end positions and reduces the processing delay of the voice blocks at the end positions by using the set first processing unit to process the voice blocks at the end positions of the voice to be processed, and ensures the overall recognition rate of the voice to be processed by using the set second processing unit to process the voice blocks at non-end positions of the voice to be processed, thereby taking into account both the recognition accuracy and speed of the voice to be processed, and effectively meeting the requirements for both recognition accuracy and recognition speed in streaming voice recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0047] Figure 1 It is a schematic flowchart of the voice processing method provided by this application;

[0048] Figure 2(a) is the structural diagram of the speech recognition model for model improvement at the encoding end provided by this application;

[0049] Figure 2(b) is the structural diagram of the speech recognition model for model improvement at the decoding end provided by this application;

[0050] Figure 2(c) is the structural diagram of the speech recognition model for model improvement at both the encoding end and the decoding end provided by this application;

[0051] Figure 3 is an example of different processing methods for the end position of the speech audio stream and speech blocks other than the end position provided by this application;

[0052] Figure 4 is the schematic flow diagram of identifying the first-word speech block provided by this application;

[0053] Figure 5 is an example of sharing part of the encoding layer by the first and second encoding units provided by this application;

[0054] Figure 6(a) is an example of the composition structure of a decoding layer in the second decoding unit under the AED architecture provided by this application;

[0055] Figure 6(b) is an example of the composition structure of two adjacent decoding layers in the first decoding unit under the AED architecture provided by this application;

[0056] Figure 7 is the composition structure diagram of the electronic device provided by this application. Specific Embodiments

[0057] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0058] This application discloses a speech processing method and an electronic device, which are applicable to but not limited to speech recognition processing in a streaming speech recognition scenario to effectively meet the requirements for recognition accuracy and recognition speed in speech processing such as streaming speech recognition.

[0059] The embodiments of this application will mainly use streaming speech recognition as an example to illustrate the solution.

[0060] The voice processing method can be applied to an electronic device with voice processing function. The electronic device applying the method of the present application can be, but is not limited to, devices under numerous general or special computing device environments or configurations with voice processing function, such as: personal computers, server computers, handheld devices or portable devices, tablet devices, AR / VR devices, multi-processor devices, and so on.

[0061] Refer to Figure 1 the flowchart of the information processing method provided. The information processing method disclosed in the embodiments of the present application includes the following processing steps:

[0062] Step 101, obtain the to-be-processed voice block of the current to-be-processed voice.

[0063] The to-be-processed voice is a voice audio stream to be processed, and specifically can be a voice audio stream corresponding to a voice sentence, or a voice audio stream corresponding to a continuous plurality of voice sentences. And it can be a voice audio stream collected in real time at the recording site (such as, meeting recording site, call recording site), or a recording audio stream corresponding to a recording file that has been recorded before, without limitation.

[0064] Streaming speech recognition requires supporting real-time return of recognition results during the process of processing the audio stream. Based on this, the present application divides the to-be-processed voice into individual voice blocks, and performs speech recognition in units of voice blocks to meet the real-time return requirements of recognition results in scenarios such as streaming speech recognition.

[0065] Among them, a voice block is composed of a certain duration of speech frames. By sequentially obtaining the preset number of speech frames to be processed in the current to-be-processed voice audio stream, a to-be-processed voice block composed of the preset number of speech frames organized in sequence is obtained. That is, by accumulating a certain number of frames of speech frames in physical time, a to-be-processed voice block is obtained, and it is input into a speech recognition model for speech recognition processing. Until when processing to the end of the to-be-processed voice audio stream, according to the actual number of speech frames included in the last voice block at its end (less than or equal to the above preset number), the last voice block composed of each speech frame including the actual number of frames is obtained.

[0066] Step 102, in response to the to-be-processed voice block being a voice block at the end position of the to-be-processed voice, process the to-be-processed voice block at least through the first processing unit of the speech recognition model to obtain a first processing result.

[0067] When end-to-end recognition uses a neural network model to complete the conversion from voice to text in speech recognition, the mainstream model architecture is the Encoder-Decoder structure. Once the model is trained, all speech chunks in the recognition stage are calculated by this fixed model in a fixed way. The more complex the model structure and the more layers it has, the higher the recognition accuracy, but the slower the recognition speed; the simpler the model structure and the fewer layers it has, the faster the recognition speed, thus it is impossible to balance both recognition accuracy and recognition speed.

[0068] This application solves the recognition rate problem from the perspective of optimizing the inference time of the speech recognition model, while ensuring that the overall sentence recognition rate of the speech does not decrease significantly.

[0069] The speech recognition model can be an AED model (i.e., an Encoder-Decoder model based on the attention mechanism), but it is not limited to this. For example, it can also be but not limited to other deep learning-based models such as RNN (Recurrent Neural Network).

[0070] Among them, in the embodiments of this application, two processing units with basically the same functions (such as both having speech encoding functions or both having speech decoding functions) but different structures are specifically set in the speech recognition model: a first processing unit and a second processing unit, to improve the structure of the speech recognition model. The structure of the first processing unit is simple, and compared with the first processing unit, the structure of the second processing unit is more complex, so that for the same speech chunk to be processed, the time required for the first processing unit to process this speech chunk to be processed is shorter than the time required for the second processing unit to process this speech chunk to be processed.

[0071] The speech recognition model includes an encoding unit (Encoder) for providing speech encoding functions in speech recognition and a decoding unit (Decoder) for providing speech decoding functions in speech recognition, and completes the conversion from voice to text in speech recognition through the encoding and decoding processing of speech chunks.

[0072] Among them, in one embodiment, the speech recognition model is improved at the encoding end according to the above-mentioned improved method provided by the present application. Referring to Fig. 2(a), specifically, two encoding units (Encoders) with different structures are set in the speech recognition model: the first encoding unit (Encoder1) and the second encoding unit (Encoder2). Compared with the second encoding unit, the composition structure of the first encoding unit is simpler, and for the same speech block, the time consumed by the first encoding unit to complete the encoding process of the speech block is lower. In this embodiment, the first encoding unit is used as the above-mentioned first processing unit, and the second encoding unit is used as the above-mentioned second processing unit. There is no change at the decoding end, and an existing decoding unit (Decoder) that can implement the decoding function can be used.

[0073] Optionally, in another embodiment, the speech recognition model is improved at the decoding end according to the above-mentioned improved method provided by the present application. Referring to Fig. 2(b), specifically, two decoding units (Decoders) with different structures are set in the speech recognition model: the first decoding unit (Decoder1) and the second decoding unit (Decoder2). Compared with the second decoding unit, the composition structure of the first decoding unit is simpler, and for the same speech block, the time consumed by the first decoding unit to complete the decoding process of the speech block is lower. In this embodiment, the first decoding unit is used as the above-mentioned first processing unit, and the second decoding unit is used as the above-mentioned second processing unit. There is no change at the encoding end, and an existing encoding unit (Encoder) that can implement the encoding function can be used.

[0074] In addition, in other embodiments, the speech recognition model can also be improved at the encoding end and the decoding end simultaneously according to the above-mentioned improved method provided by the present application. Referring to Fig. 2(c), specifically, two encoding units (Encoders) with different structures are set in the speech recognition model: the first encoding unit and the second encoding unit, and two decoding units (Decoders) with different structures are set: the first decoding unit and the second decoding unit. Among them, compared with the second encoding unit, the composition structure of the first encoding unit is simpler and the encoding time for the same speech block is lower. Similarly, compared with the second decoding unit, the composition structure of the first decoding unit is simpler and the decoding time for the same speech block is lower. In this embodiment, the first encoding unit and the first decoding unit are used as the above-mentioned first processing unit, and the second encoding unit and the second decoding unit are used as the above-mentioned second processing unit.

[0075] For the current speech chunk to be processed of the speech to be processed, in this embodiment, according to the position of the speech chunk to be processed in the speech to be processed, it is identified whether the speech chunk to be processed is a speech chunk at a preset end position of the speech to be processed, and for the two different cases where the speech chunk to be processed is a speech chunk at the end position of the speech to be processed or a speech chunk other than the end position, different processing methods for the speech chunk to be processed are correspondingly selected.

[0076] Among them, if the speech chunk to be processed is a speech chunk at the end position in the speech to be processed, the speech chunk to be processed is at least processed by the first processing unit of the speech recognition model, and accordingly, the speech chunk to be processed is subjected to speech recognition processing by the processing method corresponding to the first processing unit to obtain a first processing result.

[0077] The above-mentioned end position can be, but is not limited to, at least one of the following position one and position two:

[0078] Position one: The position corresponding to the first number of speech chunks at the head of the speech audio stream;

[0079] For example, the position corresponding to 2 speech chunks at the head of the audio stream of one speech sentence or multiple consecutive speech sentences.

[0080] Position two: The position corresponding to the second number of speech chunks at the tail of the speech audio stream.

[0081] For example, the position corresponding to 1 speech chunk at the tail of the audio stream of one speech sentence or multiple consecutive speech sentences.

[0082] Preferably, the end position is set to include both the above-mentioned position one and position two. In this case, the speech chunks other than the end position are other speech chunks other than the first number of speech chunks at the head and the second number of speech chunks at the tail in the speech audio stream, that is, the middle position speech chunks other than the head and the tail.

[0083] The above-mentioned first number and second number can be the same or different, and there is no limitation in this regard. For example, both the first number and the second number are 1, or the first number is 2 and the second number is 1, etc.

[0084] For the speech chunks at the end position of the speech audio stream to be processed, in this embodiment, the speech chunks at the end position are processed by using a first processing unit with a simpler structure (compared to the structure of the second processing unit), such as encoding and / or decoding processing, to improve the processing speed of the speech recognition model for the speech chunks at the end position of the speech audio stream.

[0085] Among them, with reference to Figure 3Example, for the implementation of improving the speech recognition model at the encoding end, if the speech block to be processed is the speech block at the end position of the speech to be processed, in this step 102, the speech block to be processed is at least processed by the first processing unit of the speech recognition model to obtain a first processing result, which can be specifically implemented as: through the first encoding method corresponding to the first encoding unit (such as the small Encoder in Figure 3 and the decoding method corresponding to the decoding unit, perform speech recognition based on encoding and decoding on the speech block to be processed to obtain a first recognition result.

[0086] Correspondingly, for the implementation of improving the speech recognition model at the decoding end, if the speech block to be processed is the speech block at the end position of the speech to be processed, in this step 102, the speech block to be processed is at least processed by the first processing unit of the speech recognition model to obtain a first processing result, which can be specifically implemented as: through the encoding method corresponding to the encoding unit of the speech recognition model and the first decoding method corresponding to the first decoding unit, perform speech recognition based on encoding and decoding on the speech block to be processed to obtain a first recognition result.

[0087] For the implementation of improving the speech recognition model at both the encoding end and the decoding end, if the speech block to be processed is the speech block at the end position of the speech to be processed, in this step 102, the speech block to be processed is at least processed by the first processing unit of the speech recognition model to obtain a first processing result, which can be specifically implemented as: through the first encoding method corresponding to the first encoding unit and the first decoding method corresponding to the first decoding unit of the speech recognition model, perform speech recognition based on encoding and decoding on the speech block to be processed to obtain a first recognition result.

[0088] Step 103: In response to the speech block to be processed being a speech block other than the end position in the speech to be processed, at least process the speech block to be processed by the second processing unit of the speech recognition model to obtain a second processing result.

[0089] For the speech block other than the end position of the speech to be processed (such as the middle position speech block other than the head and tail), by using the second processing unit with a complex structure (compared to the structure of the first processing unit) to process it (such as encoding and / or decoding), the overall recognition rate of the speech recognition model for the speech to be processed is guaranteed.

[0090] Such as Figure 3 shown, for the implementation of improving the speech recognition model at the encoding end, if the speech block to be processed is a speech block other than the end position of the speech to be processed, in this step 103, at least process the speech block to be processed by the second processing unit of the speech recognition model to obtain a second processing result, which can be specifically implemented as: through the second encoding unit of the speech recognition model (such as Figure 3The second encoding method corresponding to the big Encoder in and the decoding method corresponding to the decoding unit are used to perform codec-based speech recognition on the speech block to be processed, and a second recognition result is obtained.

[0091] Correspondingly, for the implementation manner of improving the speech recognition model at the decoding end, if the speech block to be processed is a speech block other than the end position of the speech to be processed, in this step 103, the speech block to be processed is at least processed by the second processing unit of the speech recognition model to obtain a second processing result, which can be specifically implemented as: performing codec-based speech recognition on the speech block to be processed through the encoding method corresponding to the encoding unit of the speech recognition model and the second decoding method corresponding to the second decoding unit, and obtaining a second recognition result.

[0092] For the implementation manner of improving the speech recognition model at both the encoding end and the decoding end, if the speech block to be processed is a speech block other than the end position of the speech to be processed, in this step 103, the speech block to be processed is at least processed by the second processing unit of the speech recognition model to obtain a second processing result, which can be specifically implemented as: performing codec-based speech recognition on the speech block to be processed through the second encoding method corresponding to the second encoding unit of the speech recognition model and the second decoding method corresponding to the second decoding unit, and obtaining a second recognition result.

[0093] The first recognition result or the second recognition result corresponding to the current speech block is the text information obtained by converting the speech information of the current speech block into text based on codec processing. In the corresponding implementation manner, specifically, the current speech block is converted from speech information to text information through the encoding method corresponding to the corresponding encoding unit and the decoding method corresponding to the corresponding decoding unit.

[0094] The first encoding method corresponding to the first encoding unit is determined by the composition structure of the first encoding unit. The second encoding method corresponding to the second encoding unit is correspondingly determined by the composition structure of the second encoding unit. Among them, for the same speech block to be processed, the encoding time of the first encoding method is lower, and the encoding process of the second encoding method can make the speech recognition accuracy of the same speech block higher.

[0095] Similarly, the first decoding method corresponding to the first decoding unit is determined by the composition structure of the first decoding unit. The second decoding method corresponding to the second decoding unit is correspondingly determined by the composition structure of the second decoding unit. Among them, for the same speech block to be processed, the decoding time of the first decoding method is lower, and the decoding process of the second decoding method can make the speech recognition accuracy of the same speech block higher.

[0096] As can be seen from the above solution, in the method of this embodiment, for the speech to be processed, the speech block at the end position is processed by the first processing unit of the speech recognition model, and the speech block other than the end position (i.e., the speech block at the non-end position) is processed by the second processing unit of the speech recognition model. Among them, the time required for the first processing unit in the speech recognition model to process the speech block to be processed is shorter than the time required for the second processing unit to process the speech block to be processed.

[0097] This application changes the model structure of the speech recognition model, sets different processing units in the speech recognition model, and accelerates the processing speed of the speech block at the end position and reduces the processing delay of the speech block at the end position by using the set first processing unit to process the speech block at the end position of the speech to be processed. By using the set second processing unit to process the speech block at the non-end position of the speech to be processed, the overall recognition rate of the speech to be processed is ensured, so as to balance the recognition accuracy and speed of the speech to be processed, and can effectively meet the requirements of both recognition accuracy and recognition speed for streaming speech recognition.

[0098] Streaming speech recognition generally uses the first-word delay and the last-word delay to measure the speed index of the user experience. During the process of using streaming speech recognition, users are more concerned about the first-word delay and the last-word delay, that is, it is required to have a lower first-word delay and last-word delay, and at the same time, the whole-sentence recognition rate of the speech cannot drop significantly. Among them, the first-word delay refers to the time consumed from when the user starts speaking in the audio stream to when the system recognizes the first word spoken by the user. The last-word delay is the time difference from when the user finishes speaking the last word in the audio stream to when the system recognizes the last word, which are respectively expressed as follows: First-word delay = time for accumulating one or several speech blocks at the beginning of the speech (physical time) + model inference time; Last-word delay = time for accumulating the last one or several speech blocks (physical time) + model inference time.

[0099] Based on this, in one embodiment, preferably, the first number of speech blocks at the head of the speech audio stream use the first-word speech block of the speech audio stream as the starting speech block, where the first-word speech block is specifically the speech block corresponding to the speech audio stream that contains the first speech word of the speech audio stream.

[0100] In this embodiment, for the current speech audio stream to be processed, when starting to perform recognition processing on the speech audio stream, first recognize the first-word speech block of the speech audio stream, so as to further recognize the first number of speech blocks at the head of the speech audio stream based on the recognized first-word speech block, and then use a processing unit with a simple structure (i.e., the first processing unit of the speech recognition model) to process the first number of speech blocks.

[0101] Among them, the process of recognizing the first-word speech block of the speech audio stream is as Figure 4 shown, and specifically includes:

[0102] Step 401: Process the speech audio stream starting from the first speech chunk of the speech audio stream through at least the first processing unit of the speech recognition model, and obtain the processing result of the currently processed speech chunk.

[0103] The speech chunk at the head of the speech audio stream (e.g., the first speech chunk at the head or a certain number of speech chunks at the head) may contain the actual speech content of the speaker, or may only contain silence or noise. For example, there is a certain delay between the start time of the speaker's speech and the recording start time, or there are pauses between different speech sentences during the speaker's speech, resulting in one or more speech chunks at the head of the audio stream corresponding to a certain speech sentence being silence or noise.

[0104] For this situation, in this embodiment, for the currently to-be-processed speech audio stream, at the start of speech recognition, first use a processing unit with a simple structure, i.e., the first processing unit of the speech recognition model, to process the speech chunks at the head of the speech audio stream starting from the first speech chunk, so as to obtain the recognition result of the head speech chunks starting from the first speech chunk in the speech audio stream more quickly.

[0105] Specifically, for the implementation manner of improving the speech recognition model at the encoding end, use the first encoding unit and decoding unit of the speech recognition model to process the speech chunks at the head of the speech audio stream starting from the first speech chunk; for the implementation manner of improving the speech recognition model at the decoding end, use the encoding unit and the first decoding unit of the speech recognition model to process the speech chunks at the head of the speech audio stream starting from the first speech chunk; for the implementation manner of improving the speech recognition model at both the encoding end and the decoding end, use the first encoding unit and the first decoding unit of the speech recognition model to process the speech chunks at the head of the speech audio stream starting from the first speech chunk.

[0106] Corresponding to the situation where the speech chunk at the head of the speech audio stream contains or does not contain the actual speech content, the recognition result of the current speech chunk at the head of the speech audio stream (e.g., the first speech chunk or the second speech chunk at the head, etc.) correspondingly contains or does not contain specific text characters, such as Chinese characters or English characters when converting speech to text, etc.

[0107] Step 402: Determine whether the processing result of the currently processed speech chunk contains recognized text characters.

[0108] Step 403: If the processing result of the currently processed speech chunk contains recognized text characters, determine that the currently processed speech chunk is the first-character speech chunk of the speech audio stream.

[0109] Among them, if the recognition result corresponding to the current speech chunk contains recognized text characters, it indicates that the first speech word of the speaker has appeared in the current speech chunk, and accordingly, the current speech chunk is determined as the first-word speech chunk of the speech audio stream.

[0110] Step 404: If the processing result of the currently processed speech chunk does not contain recognized text characters, at least the first processing unit processes the next speech chunk of the currently processed speech chunk, and updates the current speech chunk to the next speech chunk until the processing result of the speech chunk contains recognized text characters, and determines the speech chunk corresponding to the processing result containing text characters as the first-word speech chunk of the speech audio stream.

[0111] On the contrary, if the recognition result corresponding to the current speech chunk does not contain recognized text characters, it indicates that the first speech word of the speaker has not appeared from the first speech chunk of the speech audio stream to the current speech chunk. In this case, the first processing unit continues to quickly process the next speech chunk of the speech audio stream.

[0112] Until the processing result of the speech chunk contains recognized text characters, the speech chunk corresponding to the processing result containing text characters is determined as the first-word speech chunk of the speech audio stream.

[0113] Subsequently, further based on the first-word speech chunk that recognizes the actual speech content of the speaker, trigger the recognition of the first number of speech chunks at the head of the audio stream starting from the first-word speech chunk for the speech audio stream, and for each speech chunk in the first number of speech chunks, use the first processing unit of the speech audio stream to process it to reduce the processing time of the first number of speech chunks at the head of the audio stream and improve its processing speed. Until after processing the first number of speech chunks at the head of the audio stream starting from the first-word speech chunk, adjust to use the second processing unit of the speech recognition model to process each speech chunk in the middle position of the speech audio stream. By using the second processing unit to process each middle speech chunk, the overall recognition rate of the speech audio stream is guaranteed. Finally, when the speech recognition progress reaches the second number of speech chunks at the tail of the speech audio stream, adjust to use the first processing unit for processing.

[0114] It should be noted that the recognition of the first number of speech chunks at the head of the audio stream starting from the first-word speech chunk in the speech audio stream is synchronized with the encoding / decoding processing of the speech chunk, that is, for each obtained speech chunk, the position recognition of the speech chunk (recognizing whether it is the first number of speech chunks at the head) is completed synchronously, and the corresponding processing unit is used to perform encoding / decoding processing on it, which does not mean that after completing the recognition of all the first number of speech chunks at the head, the encoding / decoding processing of the first number of speech chunks is triggered.

[0115] Preferably, for the first speech block in the speech audio stream and each speech block before it, the encoding process of the next speech block starts only when the recognition result of the current speech block is obtained based on encoding and decoding. For the speech blocks after the first speech block in the speech audio stream, there is no restriction. The encoding process of the next speech block can start only when the recognition result of the current speech block is obtained based on encoding and decoding, or, after the encoding process of the current speech block is completed, the encoding process of the next speech block can be directly triggered without waiting for the decoding process of the current speech block to complete and obtain its recognition result to enter the encoding process flow of the next speech block.

[0116] Generally, the acquisition progress of the speech audio stream is higher than the recognition progress of the speech blocks in the speech audio stream by the speech recognition model. During the speech recognition process of the speech audio stream by the speech recognition model, for the currently to-be-recognized speech block obtained, it can be specifically determined whether the speech block is the tail (end) speech block of the speech audio stream by determining whether there are at least a preset number of subsequent speech frames after the moment corresponding to the speech block. Here, the preset number is related to the excess situation of the acquisition progress of the speech audio stream compared to the recognition progress (such as the number of excess frames), and can be specifically set according to this excess situation.

[0117] In this embodiment, by defining the first speech block and using a processing unit with a simpler structure (i.e., the first processing unit) in the speech recognition model to process the first speech block and each speech block before it in sequence, the processing rate of the speech audio stream with silence and / or noise at the starting position is further improved. It can make the silence and / or noise speech blocks at the starting position of the speech audio stream quickly pass through the speech recognition model, quickly advance the processing progress to the processing of the speech blocks containing actual speech content, and this fast processing of this part will not cause a decrease in the overall recognition rate of the speech audio stream.

[0118] In addition, in this embodiment, the first number of speech blocks at the head of the speech audio stream is defined by the defined first speech block, and on this basis, the first processing unit of the speech recognition model is used to process the first number of speech blocks at the head starting from the first speech block, reducing the first-word delay in speech recognition. At the same time, by using the first processing unit to process the tail speech blocks of the speech audio stream, the last-word delay in speech recognition is reduced. By reducing the first-word delay and the last-word delay, the speed index that measures the user experience in streaming speech recognition is improved, enhancing the user experience.

[0119] The following further illustrates the structure of the encoding and decoding unit of the speech recognition model through an embodiment.

[0120] Among them, for the implementation manners of improving the speech recognition model at least at the encoding end (such as the encoding end, or the encoding end and the decoding end), the encoding end of the speech recognition model includes a first encoding unit and a second encoding unit. The first encoding unit and the second encoding unit each include at least one encoding layer for performing speech encoding. The compositional structures of the first encoding unit and the second encoding unit are different, specifically referring to that the first encoding unit and the second encoding unit respectively form different overall encoding processing structures based on the encoding layers they contain. The encoding layer structures of different encoding layers in the same encoding unit may be the same or different, without limitation.

[0121] The first encoding unit and the second encoding unit respectively form different overall encoding processing structures based on the encoding layers they contain, which may be at least one of the following situations:

[0122] 11) The number of encoding layers of the first encoding unit is lower than the number of encoding layers of the second encoding unit;

[0123] 12) The compositional structure complexity of the encoding layers of the first encoding unit is lower than the compositional structure complexity of the encoding layers of the second encoding unit.

[0124] That is to say, in the embodiments of the present application, by streamlining in terms of the number of encoding layers and / or the compositional structure of the encoding layers included in the encoding unit, the first encoding unit has a simpler structure and a faster processing speed compared to the second encoding unit, and the premise is that the first encoding unit can still meet the basic encoding processing requirements in speech recognition processing and will not cause the inability to implement speech recognition (but it is allowed that there is an error recognition rate within the set tolerance limit in speech recognition based on the first encoding unit).

[0125] The first encoding unit and the second encoding unit may share at least one encoding layer, or may not share any encoding layer.

[0126] See Figure 5 For the example of, in the speech recognition model provided by this example, each encoding layer corresponding to Nx constitutes the first encoding unit, and each encoding layer corresponding to Nx and Mx constitutes the second encoding unit. The first encoding unit and the second encoding unit share each encoding layer corresponding to Nx (that is, each encoding layer included in the first encoding unit). This example simplifies the number of encoding layers, making the structure of the first encoding unit simpler than that of the second encoding unit and having a faster processing speed for speech blocks.

[0127] Similarly, for embodiments that improve the speech recognition model at least at the decoding end (such as the decoding end, or both the encoding end and the decoding end), the decoding end of the speech recognition model includes a first decoding unit and a second decoding unit. The first decoding unit and the second decoding unit each include at least one decoding layer for performing speech decoding. The compositional structures of the first decoding unit and the second decoding unit are different, specifically referring to the fact that the first decoding unit and the second decoding unit respectively form different overall decoding processing structures based on the decoding layers they contain. The decoding layer structures of different decoding layers in the same decoding unit can be the same or different, without limitation.

[0128] The first decoding unit and the second decoding unit respectively form different overall decoding processing structures based on the decoding layers they contain, which can be at least one of the following cases:

[0129] 21) The number of decoding layers in the first decoding unit is lower than the number of decoding layers in the second decoding unit;

[0130] 22) The compositional structure complexity of the decoding layers in the first decoding unit is lower than the compositional structure complexity of the decoding layers in the second decoding unit.

[0131] That is, by streamlining in terms of the number of decoding layers and / or the compositional structure of the decoding layers included in the decoding unit, the first decoding unit is simpler in structure and faster in processing speed compared to the second decoding unit, and the premise is that the first decoding unit can still meet the basic decoding processing requirements in speech recognition and will not cause the inability to achieve speech recognition (but it is allowed that there is an error rate in speech recognition based on the first decoding unit that does not exceed the set tolerance limit).

[0132] The first decoding unit and the second decoding unit can also share at least one decoding layer, or do not share any decoding layers.

[0133] Referring to FIG. 6(a), in the AED (Attention-based Encoder-Decoder) architecture provided by the embodiments of the present application, in an embodiment where at least the decoding end is improved, the compositional structure of a decoding layer in the second decoding unit, where the structures of each decoding layer in the second decoding unit are the same. FIG. 6(b) shows the compositional structures of two adjacent decoding layers in the first decoding unit in this architecture. Compared with the second decoding unit, the first decoding unit removes the calculation of the multi-head cross-attention every other decoding layer. As shown in the two-layer decoding layer structure in FIG. 6(b), the first decoding layer (the first layer) contains a sub-layer of the multi-head cross-attention (the sub-layers corresponding to the functional boxes with gray backgrounds), and the second decoding layer (the second layer) does not contain this sub-layer. And experiments prove that this sub-layer can be removed one by one every other decoding layer, which can achieve the purpose of accelerating the calculation and does not affect the basic decoding function of the decoding unit.

[0134] On the premise that the basic encoding / decoding functions can be achieved, this embodiment streamlines the number of encoding / decoding layers and / or the structure of the encoding / decoding layers of the encoding / decoding unit, so as to at least improve the processing speed of the speech blocks at the end positions of the speech audio stream, reduce the processing delay of the first / last words of the speech audio stream. At the same time, by using the encoding / decoding unit with a normal structure to process the speech blocks at the middle positions (positions other than the end positions defined above) of the speech audio stream, the overall recognition rate of the speech audio stream is ensured, taking into account the requirements for both recognition accuracy and recognition speed in streaming speech recognition.

[0135] In addition, the embodiment of the present application also discloses an electronic device, which may be, but is not limited to, a device in many general or specific computing device environments or configurations, such as: a personal computer, a server computer, a handheld device or a portable device, a tablet device, a multi-processor device, and so on.

[0136] The composition structure of the electronic device, as Figure 7 shown, at least includes:

[0137] A memory 10 for storing a computer instruction set;

[0138] The computer instruction set can be implemented in the form of a computer program.

[0139] A processor 20 for implementing the speech processing method disclosed in any of the above method embodiments by executing the computer instruction set.

[0140] The processor 20 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, etc.

[0141] The electronic device is equipped with a display device and / or has a display interface and can externally connect a display device.

[0142] Optionally, the electronic device further includes a camera component and / or is connected to an external camera component.

[0143] In addition, the electronic device may further include components such as a communication interface and a communication bus. The memory, the processor and the communication interface complete the communication with each other through the communication bus.

[0144] The communication interface is used for communication between the electronic device and other devices. The communication bus can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc.

[0145] It should be noted that the various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.

[0146] For the convenience of description, when describing the above system, device or equipment, it is divided into various modules or units according to functions for description. Of course, when implementing this application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0147] From the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0148] Finally, it should also be noted that in this article, relational terms such as first, second, third, and fourth are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0149] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A voice processing method, including: obtaining a voice block to be processed in the current voice to be processed; in response to the voice block to be processed being a voice block at an end position in the voice to be processed, performing voice recognition processing on the voice block to be processed at least through a first processing unit of a voice recognition model to obtain a first processing result; in response to the voice block to be processed being a voice block other than the end position in the voice to be processed, performing voice recognition processing on the voice block to be processed at least through a second processing unit of the voice recognition model to obtain a second processing result; wherein, the duration required for the first processing unit to process the voice block to be processed is lower than the duration required for the second processing unit to process the voice block to be processed, and the structural complexity of the first processing unit is lower than the structural complexity of the second processing unit.

2. The method according to claim 1, wherein the end position includes at least one of the following: the positions corresponding to a first number of voice blocks at the head of a voice audio stream; the positions corresponding to a second number of voice blocks at the tail of the voice audio stream; wherein, the first number of voice blocks at the head of the voice audio stream start from the first word voice block of the voice audio stream, and the first word voice block is the voice block corresponding to the voice audio stream and containing the first voice word of the voice audio stream.

3. The method according to claim 2, in the process of identifying the first word voice block of the voice audio stream, including: performing processing on the voice audio stream at least through the first processing unit starting from the first voice block of the voice audio stream to obtain a processing result of the currently processed voice block; if the processing result of the currently processed voice block contains recognized text characters, determining the currently processed voice block as the first word voice block of the voice audio stream; if the processing result of the currently processed voice block does not contain recognized text characters, performing processing on the next voice block of the currently processed voice block at least through the first processing unit until the processing result of the voice block contains recognized text characters, and determining the voice block corresponding to the processing result containing text characters as the first word voice block of the voice audio stream.

4. The method according to claim 1, wherein the voice recognition model includes a first encoding unit, a second encoding unit and a decoding unit, the first processing unit is the first encoding unit, and the second processing unit is the second encoding unit; the performing voice recognition processing on the voice block to be processed at least through the first processing unit of the voice processing model to obtain a first processing result includes: performing codec-based voice recognition on the voice block to be processed through a first encoding method corresponding to the first encoding unit of the voice recognition model and a decoding method corresponding to the decoding unit to obtain a first recognition result; the performing voice recognition processing on the voice block to be processed at least through the second processing unit of the voice recognition model to obtain a second processing result includes: Based on the second encoding method corresponding to the second encoding unit and the decoding method corresponding to the decoding unit of the voice recognition model, perform codec-based voice recognition on the to-be-processed voice block to obtain a second recognition result.

5. The method according to claim 1, wherein the voice recognition model includes an encoding unit, a first decoding unit, and a second decoding unit, the first processing unit is the first decoding unit, and the second processing unit is the second decoding unit; The at least performing voice recognition processing on the to-be-processed voice block through the first processing unit of the voice processing model to obtain a first processing result includes: Based on the encoding method corresponding to the encoding unit and the first decoding method corresponding to the first decoding unit of the voice recognition model, perform codec-based voice recognition on the to-be-processed voice block to obtain a first recognition result; The at least performing voice recognition processing on the to-be-processed voice block through the second processing unit of the voice recognition model to obtain a second processing result includes: Based on the encoding method corresponding to the encoding unit and the second decoding method corresponding to the second decoding unit of the voice recognition model, perform codec-based voice recognition on the to-be-processed voice block to obtain a second recognition result.

6. The method according to claim 1, wherein the voice recognition model includes a first encoding unit, a second encoding unit, a first decoding unit, and a second decoding unit, the first processing unit includes the first encoding unit and the first decoding unit, and the second processing unit includes the second encoding unit and the second decoding unit; The at least performing voice recognition processing on the to-be-processed voice block through the first processing unit of the voice processing model to obtain a first processing result includes: Based on the first encoding method corresponding to the first encoding unit and the first decoding method corresponding to the first decoding unit of the voice recognition model, perform codec-based voice recognition on the to-be-processed voice block to obtain a first recognition result; The at least performing voice recognition processing on the to-be-processed voice block through the second processing unit of the voice recognition model to obtain a second processing result includes: Based on the second encoding method corresponding to the second encoding unit and the second decoding method corresponding to the second decoding unit of the voice recognition model, perform codec-based voice recognition on the to-be-processed voice block to obtain a second recognition result.

7. The method according to claim 4 or 6, wherein, The first encoding unit and the second encoding unit each include at least one encoding layer for performing voice encoding; The first encoding unit and the second encoding unit satisfy at least one of the following conditions: The number of encoding layers of the first encoding unit is lower than the number of encoding layers of the second encoding unit; The compositional structure complexity of the encoding layers of the first encoding unit is lower than the compositional structure complexity of the encoding layers of the second encoding unit.

8. The method according to claim 5 or 6, wherein, The first decoding unit and the second decoding unit each include at least one decoding layer for performing voice decoding; The first decoding unit and the second decoding unit satisfy at least one of the following conditions: The number of decoding layers of the first decoding unit is lower than that of the second decoding unit; The compositional structure complexity of the decoding layers of the first decoding unit is lower than that of the decoding layers of the second decoding unit.

9. The method according to claim 7, wherein the first encoding unit and the second encoding unit share at least one encoding layer.

10. An electronic device, comprising: a memory for storing at least a set of instruction sets; a processor for calling and executing the instruction sets in the memory, and implementing the voice processing method according to any one of claims 1-9 by executing the instruction sets.

Citation Information

Patent Citations

  • Method of recognizing speech and electronic device thereof

    CN103544955A

  • Method and system for automatically recognizing voice

    CN103971686A