Speech recognition method, server and computer readable storage medium

By dynamically truncating the modeling unit in the streaming speech recognition system, the misrecognition problem caused by the window boundary cutting of the pronunciation unit is solved, and the accuracy and decoding efficiency of speech recognition are improved.

CN120299459APending Publication Date: 2025-07-11GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510610777.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the streaming speech recognition system leads to fragmentation of acoustic features when cutting the pronunciation unit at the window boundary, resulting in misrecognition and affecting the user experience.

Method used

By determining the voice request segment in the current time window, combining the preset classification model and the connection timing classification model, dynamically truncate the modeling unit located within the predetermined time period, delaying to the next time window decoding, avoiding mechanical splitting of the modeling unit.

Benefits of technology

It improves the recognition accuracy of voice request clips, optimizes decoding efficiency and accuracy, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299459A_ABST
    Figure CN120299459A_ABST
Patent Text Reader

Abstract

The invention discloses a voice recognition method, a server and a computer readable storage medium. The method comprises the following steps: determining a first voice request segment in a current time window; determining a first initial modeling unit set associated with the first voice request segment according to the first voice request segment; under the condition that a first target modeling unit exists in the first initial modeling unit set, the first modeling unit set is determined according to the first target modeling unit and the first initial modeling unit set, and the first target modeling unit is located in a preset time period away from the end moment of the current time window. And determining a speech recognition text in the current time window according to the first modeling unit set. Thus, in the streaming speech recognition processing process, the first target modeling unit is dynamically cut off and is delayed to the next time window for decoding, recognition errors caused by mechanical segmentation of the modeling unit are avoided, and therefore the recognition accuracy of the speech request segment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of voice interaction, and particularly to a voice recognition method, a server, and a computer-readable storage medium. Background Art

[0002] In the related art, when using a streaming voice recognition system to recognize a user's voice request, a voice request is often intercepted through a sliding window of a fixed length, and the intercepted voice request is decoded to obtain a voice recognition text corresponding to the user's voice request. However, in this case, when the window boundary exactly cuts an articulatory unit, it may cause fragmentation of acoustic features, resulting in the loss of information of the articulatory unit, leading to misrecognition and poor user experience. Summary of the Invention

[0003] The present application provides a voice recognition method, a server, and a computer-readable storage medium.

[0004] An embodiment of the present application provides a voice recognition method, the method including:

[0005] Determine a first voice request segment within a current time window;

[0006] Determine a first initial modeling unit set associated with the first voice request segment according to the first voice request segment;

[0007] In the case where there is a first target modeling unit in the first initial modeling unit set, determine a first modeling unit set according to the first target modeling unit and the first initial modeling unit set, where the first target modeling unit is located within a predetermined time period from the end moment of the current time window;

[0008] Determine a voice recognition text within the current time window according to the first modeling unit set.

[0009] In this way, the server determines the first voice request segment within the current time window. Then, based on the first voice request segment, the server determines the first set of initial modeling units associated with the first voice request segment. Subsequently, when there is a first target modeling unit in the first set of initial modeling units, based on the first target modeling unit and the first set of initial modeling units, the server determines the first set of modeling units, where the first target modeling unit is within a predetermined time period from the end time of the current time window. Finally, the server determines the speech recognition text within the current time window according to the first set of modeling units. In this way, during the streaming speech recognition process, by dynamically truncating the first target modeling unit within a predetermined time period from the end time of the current time window and delaying its decoding to the next time window, the recognition error caused by the mechanical segmentation of the modeling unit is avoided, thereby improving the recognition accuracy of the voice request segment.

[0010] In some embodiments, the determining, according to the first voice request segment, the first set of initial modeling units associated with the first voice request segment includes:

[0011] When there is a second target modeling unit in the second set of initial modeling units, based on the second voice request sub-segment and the first voice request segment, the target voice request segment is determined, where the second set of initial modeling units is determined according to the second voice request segment within the previous time window, the previous time window is adjacent to the current time window, the second target modeling unit is within the predetermined time period from the end time of the previous time window, and the second voice request sub-segment is the voice request sub-segment within the predetermined time period from the end time of the previous time window;

[0012] Based on the target voice request segment, the first set of initial modeling units is determined.

[0013] Thus, in the case where there is a second target modeling unit in the second initial modeling unit set, the server determines a target voice request segment according to the second voice request sub-segment and the first voice request segment, where the second initial modeling unit set is determined according to the second voice request segment within the previous time window, the previous time window is adjacent to the current time window, the second target modeling unit is located within a predetermined time period from the end time of the previous time window, and the second voice request sub-segment is a voice request sub-segment within a predetermined time period from the end time of the previous time window. Then, the server determines a first initial modeling unit set according to the target voice request segment. In this way, in the case where there is a second target modeling unit in the second initial modeling unit set, by combining the second voice request sub-segment associated with the second target modeling unit and the first voice request segment, the target voice request segment is determined to process the second voice request sub-segment and the first voice request segment delayed to the current time window, and the first initial modeling unit set is determined, avoiding the mechanical segmentation of the modeling unit and improving the recognition accuracy of the second voice request sub-segment.

[0014] In some embodiments, the determining the first initial modeling unit set according to the target voice request segment includes:

[0015] Based on a preset classification model, parsing the target voice request segment to determine a modeling unit associated with the current time window, where the modeling unit includes Chinese characters and / or English words;

[0016] Determining the first initial modeling unit set according to the modeling unit.

[0017] Thus, based on the preset classification model, the server parses the target voice request segment to determine a modeling unit associated with the current time window, and the modeling unit includes Chinese characters and / or English words. Then, the server determines the first initial modeling unit set according to the modeling unit. In this way, by parsing the target voice request segment through the preset classification model to determine the first initial modeling unit set, the recognition accuracy of subsequent voice segments can be improved. Moreover, the determined first initial modeling unit can be used in subsequent steps.

[0018] In some embodiments, the determining the first initial modeling unit set associated with the first voice request segment according to the first voice request segment includes:

[0019] In the case where there is no second target modeling unit in the second initial modeling unit set, based on the preset classification model, parsing the first voice request segment to determine the modeling unit;

[0020] Determining the first initial modeling unit set according to the modeling unit.

[0021] Thus, in the case where there is no second target modeling unit in the second initial modeling unit set, based on a preset classification model, the server analyzes the first voice request segment to determine a modeling unit. Then, according to the modeling unit, a first initial modeling unit set is determined. In this way, in the case where there is no second target modeling unit in the second initial modeling unit set, directly analyzing the first voice request segment through the preset classification model can improve the accuracy of subsequent voice segment recognition.

[0022] In some embodiments, in the case where there is a first target modeling unit in the first initial modeling unit set, determining a first modeling unit set according to the first target modeling unit and the first initial modeling unit set includes:

[0023] Determining the first modeling unit set according to the modeling units in the first initial modeling unit set except the first target modeling unit.

[0024] Thus, the server determines the first modeling unit set according to the modeling units in the first initial modeling unit set except the first target modeling unit. In this way, by removing redundant modeling units that need to be processed across windows in the current window, the quality of the first initial modeling unit set is optimized, and recognition errors caused by mechanical segmentation of modeling units are avoided, thereby improving the decoding efficiency and accuracy.

[0025] In some embodiments, determining the speech recognition text within the current time window according to the first modeling unit set includes:

[0026] Determining the speech recognition text according to the first modeling unit set and a second modeling unit set, where the second modeling unit set is determined according to the second initial modeling unit set.

[0027] Thus, the server determines the speech recognition text according to the first modeling unit set and the second modeling unit set, and the second modeling unit set is determined according to the second initial modeling unit set. In this way, by combining the first modeling unit set and the second modeling unit set, the model can comprehensively understand the voice segment, thereby improving the recognition accuracy of the speech recognition text.

[0028] In some embodiments, determining the speech recognition text according to the first modeling unit set and the second modeling unit set includes:

[0029] Determining a historical modeling unit set according to the second modeling unit set and the first modeling unit set;

[0030] Perform auxiliary decoding processing on the first set of modeling units according to the set of historical modeling units to determine the speech recognition text.

[0031] In this way, the server determines the set of historical modeling units according to the second set of modeling units and the first set of modeling units. Then, the server performs auxiliary decoding processing on the first set of modeling units according to the set of historical modeling units to determine the speech recognition text. In this way, by using the set of historical modeling units, the ability to understand the speech segment within the current time window is improved, the decoding result is optimized, and thus the recognition accuracy is improved.

[0032] In some embodiments, the method further includes:

[0033] In the case where the first target modeling unit does not exist in the first set of initial modeling units, determine the first set of initial modeling units as the first set of modeling units.

[0034] In this way, in the case where the first target modeling unit does not exist in the first set of initial modeling units, the server determines the first set of initial modeling units as the first set of modeling units. In this way, when the first target modeling unit does not exist in the first set of initial modeling units, directly determining the first set of initial modeling units as the first set of modeling units can simplify the decoding process and improve the decoding efficiency.

[0035] An embodiment of the present application provides a server, which includes a processor and a memory. A computer program is stored on the memory. When the computer program is executed by the processor, the above method is implemented.

[0036] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.

[0037] The additional aspects and advantages of the embodiments of the present application will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where:

[0039] Figure 1 is one of the schematic flowcharts of the speech recognition method of some embodiments of the present application;

[0040] Figure 2 is the schematic diagram of speech request segment segmentation of some embodiments of the present application;

[0041] Figure 3 This is the second flow chart of the speech recognition method of certain embodiments of the present application;

[0042] Figure 4 This is a third flow chart of a speech recognition method according to some embodiments of the present application;

[0043] Figure 5 This is a fourth flow chart of a speech recognition method according to some embodiments of the present application;

[0044] Figure 6 This is a fifth flow chart of a speech recognition method according to certain embodiments of the present application;

[0045] Figure 7 This is the sixth flow chart of the speech recognition method of certain embodiments of the present application;

[0046] Figure 8 FIG7 is a flowchart of a speech recognition method according to some embodiments of the present application;

[0047] Figure 9 This is the eighth flow chart of the speech recognition method of certain embodiments of the present application. DETAILED DESCRIPTION

[0048] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of the present application, and cannot be understood as limiting the embodiments of the present application.

[0049] In the related art, the streaming speech recognition system usually uses a sliding window of fixed length (such as 200-300ms) to segment the user voice stream, and performs real-time decoding based on the acoustic features in the window (such as MFCC and Mel spectrum, etc.) to generate incremental recognition results.

[0050] However, although this processing method can ensure the real-time performance of speech recognition, it also has obvious limitations. When the boundary of the sliding window happens to be in the middle of a pronunciation unit, the pronunciation unit will be divided into two parts, each falling into a different window. Since the segmented segments lack complete acoustic features, it is difficult for the model to accurately identify the pronunciation unit, which leads to misrecognition. For example, when recognizing "你先送他一枝花", if the window boundary happens to be in the middle of the pronunciation of the word "先(xian)", the model may mistakenly recognize the word "先(xian)" as "嘻(xi)", and recognize "安(an)" in the next time window, causing the semantics of the entire sentence to change, affecting the user experience.

[0051] Based on the above problems, please refer to Figure 1 , an embodiment of the present application provides a speech recognition method, and the method includes:

[0052] 01: Determine a first speech request segment within the current time window;

[0053] 02: Determine a first initial modeling unit set associated with the first speech request segment according to the first speech request segment;

[0054] 03: When there is a first target modeling unit in the first initial modeling unit set, determine a first modeling unit set according to the first target modeling unit and the first initial modeling unit set;

[0055] 04: Determine the speech recognition text within the current time window according to the first modeling unit set.

[0056] An embodiment of the present application further provides a server, including a memory and a processor. The speech recognition method of the embodiment of the present application can be implemented by the server of the embodiment of the present application. Specifically, a computer program is stored in the memory, and the processor is used to determine a first speech request segment within the current time window. And determine a first initial modeling unit set associated with the first speech request segment according to the first speech request segment. The processor is further used to determine a first modeling unit set according to the first target modeling unit and the first initial modeling unit set when there is a first target modeling unit in the first initial modeling unit set. And determine the speech recognition text within the current time window according to the first modeling unit set.

[0057] An embodiment of the present application further provides a speech recognition device. The speech recognition method of the embodiment of the present application can be implemented by the speech recognition device of the embodiment of the present application. Specifically, the speech recognition device includes a determination module. The determination module is used to determine a first speech request segment within the current time window. And determine a first initial modeling unit set associated with the first speech request segment according to the first speech request segment. The determination module is further used to determine a first modeling unit set according to the first target modeling unit and the first initial modeling unit set when there is a first target modeling unit in the first initial modeling unit set. And determine the speech recognition text within the current time window according to the first modeling unit set.

[0058] Specifically, the first speech request segment refers to the speech request segment currently being processed by the streaming speech recognition system. The streaming speech recognition system usually uses a sliding window with a fixed length (such as 200 - 300 ms) to segment and intercept the user speech stream to determine audio data with a fixed duration. For example, please refer to Figure 2, the voice request is "you give her a flower first", the streaming speech recognition system may divide the above voice request into three voice request segments, the voice request segment in the [0, T1) time window is "you + the first half of xian (xi)", the voice request segment in the [T1-T2) time window is "the second half of xian (an) + send her a + zhi", and the voice request segment in the [T2-T3) time window is "flower". If the current time window being processed is the [T1-T2) time window, then the first voice request segment is "the second half of xian (an) + send her a + zhi". It should be noted that the above-mentioned intercepted voice request segments are not actual specific Chinese characters or English words, but are actually composed of audio data corresponding to Chinese characters or English words, which are specifically manifested in data forms such as original waveforms and acoustic features.

[0059] The first initial modeling unit set refers to the processing of the first voice request segment based on the introduced Connectionist Temporal Classification (CTC) model to determine the preliminary modeling unit set associated with the first voice request segment. The Connectionist Temporal Classification model is used to determine whether there is a modeling unit in the current frame. If the current frame corresponds to a basic modeling unit, a peak is output, otherwise, blank is output. Continuing with the above example, the first voice request segment is "the second half of xian (an) + send her one + zhi", and the output of the Connectionist Temporal Classification model may be "peak peak peak__peak peak..._peak peak", then the associated first initial modeling unit set is ["first", "send", "her", "one" and "zhi"]. Among them, the modeling unit "zhi" is incomplete in semantics and may be a homophone, so the specific Chinese character cannot be determined based on the current information alone.

[0060] The modeling unit refers to the basic unit used to map continuous speech signals into discrete semantic symbols, including Chinese characters or English words.

[0061] The first target modeling unit refers to the modeling units in the first initial modeling unit set that are within a predetermined time period from the end time of the current time window. That is, as long as the time frames corresponding to the modeling units parsed from the first speech request segment are within the predetermined time period from the end time of the current time window, then these modeling units are considered as the first target modeling units. It should be noted that the predetermined time period can be determined according to actual needs. Continuing with the above example, the first initial modeling unit set is ["xian", "song", "ta", "yi", and "zhi"]. If only the modeling unit "zhi" is within the predetermined time period from the end time of the current time window, that is, the peak corresponding to the modeling unit "zhi" is within the predetermined time period from the end time of the current time window, then the first target modeling unit is "zhi". If both the modeling unit "zhi" and the modeling unit "yi" are within the predetermined time period from the end time of the current time window, then the first target modeling units are "zhi" and "yi". The following takes the first target modeling unit as "zhi" to illustrate the speech recognition method provided by the embodiments of the present application.

[0062] The first modeling unit set refers to the modeling unit set determined after a certain process on the first initial modeling unit set. Continuing with the above example, the first initial modeling unit set is ["xian", "song", "ta", "yi", and "zhi"], then the corresponding first modeling unit set is ["xian", "song", "ta", and "yi"].

[0063] The speech recognition text refers to the text that can be understood and processed by the large language model determined after decoding the first speech request segment within the current time window. Continuing with the above example, the first speech request segment is "the second half of 'xian' (an) + song ta yi + zhi", and the corresponding speech recognition text is "xian song ta yi".

[0064] First, the server intercepts a speech request segment from the continuous speech stream through a sliding window with a fixed duration as the first speech request segment for current analysis and decoding, such as "the second half of 'xian' (an) + song ta yi + zhi".

[0065] Next, the server performs parsing processing on the first speech request segment to determine all the modeling units associated with the first speech request segment and determines the first initial modeling unit set, such as ["xian", "song", "ta", "yi", and "zhi"].

[0066] Then, it is determined whether there is a first target modeling unit in the first set of initial modeling units, and the first target modeling unit is within a predetermined time period from the end time of the first voice request segment. If there is a first target modeling unit in the first set of initial modeling units, then based on the first target modeling unit and the first set of initial modeling units, a first set of modeling units for converting the speech recognition text is determined. Such as ["first", "send", "her", and "one"].

[0067] Finally, the server determines the speech recognition text within the current time window according to the first set of modeling units.

[0068] In summary, in the speech recognition method and server provided by the embodiments of the present application, the server determines the first voice request segment within the current time window. Then, the server determines a first set of initial modeling units associated with the first voice request segment according to the first voice request segment. Then, in the case where there is a first target modeling unit in the first set of initial modeling units, based on the first target modeling unit and the first set of initial modeling units, the server determines a first set of modeling units, where the first target modeling unit is within a predetermined time period from the end time of the current time window. Finally, the server determines the speech recognition text within the current time window according to the first set of modeling units. In this way, during the streaming speech recognition process, by dynamically truncating the first target modeling unit within a predetermined time period from the end time of the current time window and delaying it to the next time window for decoding, the recognition error caused by the mechanical segmentation of the modeling unit is avoided, thereby improving the recognition accuracy of the voice request segment.

[0069] Please refer to Figure 3 , in some embodiments, step 02 (determining a first set of initial modeling units associated with the first voice request segment according to the first voice request segment) includes:

[0070] 021: In the case where there is a second target modeling unit in the second set of initial modeling units, determine a target voice request segment according to the second sub-voice request segment and the first voice request segment;

[0071] 022: Determine the first set of initial modeling units according to the target voice request segment.

[0072] In some embodiments, the determining module is further configured to, in the case where there is a second target modeling unit in the second set of initial modeling units, determine a target voice request segment according to the second sub-voice request segment and the first voice request segment. And determine the first set of initial modeling units according to the target voice request segment.

[0073] In some embodiments, the processor is further configured to, in the case where a second target modeling unit exists in the second set of initial modeling units, determine a target voice request segment according to the second sub-segment of the voice request and the first voice request segment. And determine the first set of initial modeling units according to the target voice request segment.

[0074] Specifically, the second set of initial modeling units refers to a set of preliminary modeling units associated with the second voice request segment determined by processing the second voice request segment in the previous time window adjacent to the current time window based on the introduced connectionist temporal classification model. Continuing with the above example, if the second voice request segment is "the first half of 'ni + xian' (xi)", then the associated first set of initial modeling units is [ni and "xi"]. Please refer to Figure 2 and it can be found that the peak corresponding to "xian" is within a predetermined time period from the end time of the previous time window, so "xian" is the second target modeling unit.

[0075] The second target modeling unit refers to an undecoded modeling unit in the second set of initial modeling units that is within a predetermined time period from the end time of the previous time window.

[0076] The second sub-segment of the voice request refers to the voice segment within a predetermined time period at the end of the previous time window, that is, the voice request segment in the second voice request segment that has not completed voice recognition, including the undecoded modeling units that cooperate with the first set of initial modeling units, which can reduce semantic errors and semantic fragmentation caused by mechanical segmentation. The second voice request segment refers to the voice request segment in the previous time window.

[0077] The target voice request segment refers to the voice segment finally determined for decoding, and in the case where a second target modeling unit exists in the second set of initial modeling units, it includes information from the first voice request segment and the second sub-segment of the voice request.

[0078] First, check the second set of initial modeling units in the previous time window. If a second target modeling unit exists in the second set of initial modeling units in the previous time window, then splice the second sub-segment of the voice request and the first voice request segment to determine the target voice request segment finally used for decoding.

[0079] Then, the server determines the first set of initial modeling units according to the target voice request segment.

[0080] Continuing with the above example, it can be found that there is a second target modeling unit "the first half of xian (xi)" in the second set of initial modeling units. The second voice request segment is "you + the first half of xian (xi)", and the first voice request segment is "the second half of xian (an) + send her one + zhi". Then the target voice request segment is the first voice request segment "the second half of xian (an) + send her one + zhi" + the second voice request sub-segment "the first half of xian (xi)", that is, the voice request segment corresponding to "xian + send her + zhi". Furthermore, the first set of initial modeling units is determined as ["xian", "send", "her", "one", and "zhi"].

[0081] In this way, when there is a second target modeling unit in the second set of initial modeling units, the server determines the target voice request segment based on the second voice request sub-segment and the first voice request segment. Among them, the second set of initial modeling units is determined according to the second voice request segment within the previous time window. The previous time window is adjacent to the current time window. The second target modeling unit is located within a predetermined time period from the end time of the previous time window. The second voice request sub-segment is the voice request sub-segment within a predetermined time period from the end time of the previous time window. Then, the server determines the first set of initial modeling units according to the target voice request segment. In this way, when there is a second target modeling unit in the second set of initial modeling units, by combining the second voice request sub-segment associated with the second target modeling unit and the first voice request segment, the target voice request segment is determined to process the second voice request sub-segment and the first voice request segment delayed to the current time window, and the first set of initial modeling units is determined, avoiding the mechanical segmentation of modeling units and improving the accuracy of recognizing the second voice request sub-segment.

[0082] Please refer to Figure 4 , in some embodiments, step 022 (determining the first set of initial modeling units according to the target voice request segment) includes:

[0083] 0221: Based on a preset classification model, parse the target voice request segment to determine the modeling units associated with the current time window;

[0084] 0222: Determine the first set of initial modeling units according to the modeling units.

[0085] In some embodiments, the determining module is further configured to parse the target voice request segment based on a preset classification model to determine the modeling units associated with the current time window. And determine the first set of initial modeling units according to the modeling units.

[0086] In some embodiments, the processor is further configured to parse the target voice request segment based on a preset classification model to determine the modeling units associated with the current time window, and determine a first initial modeling unit set according to the modeling units.

[0087] Specifically, the preset classification model refers to a Connectionist Temporal Classification (CTC) model, which can determine whether the current frame corresponds to a modeling unit. If the current frame corresponds to a basic modeling unit, it outputs "peak"; otherwise, it outputs "blank". Moreover, "blank" is used to connect different "peaks" to represent the separation between different modeling units.

[0088] After inputting the target voice request segment into the preset classification model, the preset classification model will parse the target voice request segment to identify all the modeling units associated with the current time window.

[0089] Then, a first initial modeling unit is determined according to the identified modeling units.

[0090] Continuing with the above example, the Connectionist Temporal Classification model parses the voice request segment corresponding to "xian + song ta + zhi", and determines the modeling units associated with the current time window, including "xian", "song", "ta", "yi", and "zhi".

[0091] Then, according to the above modeling units, the first initial modeling unit set is determined as ["xian", "song", "ta", "yi", and "zhi"].

[0092] In this way, based on the preset classification model, the server parses the target voice request segment to determine the modeling units associated with the current time window. The modeling units include Chinese characters and / or English words. Then, the server determines a first initial modeling unit set according to the modeling units. In this way, by parsing the target voice request segment through the preset classification model to determine the first initial modeling unit set, the accuracy of voice segment recognition can be improved. Moreover, the determined first initial modeling units can be used in subsequent steps.

[0093] Please refer to Figure 5 , in some embodiments, step 02 (determining a first initial modeling unit set associated with the first voice request segment according to the first voice request segment) includes:

[0094] 023: In the case where the second target modeling unit does not exist in the second initial modeling unit set, parse the first voice request segment based on the preset classification model to determine the modeling units;

[0095] 024: Determine a first initial modeling unit set according to the modeling units.

[0096] In some embodiments, the determining module is further configured to, when a second target modeling unit does not exist in the second initial modeling unit set, parse the first voice request segment based on a preset classification model to determine the modeling units, and determine a first initial modeling unit set according to the modeling units.

[0097] In some embodiments, the processor is further configured to, when a second target modeling unit does not exist in the second initial modeling unit set, parse the first voice request segment based on a preset classification model to determine the modeling units, and determine a first initial modeling unit set according to the modeling units.

[0098] Specifically, when a second target modeling unit does not exist in the second initial modeling unit set, that is, all the modeling units in the second initial modeling unit set are not within a predetermined time period from the end moment of the previous time window, then directly parse the first voice request segment based on the preset classification model to determine the modeling units. Furthermore, determine a first initial modeling unit set according to the modeling units.

[0099] Continuing with the above example, if it is determined that a second target modeling unit does not exist in the second initial modeling unit set, then directly parse the first voice request segment "the second half of 'xian' (an) + send her one + zhi" based on the preset classification model, and determine that the first initial modeling unit set is ["an", "send", "her", "one", and "zhi"].

[0100] In this way, when a second target modeling unit does not exist in the second initial modeling unit set, the server parses the first voice request segment based on the preset classification model to determine the modeling units. Then, determine a first initial modeling unit set according to the modeling units. In this way, when a second target modeling unit does not exist in the second initial modeling unit set, directly parsing the first voice request segment through the preset classification model can improve the accuracy of voice segment recognition.

[0101] Please refer to Figure 6 , in some embodiments, step 03 (when a first target modeling unit exists in the first initial modeling unit set, determine a first modeling unit set according to the first target modeling unit and the first initial modeling unit set) includes:

[0102] 031: Determine a first modeling unit set according to the modeling units in the first initial modeling unit set except the first target modeling unit.

[0103] In some embodiments, the determination module is used to determine the first modeling unit set based on the modeling units in the first initial modeling unit set except the first target modeling unit.

[0104] In some embodiments, the processor is further configured to determine the first modeling unit set based on the modeling units in the first initial modeling unit set except the first target modeling unit.

[0105] Specifically, after determining the first initial modeling unit set, if there is a first target modeling unit in the first initial modeling unit set, the first modeling unit set is determined based on the modeling units in the first initial modeling unit set except the first target modeling unit. In other words, the first modeling unit set is the required first modeling unit set after the first target modeling unit is removed from the first initial modeling unit set.

[0106] Continuing with the above example, the first initial modeling unit set is [“安”, “送”, “她”, “一” and “zhi”], and there is a first target modeling unit [“zhi”], then the first modeling unit set is [“安”, “送”, “她” and “一”].

[0107] In this way, the server determines the first modeling unit set based on the modeling units in the first initial modeling unit set except the first target modeling unit. In this way, by removing the redundant modeling units that need to be processed across windows in the current window, the quality of the first initial modeling unit set is optimized, and the decoding efficiency and accuracy are improved.

[0108] See also Figure 7 In some embodiments, step 04 (determining the speech recognition text within the current time window according to the first set of modeling units) includes:

[0109] 041: Determine a speech recognition text according to the first modeling unit set and the second modeling unit set.

[0110] In some embodiments, the determination module is further configured to determine the speech recognition text based on the first set of modeling units and the second set of modeling units.

[0111] In some embodiments, the processor is further configured to determine speech recognition text based on the first set of modeling units and the second set of modeling units.

[0112] Specifically, the second modeling unit set refers to a modeling unit set determined after the second initial modeling unit set is processed in a certain manner.

[0113] To more accurately identify the speech segment within the current time window, it is necessary to consider its relevance to the speech segment within the previous time window. Specifically, we need to use the set of second modeling units corresponding to the speech segment recognized in the previous time window as context information to help understand and identify the first speech request segment in the current time window. That is to say, it is necessary to comprehensively consider the first modeling unit set and the second modeling unit set, understand and identify the first modeling unit set, and determine the speech recognition text.

[0114] In this way, the server determines the speech recognition text based on the first modeling unit set and the second modeling unit set, and the second modeling unit set is determined according to the second initial modeling unit set. In this way, by combining the first modeling unit set and the second modeling unit set, the model can comprehensively understand the speech segment, thereby improving the recognition accuracy of the speech recognition text.

[0115] Please refer to Figure 8 , in some embodiments, step 041 (determine the speech recognition text according to the first modeling unit set and the second modeling unit set) includes:

[0116] 0411: Determine the historical modeling unit set according to the second modeling unit set and the first modeling unit set;

[0117] 0412: Perform auxiliary decoding processing on the first modeling unit set according to the historical modeling unit set to determine the speech recognition text.

[0118] In some embodiments, the determination module is further configured to determine the historical modeling unit set according to the second modeling unit set and the first modeling unit set, and perform auxiliary decoding processing on the first modeling unit set according to the historical modeling unit set to determine the speech recognition text.

[0119] In some embodiments, the processor is further configured to determine the historical modeling unit set according to the second modeling unit set and the first modeling unit set, and perform auxiliary decoding processing on the first modeling unit set according to the historical modeling unit set to determine the speech recognition text.

[0120] Specifically, the decoding process refers to the process of converting the modeling unit set output by the speech recognition model into the final speech recognition text.

[0121] Determine the historical modeling unit set according to the modeling unit set corresponding to the speech segment recognized in the previous time window (the second modeling unit set) and the modeling unit set corresponding to the speech segment recognized in the current time window (the first modeling unit set). The historical modeling unit set includes historical information related to the speech segment in the current time window.

[0122] Next, by using the historical modeling unit set, perform auxiliary decoding processing on the first modeling unit set, and then determine the final speech recognition text according to the first modeling unit set after the auxiliary decoding processing.

[0123] In this way, the server determines the historical modeling unit set according to the second modeling unit set and the first modeling unit set. Next, the server performs auxiliary decoding processing on the first modeling unit set according to the historical modeling unit set to determine the speech recognition text. In this way, by using the historical modeling unit set, the understanding ability of the speech segment within the current time window is improved, the decoding result is optimized, and thus the recognition accuracy is improved.

[0124] Please refer to Figure 9 , in some embodiments, the method further includes:

[0125] 05: In the case where the first target modeling unit does not exist in the first initial modeling unit set, determine the first initial modeling unit set as the first modeling unit set.

[0126] In some embodiments, the determining module is used to determine the first initial modeling unit set as the first modeling unit set in the case where the first target modeling unit does not exist in the first initial modeling unit set.

[0127] In some embodiments, the processor is further used to determine the first initial modeling unit set as the first modeling unit set in the case where the first target modeling unit does not exist in the first initial modeling unit set.

[0128] Specifically, recognize the first speech request segment within the current time window to obtain the first initial modeling unit set. Then, determine whether the first target modeling unit exists in the first initial modeling unit set. If the first target modeling unit does not exist in the first initial modeling unit set, then determine the first initial modeling unit set as the first modeling unit set. Use the historical modeling unit set to perform auxiliary decoding on the first modeling unit set to obtain the final speech recognition text.

[0129] In this way, in the case where the first target modeling unit does not exist in the first initial modeling unit set, the server determines the first initial modeling unit set as the first modeling unit set. In this way, when the first target modeling unit does not exist in the first initial modeling unit set, directly determine the first initial modeling unit set as the first modeling unit set, which can simplify the decoding process and improve the decoding efficiency.

[0130] The following uses a complete example to explain the speech recognition method provided by the embodiments of the present application. Please refer to Figure 2 , where the preset time period is t milliseconds.

[0131] For the first time window [0, T), the output sequence of the connected time series classification model may appear as: peakpeak___…_peak peak. It can be found that the peak corresponding to the last modeling unit appears at the end of the window t milliseconds. The modeling unit at the end of the time window will not be decoded and will be decoded in the next time window. At this time, the decoder only performs a single-step decoding operation, and the decoding result of the first audio segment is the single word "you";

[0132] For the second window interval [T, 2T), the output sequence of the connected time series classification model may appear as: peakpeak peak__peak peak…_peak peak. At this stage, the model traces back the historical information of the acoustic features of the previous time window [0, T). After obtaining the complete context, the decoder begins to decode. After analysis, it is determined that the first word should be "先" (rather than the two-word combination of "嘻安"). At the same time, since the modeling unit peak within t milliseconds at the end of the time window [T, 2T) also appears at t milliseconds at the end of the window, only four-step decoding operations are performed in the current interval, and the decoding result of the second audio segment is "先送他一";

[0133] For the third window interval [2T, 3T), the speech recognition system integrates the acoustic feature history information of the previous time window [T, 2T) to identify the feature "zhi" that has not yet been decoded in the last modeling unit of the previous segment. After comprehensively analyzing the acoustic features of the current window [2T, 3T), it accurately determines that the undecoded unit corresponds to "枝" (as the counting unit of "花", which is different from the homophone "只"). Based on this, the model sets the current audio decoding step size to 3 steps, and finally generates the decoding result of the third audio segment as "枝花的".

[0134] The present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned speech recognition method are implemented.

[0135] It is understood that a computer program includes computer program code. The computer program code may be in source code form, object code form, executable file or some intermediate form. Computer readable storage media may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution medium.

[0136] In the description of this specification, the descriptions referring to terms such as "specifically", "furthermore", "specially", "understandably", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the illustrative expressions of the above terms are not necessarily intended to refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0137] Any process or method description shown in the flowchart or described in other ways herein can be understood to represent a module, segment, or portion of executable instructions including one or more steps for implementing a specific logical function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present application pertain.

[0138] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A voice recognition method, characterized in that, The method includes: Determine a first voice request segment within a current time window; According to the first voice request segment, determine a first set of initial modeling units associated with the first voice request segment; When there is a first target modeling unit in the first set of initial modeling units, according to the first target modeling unit and the first set of initial modeling units, determine a first set of modeling units, where the first target modeling unit is within a predetermined time period from the end time of the current time window; According to the first set of modeling units, determine the speech recognition text within the current time window.

2. The method according to claim 1, characterized in that, The step of, according to the first voice request segment, determining a first set of initial modeling units associated with the first voice request segment includes: When there is a second target modeling unit in a second set of initial modeling units, according to a second voice request sub-segment and the first voice request segment, determine a target voice request segment, where the second set of initial modeling units is determined according to a second voice request segment within a previous time window, the previous time window is adjacent to the current time window, the second target modeling unit is within the predetermined time period from the end time of the previous time window, and the second voice request sub-segment is a voice request sub-segment within the predetermined time period from the end time of the previous time window; According to the target voice request segment, determine the first set of initial modeling units.

3. The method according to claim 2, wherein The step of, according to the target voice request segment, determining the first set of initial modeling units includes: Based on a preset classification model, parse the target voice request segment to determine a modeling unit associated with the current time window, where the modeling unit includes Chinese characters and / or English words; According to the modeling unit, determine the first set of initial modeling units.

4. The method according to claim 3, wherein The step of, according to the first voice request segment, determining a first set of initial modeling units associated with the first voice request segment includes: When there is no second target modeling unit in the second set of initial modeling units, based on the preset classification model, parse the first voice request segment to determine the modeling unit; According to the modeling unit, determine the first set of initial modeling units.

5. The method according to claim 3, wherein The step of, when there is a first target modeling unit in the first set of initial modeling units, according to the first target modeling unit and the first set of initial modeling units, determining a first set of modeling units includes: According to the modeling units in the first set of initial modeling units except the first target modeling unit, determine the first set of modeling units.

6. The method according to claim 2, wherein The step of, according to the first set of modeling units, determining the speech recognition text within the current time window includes: According to the first set of modeling units and a second set of modeling units, determine the speech recognition text, where the second set of modeling units is determined according to the second set of initial modeling units.

7. The method according to claim 6, wherein The step of, according to the first set of modeling units and the second set of modeling units, determining the speech recognition text includes: Determine a historical modeling unit set according to the second modeling unit set and the first modeling unit set; Perform auxiliary decoding processing on the first modeling unit set according to the historical modeling unit set to determine the speech recognition text.

8. The method according to claim 1, wherein The method further includes: In the case where the first target modeling unit does not exist in the first initial modeling unit set, determine the first initial modeling unit set as the first modeling unit set.

9. A server, characterized in that, The server includes a processor and a memory, and a computer program is stored on the memory. When the computer program is executed by the processor, the method according to any one of claims 1-8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the steps of the method according to any one of claims 1-8 are implemented.