Segment based Continuous Speech Recognition Method
Patent Information
- Application Number
- KR1020250032051
- Authority / Receiving Office
- KR · KR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2026-09-21
Smart Images

Figure PAT00001_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a speech recognition device and method, and more specifically, to a segment-unit continuous speech recognition device and method that recognizes and buffers segments of a preset size from a speech signal, removes overlapping sections from the recognized segments and the buffered segments, and then outputs the result. Background Technology
[0002] Speech recognition technology has continued to advance, demonstrating high recognition performance for conversational speech, such as recordings of broadcasts and conference speech, as well as phone conversations between people, and is being widely utilized in various other fields.
[0003] In particular, transformer-based end-to-end speech recognition systems, which are being actively developed recently, demonstrate higher recognition performance than conventional speech recognition technologies that combine language and acoustic models by training a single neural network using speech signals and corresponding text, thereby eliminating the need for language-specific expertise.
[0004] Although transformer-based speech recognition technology demonstrates excellent performance, it has the problem that calculation can only begin when the entire utterance is input, and also has the problem that the amount of calculation is proportional to the square of the input signal length, so the calculation speed increases significantly as the input gets longer. The problem to be solved
[0005] The present invention was conceived from the technical background described above and aims to improve responsiveness by using the same structure and model as existing speech recognition systems, but performing recognition in segment units to quickly present recognition results to the user, and to perform speech recognition for long utterances more quickly with a small amount of computation while maintaining recognition performance.
[0006] The objectives of the present invention are not limited to those mentioned above, and other unmentioned objectives will be clearly understood by those skilled in the art from the description below. means of solving the problem
[0007] A segment-unit continuous speech recognition device according to one aspect of the present invention for achieving the aforementioned purpose comprises: a segment unit that divides an input signal into segments by creating overlapping sections within the input signal; a recognition unit that derives a result of recognizing each segment using an existing speech recognition device as is; a buffering unit that buffers the result of the recognized segment for subsequent use; a duplicate removal unit that removes a duplicate portion of the recognition result of the current segment from the recognition result of the previous segment; and an output unit that continuously outputs the recognition result with the duplicate sections removed.
[0008] Since the above duplicate removal unit has significant overlap between adjacent segments, there is significant overlap in the recognition results as well.
[0010] A segment-unit continuous speech recognition method according to another aspect of the present invention comprises: a step of dividing an input signal into segments of a predetermined length in a speech-unit speech recognition system; a step of performing speech recognition for each segment by creating overlapping intervals between each segment; a step of temporarily buffering the recognition results of each segment; a step of comparing the recognition results of each current segment with the recognition results of each previously buffered segment based on syllable-unit similarity; and a step of continuously performing recognition for the entire speech after deleting overlapping results from the comparison. Effects of the invention
[0011] According to the present invention, by dividing a voice signal into segments and buffering them, and then comparing the segments divided from the voice signal with the buffered segments to remove overlapping sections and outputting a recognition result, a recognition result can be derived even in the middle of a speech, and through repetitive recognition for short speeches, a high speed can be maintained regardless of the total length of the speech. Brief explanation of the drawing
[0012] FIG. 1 is a block diagram of a segment-unit continuous speech recognition device according to one embodiment of the present invention. FIG. 2 is a flowchart of a segment-unit continuous speech recognition method according to an embodiment of the present invention. FIG. 3 is an example of removing duplicate parts of the recognition results of the current segment and the previous segment according to an embodiment of the present invention and outputting only the correct recognition results. Specific details for implementing the invention
[0013] The advantages and features of the present invention and the methods for achieving them will become clear by referring to the embodiments described below in detail together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below but may be implemented in various different forms. These embodiments are provided merely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined only by the scope of the claims. Meanwhile, the terms used in this specification are for describing the embodiments and are not intended to limit the present invention. In this specification, the singular form includes the plural form unless specifically stated otherwise in the text. The terms "comprises" and / or "comprising" as used in this specification do not exclude the presence or addition of one or more other components, steps, actions, and / or elements in addition to the mentioned components, steps, actions, and / or elements.
[0015] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the attached drawings. FIG. 1 is a block diagram of a segment-unit continuous speech recognition device of the invention.
[0016] A segment-unit continuous speech recognition device according to one embodiment of the present invention includes a segment unit (100), a recognition unit (200), a buffering unit (300), a duplicate removal unit (400), and an output unit (500).
[0017] The segment unit (100) divides the input signal into segments of a preset length (S101). At this time, the segment unit (100) divides the first input signal into segments of a preset length, and from the subsequent segments onwards, divides the segments with an overlapping section in which a part of the previous segment overlaps. Here, the length of the segment is set to a preset time of about 2 seconds, and the overlapping section is set to about 1 second.
[0018] The recognition unit (200) recognizes each segment using a speech recognition device.
[0019] The buffering unit (300) buffers the result of the recognized segment for use next.
[0020] The duplicate removal unit (400) removes duplicate parts from the recognition results of the current segment in the recognition results of the buffered segment. That is, since there are significant duplicate sections between adjacent segments, the duplicate removal unit (400) also has significant duplicates in the recognition results.
[0021] Figure 2 is an example diagram showing an example of segmenting an input signal. Figure 2 shows an example of performing recognition on Korean with a total segment length of 2 seconds and an overlapping interval of 1 second.
[0022] Of these, “since China emitted 30 percent” (S10) is buffered as the previous segment output, and “since it borrowed 30 percent of global warming” (S20) is the current segment output.
[0023] In this way, the recognition results have a similarity greater than a preset level or significant overlap, such as the latter part of the previous segment's recognition result, "because it was discharged," and the first part of the current segment's recognition result, "because it was loaned," although not identical.
[0024] Accordingly, the duplicate removal unit (400) lists each recognition result in individual phoneme units, determines the duplicate section, and passes the result of removing it from the previous segment, “30 percent of warming,” to the next step.
[0025] The output unit (500) outputs the recognition result “30 percent of global warming” with the overlapping section removed.
[0026] According to one embodiment of the present invention, in order to perform speech recognition before speech ends, an input speech signal is segmented into pre-set lengths, buffered, and then output after removing the overlapping portion between the next segmented segment and the buffered previous segment, thereby having the effect of solving the problem of increased computational load due to speech length.
[0027] According to one embodiment of the present invention, there is an effect of solving the problem of difficulty in speech recognition even in the speech boundary region of a speech signal.
[0029] FIG. 3 is a flowchart of a segment-unit continuous speech recognition method according to an embodiment of the present invention.
[0030] The segment unit (100) divides the input signal into segments of a preset length (S101). At this time, the segment unit (100) divides the first input signal into segments of a preset length, and from the subsequent segments onwards, divides the segments with an overlapping section in which a part of the previous segment overlaps. Here, the length of the segment is set to a preset time of about 2 seconds, and the overlapping section is set to about 1 second.
[0031] The recognition unit (200) derives the result of recognizing each segment using a speech recognition device (S102).
[0032] The buffering unit (300) buffers the result of the recognized segment for use next (S103).
[0033] The duplicate removal unit (400) removes the duplicated parts of the recognition results of the current segment from the recognition results of the previous segment (S104). That is, since there is a significant overlap between adjacent segments, the duplicate removal unit (400) also has significant overlap in the recognition results. FIG. 2 shows an example of performing recognition on Korean with a segment length of 2 seconds and an overlap period of 1 second, where the latter part of the recognition results of the previous segment and the former part of the recognition results of the current segment have significant overlap in the recognition results. Accordingly, the duplicate removal unit (400) lists each recognition result in individual phoneme units, determines the overlap section, and passes the result of removing it from the previous segment to the next step.
[0034] The output unit (500) continuously outputs the recognition result with the duplicate section removed (S105).
[0036] Although the configuration of the present invention has been described in detail with reference to the accompanying drawings, this is merely illustrative, and it is obvious that those skilled in the art can make various modifications and changes within the scope of the technical concept of the present invention. Accordingly, the scope of protection of the present invention should not be limited to the aforementioned embodiments but should be determined by the description in the following claims.
Claims
Claim 1 A segment-unit continuous speech recognition method comprising: a step of dividing an input signal into segments of a fixed length in a speech unit speech recognition system; a step of performing speech recognition for each segment by placing overlapping intervals between each segment; a step of temporarily buffering the recognition results of each segment; a step of comparing the recognition results of each segment currently with the recognition results of each segment previously buffered based on syllable-unit similarity; and a step of continuously performing recognition for the entire speech after deleting overlapping results from the comparison.