Spoken dialogue reconstruction method, spoken dialogue reconstruction device, recording medium and computer program

By dividing and merging speaker-specific voice recognition data into blocks and reconstructing them into a dialogue format, the method effectively addresses the challenge of accurately representing multi-speaker conversations, achieving a dialogue structure that closely resembles actual dialogue flow with real-time confirmation and high readability.

JP7681266B2Active Publication Date: 2025-05-22LLSOLLU CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2021038052
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-10
Filing Date
2021-03-10
Publication Date
2025-05-22
Estimated Expiration
2041-03-10

AI Technical Summary

Technical Problem

Existing technologies face challenges in reconstructing speaker-specific speech recognition data into a dialogue format that accurately reflects the flow of an actual dialogue, especially in multi-speaker conversations with overlapping words.

Method used

A method and apparatus that acquire speaker-specific voice recognition data, divide it into blocks using token boundaries, sort and merge these blocks in time order regardless of speaker, and then reconstruct them into a dialogue format by classifying them by time order and speaker.

Benefits of technology

This approach enables a dialogue structure that closely mimics the flow of an actual dialogue, allowing for real-time confirmation of voice recognition results with minimal disruption to the dialogue structure and high readability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007681266000001
    Figure 0007681266000001
  • Figure 0007681266000002
    Figure 0007681266000002
  • Figure 0007681266000003
    Figure 0007681266000003
Patent Text Reader

Abstract

To provide a voice interaction reconstitution method for providing an interaction constitution nearest to a flow of interaction as much as possible.SOLUTION: A method comprises the steps of: acquiring voice recognition data by speaker for voice interaction; dividing the acquired voice recognition data by speaker into a plurality of blocks by using a boundary between tokens with division reference which is previously set; aligning the plurality of divided blocks in a time order regardless of the speaker; merging the plurality of blocks by continuous utterance by the same speaker for the plurality of aligned blocks; and separating the time order and the speaker and interactively reconstituting the plurality of blocks on which a result of merging is reflected.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a method and apparatus for reconstructing speaker-specific speech recognition data for a speech dialogue into a dialogue format. [Background technology]

[0002] Among the technologies for input processing of natural language, speech-to-text conversion (STT) is a speech recognition technology that converts speech into text.

[0003] These voice recognition technologies can be divided into two types based on their real-time nature: one type is a method that receives the voice to be converted all at once and converts it all at once, and the other type is a method that receives real-time generated voice in a predetermined unit (e.g., in units of less than one second) and converts it in real time.

[0004] Among them, the batch conversion method usually recognizes the entire input speech and then generates the result at once, while the real-time conversion method must define the time when the speech recognition result is generated.

[0005] There are three main ways to define the time when the recognition result is generated for the real-time conversion method. First, the recognition result can be generated when a special end signal (e.g., pressing the recognition / call end button) is input. Second, the recognition result can be generated when an EPD (End Point Detection) occurs, such as silence of a certain length (e.g., 0.5 seconds). Third, the recognition result can be generated at regular intervals.

[0006] Among them, the third method of defining the time when the recognition result is generated has the characteristic that the time when the recognition result is generated is the time when the connected words are not finished, i.e., in the middle of speaking. Therefore, it is mainly used when trying to temporarily obtain the recognition result from a certain point until the present, rather than when generating a formal result, and the result obtained in this way is called an incomplete result, not a completed recognition result.

[0007] Unlike recognition results based on EPD boundaries, incomplete results may contain previous generation results in the currently generated result. For example, EPD-based recognition results generate results such as "A, B, C", "D, E", and "F, G, H" to recognize "A, B, C, D, E, F, G, H", but incomplete results usually contain previous generation results such as "A", "AB", "ABC", "D", "D, E", "F", "F, G", and "F, G, H" unless an EPD occurs.

[0008] Meanwhile, although the accuracy of voice recognition technology has improved considerably in recent years, when recognizing conversations with multiple speakers, there are still challenges, such as the problem of recognizing voices in sections where words overlap when two or more people are speaking at the same time, and the problem of speaker identification, which requires distinguishing whose voice belongs to whom.

[0009] Therefore, in a commonly used system, a method is used in which a speaker uses a different input device to recognize a voice of each speaker, and a speaker-specific voice recognition data is generated and acquired.

[0010] In this way, when generating and acquiring voice recognition data for each speaker for a voice dialogue, it is necessary to reconstruct the acquired speaker-specific voice recognition data into a dialogue format, and technology for reconstructing speaker-specific voice recognition data into a dialogue format is being continuously researched. [Prior art documents] [Patent documents]

[0011] (Patent Document 1) Korean Patent Publication No. 10-2014-0078258 (published on June 25, 2014) Summary of the Invention [Problem to be solved by the invention]

[0012] According to an embodiment of the present invention, there is provided a method and apparatus for reconstructing a voice dialogue, which provides a dialogue structure as close as possible to the flow of an actual dialogue, in reconstructing speaker-specific voice recognition data for a voice dialogue into a dialogue format.

[0013] The problems to be solved by the present invention are not limited to those mentioned above, and other problems not mentioned or to be solved by the present invention will be clearly understood by those having ordinary skill in the art to which the present invention pertains from the following description. [Means for solving the problem]

[0014] A voice dialogue reconstructing method of the voice dialogue reconstructing device according to the first aspect includes the steps of: acquiring speaker-specific voice recognition data for a voice dialogue; dividing the acquired speaker-specific voice recognition data into a plurality of blocks using boundaries between tokens according to a preset division criterion; sorting the divided plurality of blocks in time order regardless of speaker; merging a plurality of blocks resulting from successive utterances by the same speaker into the sorted plurality of blocks; and reconstructing the plurality of blocks reflecting the result of the merging into a dialogue format by dividing them into the time order and speaker.

[0015] A voice dialogue reconstructing apparatus according to a second aspect includes an input unit to which a voice dialogue is input, and a processing unit to process voice recognition for the voice dialogue input through the input unit, wherein the processing unit acquires speaker-specific voice recognition data for the voice dialogue, divides the acquired speaker-specific voice recognition data into a plurality of blocks using boundaries between tokens according to a preset division criterion, sorts the divided plurality of blocks in time order regardless of speakers, merges a plurality of blocks resulting from successive utterances by the same speaker into the sorted plurality of blocks, and reconstructs the plurality of blocks reflecting the result of the merging into a dialogue format by classifying the blocks into the time order and the speakers.

[0016] According to a third aspect, a computer-readable recording medium storing a computer program includes instructions for causing the processor to execute a method, when the computer program is executed by a processor, comprising: acquiring speaker-specific speech recognition data for a voice dialogue; dividing the acquired speaker-specific speech recognition data into a plurality of blocks using inter-token boundaries according to a preset division criterion; sorting the divided plurality of blocks in time order regardless of speaker; merging a plurality of blocks resulting from consecutive utterances by the same speaker into the sorted plurality of blocks; and reconstructing the plurality of blocks reflecting the result of the merging into a dialogue format by dividing the blocks into the time order and speaker.

[0017] According to a fourth aspect, a computer program stored in a computer-readable recording medium includes instructions for causing a processor to execute a method, when executed by a processor, comprising: acquiring speaker-specific speech recognition data for a voice dialogue; dividing the acquired speaker-specific speech recognition data into a plurality of blocks using inter-token boundaries according to a preset division criterion; sorting the divided plurality of blocks in time order regardless of speaker; merging a plurality of blocks resulting from successive utterances by the same speaker into the sorted plurality of blocks; and reconstructing the plurality of blocks reflecting the result of the merging into a dialogue format by dividing the blocks into the time order and speaker. Effect of the Invention

[0018] According to an embodiment of the present invention, in reconstructing speaker-specific voice recognition data for a voice dialogue into a dialogue format, it is possible to provide a dialogue structure that is as close as possible to the flow of an actual dialogue.

[0019] In addition, the dialogue is reconstructed by reflecting partial results, which are voice recognition results generated at regular intervals during the voice dialogue, so that the converted dialogue can be confirmed in real time. Since the voice recognition results are reflected in real time, when such voice recognition results are output to the screen, the amount of dialogue updated at one time is small, so that the dialogue structure is not disrupted, and there is relatively little change in the reading position on the screen, providing high readability and recognizability. [Brief description of the drawings]

[0020] [Figure 1] 1 is a configuration diagram of a spoken dialogue reconstructing device according to an embodiment; [Diagram 2] 1 is a flow chart illustrating a method for reconstructing a voice dialogue according to an embodiment. [Diagram 3] 4 is a flow chart illustrating a process of acquiring voice recognition data for each speaker in a method for reconstructing a voice dialogue according to an embodiment of the present invention; [Figure 4] 1 is a diagram illustrating a result of voice dialogue reconstruction by the voice dialogue reconstructing device according to an embodiment; DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0021] The advantages and features of the present invention, as well as the methods for achieving them, will become clear from the following examples described with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below, and may be embodied in various different forms, and the embodiments are provided to fully disclose the present invention and to fully inform those skilled in the art of the invention, and the present invention is defined only by the scope of the claims.

[0022] The terms used in this specification will be briefly explained and the present invention will be specifically described.

[0023] The terms used in the present invention are selected as general terms currently widely used as much as possible while taking into consideration the functions of the present invention, but this may vary depending on the intentions or precedents of engineers in the field, the emergence of new technologies, etc. In addition, in certain cases, the applicant may arbitrarily select terms, and in such cases, their meanings will be described in detail in the corresponding description of the invention. Therefore, the terms used in the present invention must be defined based on the meanings of the terms and the overall content of the present invention, rather than simply the names of the terms.

[0024] Throughout the specification, when a part "comprises" a certain element, this does not mean to the exclusion of other elements, but may further include other elements, unless specifically stated to the contrary.

[0025] In addition, the term "module" as used in the specification means software or hardware components such as FPGAs and ASICs, and a "module" only performs a certain function, and is not limited to software or hardware. A "module" may be configured to reside on an addressable storage medium or to execute one or more processors. Thus, by way of example, a "module" may include components such as software components, object-oriented software components, class components, and task components, as well as processors, functions, attributes, procedures, subroutines, program code segments, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functionality provided in the components and modules may be combined into fewer components and modules, or may be further separated into additional components and modules.

[0026] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings so that those having ordinary knowledge in the technical field to which the present invention pertains can easily implement them. And, in order to clearly explain the present invention in the drawings, parts not related to the explanation are omitted.

[0027] Figure 1 is a configuration diagram of a voice dialogue reconstruction device according to an embodiment.

[0028] According to Figure 1, the voice dialogue reconstruction device 100 includes an input unit 110 and a processing unit 120, and can further include an output unit 130 and / or a storage unit 140. The processing unit 120 can include a speaker-specific data processing unit 121, a block division unit 122, a block alignment unit 123, a block merging unit 124, and a dialogue reconstruction unit 125.

[0029] The input unit 110 receives a voice dialogue. Such an input unit 110 can input voice data by voice dialogue separated for each speaker. For example, the input unit 110 can include a number of microphones corresponding one-to-one to the number of speakers.

[0030] The processing unit 120 processes speech recognition for the voice dialogue input through the input unit 110. For example, the processing unit 120 can include computer arithmetic means such as a microprocessor.

[0031] The speaker-specific data processing unit 121 of the processing unit 120 acquires speaker-specific voice recognition data for a voice dialogue. For example, the speaker-specific data processing unit 121 may include an ASR (Automatic Speech Recognition), which may extract a character string after removing noise from the speaker-specific voice data input through the input unit 110 through a pre-processing process. When acquiring speaker-specific voice recognition data, the speaker-specific data processing unit 121 may apply a plurality of recognition result generation time points. For example, the speaker-specific data processing unit 121 may generate a speaker-specific first recognition result for a voice dialogue in an EPD (End Point Detection) unit, and may generate a speaker-specific second recognition result for each preset time. For example, the speaker-specific second recognition result may be generated after the EPD for generating the speaker-specific first recognition result is last generated. The speaker-specific data processing unit 121 may generate speaker-specific voice recognition data only after collecting the speaker-specific first recognition result and the speaker-specific second recognition result for each speaker without overlapping or duplication. Of course, the speaker-specific data processing unit 121 may apply a single recognition result generation time point in acquiring speaker-specific voice recognition data, for example, only one of the speaker-specific first recognition result and the speaker-specific second recognition result may be generated.

[0032] The block division unit 122 of the processing unit 120 divides the speaker-specific speech recognition data acquired by the speaker-specific data processing unit 121 into a plurality of blocks using boundaries between tokens according to a preset division criterion. For example, the preset division criterion may be a silent section of a certain duration or a morphological characteristic with the previous token.

[0033] The block alignment unit 123 of the processing unit 120 aligns the multiple blocks divided by the block division unit 122 in time order, regardless of the speaker.

[0034] The block merging unit 124 of the processing unit 120 merges a plurality of blocks resulting from successive utterances by the same speaker for the plurality of blocks aligned by the block alignment unit 123 .

[0035] The dialogue reconstruction unit 125 of the processing unit 120 classifies the blocks reflecting the results of merging by the block merging unit 124 into a dialogue format by dividing them into time order and speaker.

[0036] The output unit 130 outputs the processing result by the processing unit 120. For example, the output unit 130 may include an output interface, and may output the converted data provided by the processing unit 120 to another electronic device connected to the output interface under the control of the processing unit 120. Alternatively, the output unit 130 may include a network card, and may transmit the converted data provided by the processing unit 120 through a network under the control of the processing unit 120. Alternatively, the output unit 130 may include a display device that can display the processing result by the processing unit 120 on a screen, and may display the voice recognition data for the voice dialogue reconstructed into a dialogue format by the dialogue reconstruction unit 125 on a screen in chronological order by dividing the speakers.

[0037] The storage unit 140 may store an operating system program for the spoken dialogue reconstruction device 100, and may also store processing results by the processing unit 120. For example, the storage unit 140 may be a computer-readable recording medium such as a magnetic medium such as a hard disk, a floppy disk, or a magnetic tape, an optical medium such as a CD-ROM or a DVD, a magneto-optical medium such as a floptical disk, or a hardware device specially configured to store and execute program instructions, such as a flash memory.

[0038] FIG. 2 is a flowchart illustrating a method for reconstructing a voice dialogue according to an embodiment, FIG. 3 is a flowchart illustrating a process of acquiring voice recognition data for each speaker in the method for reconstructing a voice dialogue according to an embodiment, and FIG. 4 is a diagram illustrating an example of a result of reconstructing a voice dialogue by the device for reconstructing a voice dialogue according to an embodiment.

[0039] Hereinafter, a voice dialogue reconstructing method executed by the voice dialogue reconstructing device 100 according to an embodiment of the present invention will be described in detail with reference to FIGS.

[0040] First, the input unit 110 receives voice data from a voice dialogue separated for each speaker, and provides the input voice data for each speaker to the processing unit 120 .

[0041] The speaker-specific data processing unit 121 of the processing unit 120 obtains speaker-specific voice recognition data for the voice dialogue. For example, the ASR included in the speaker-specific data processing unit 121 can obtain speaker-specific voice recognition data consisting of character strings by removing noise through a preprocessing process for the speaker-specific voice data input through the input unit 110 and then extracting character strings (S210).

[0042] Here, the speaker-specific data processing unit 121 applies a plurality of recognition result generation time points in acquiring speaker-specific voice recognition data. The speaker-specific data processing unit 121 generates a speaker-specific first recognition result in EPD units for the voice dialogue. At the same time, the speaker-specific data processing unit 121 generates a speaker-specific second recognition result at preset times after the last EPD for generating the speaker-specific first recognition result is generated (S211). The speaker-specific data processing unit 121 then collects the speaker-specific first recognition result and the speaker-specific second recognition result by speaker without overlapping or duplication to finally generate speaker-specific voice recognition data (S212).

[0043] In this way, the speaker-specific voice recognition data acquired by the speaker-specific data processing unit 121 is reconstructed into a dialogue format by the dialogue reconstruction unit 125. However, when reconstructing a dialogue format in text form, unlike voice, if a situation is assumed in which a second speaker's words come out, albeit briefly, while a first speaker is speaking, in order to express such a situation in text, it is necessary to determine whether to cut off the words in the middle and where to cut off. For example, after cutting off the words based on a silent section for the entire dialogue, data of all speakers can be collected and arranged in chronological order. In this case, if additionally recognized text occurs based on the EPD, only the length of the text is added to the screen at once, which may cause a problem that the position the user was reading is lost or the dialogue structure is changed. In addition, if the dialogue structure unit is not natural, the context of the dialogue may be lost. For example, if a second speaker says "yes" while a first speaker is speaking continuously, "yes" may not be expressed in the actual context position, but may be added to the end of the long continuous speech of the first speaker. In addition, if real-time performance is added, the recognition result cannot be confirmed on the screen until the EPD occurs, even if the speaker is speaking and recognition is being performed. In fact, even if the first speaker speaks first, the second speaker who speaks later has a short speech and finishes first, so the first speaker's speech is not displayed on the screen, and only the second speaker's speech is displayed. In order to deal with such various situations, the spoken dialogue reconstruction device 100 according to an embodiment performs a division process by the block division unit 122, an alignment process by the block alignment unit 123, and a merging process by the block merging unit 124. The division process and alignment process are for inserting speech of another speaker between speeches according to the flow of the original dialogue, and the merging process is for preventing sentences constituting the dialogue from being cut too short due to the division performed for the insertion.

[0044] The block division unit 122 of the processing unit 120 divides the speaker-specific voice recognition data acquired by the speaker-specific data processing unit 121 into a plurality of blocks using boundaries between tokens (e.g., words / phrases / morphemes) according to a preset division criterion, and provides the divided blocks to the block alignment unit 122 of the processing unit 120. For example, the preset division criterion may be a silent section of a certain duration or a morphological characteristic with a previous token (e.g., between phrases), and the block division unit 122 can divide the speaker-specific voice recognition data into a plurality of blocks using the silent section of a certain duration or the morphological characteristic with a previous token as a division criterion (S220).

[0045] Next, the block arrangement unit 123 of the processing unit 120 arranges the blocks divided by the block division unit 122 in chronological order regardless of the speaker, and provides the arranged blocks to the block merging unit 124 of the processing unit 120. For example, the block arrangement unit 123 can arrange the blocks based on the start time of each block, or can arrange the blocks based on the midpoint time of each block (S230).

[0046] Then, the block merging unit 124 of the processing unit 120 merges a plurality of blocks resulting from consecutive speech by the same speaker among the plurality of blocks aligned by the block alignment unit 123, and provides speaker-specific speech recognition data reflecting the result of the block merging to the dialogue reconstruction unit 125. For example, the block merging unit 124 can determine consecutive speech by the same speaker by using a silent period of a certain duration or less existing between the previous block and the block, or a syntactic characteristic with the previous block (e.g., when the previous block is the end of a sentence, etc.) (S240).

[0047] Next, the dialogue reconstruction unit 125 of the processing unit 120 reconstructs the multiple blocks reflecting the merging results by the block merging unit 124 into a dialogue format by dividing them into time order and speaker, and provides the reconstructed voice recognition data to the output unit 130 (S250).

[0048] Accordingly, the output unit 130 outputs the processing result by the processing unit 120. For example, the output unit 130 may output the converted data provided by the processing unit 120 to another electronic device connected to the output interface under the control of the processing unit 120. Alternatively, the output unit 130 may transmit the converted data provided by the processing unit 120 through a network under the control of the processing unit 120. Alternatively, the output unit 130 may display the processing result by the processing unit 120 on a screen of a display device as illustrated in FIG. 4. As illustrated in FIG. 4, the output unit 130 may display the voice recognition data for the voice dialogue reconstructed into a dialogue format by the dialogue reconstruction unit 125 on a screen in chronological order by dividing the speakers. Here, when updating and outputting the reconstructed voice recognition data, the output unit 130 may update and output a screen reflecting the first recognition result for each speaker generated in step S211. That is, in step S250, the dialogue reconstruction unit 125 provides the voice recognition data reflecting the first recognition result for each speaker to the output unit 130 (S260).

[0049] Meanwhile, each step included in the spoken dialogue reconstruction method according to the above-mentioned embodiment can be embodied in a computer-readable recording medium recording a computer program including instructions for executing such steps.

[0050] In addition, each step included in the spoken dialogue reconstruction method according to the above-mentioned embodiment may be embodied in the form of a computer program stored in a computer-readable recording medium programmed to include instructions for executing such steps.

[0051] As described above, according to the embodiment of the present invention, in reconstructing speaker-specific voice recognition data for a voice dialogue into a dialogue format, it is possible to provide a dialogue structure that is as close as possible to the flow of an actual dialogue.

[0052] And, since the conversation is reconstructed by reflecting the incomplete results, which are the speech recognition results generated at regular intervals during the voice conversation, it is possible to check the conversation converted in real time. Since the real-time speech recognition results are reflected, when such speech recognition results are output to the screen, the amount of conversation updated at one time is small, the structure of the conversation does not collapse, and the degree of change in the reading position on the screen is relatively small, providing high readability and recognizability.

[0053] The combination of each step of each flowchart attached to the present invention can also be executed by computer program instructions. Since these computer program instructions can be loaded onto the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing equipment, the instructions executed through the processor of the computer or other programmable data processing equipment generate means for executing the functions described in each step of the flowchart. Since these computer program instructions can be stored in a computer-usable or computer-readable recording medium that can be directed to a computer or other programmable data processing equipment in order to embody the functions in a specific manner, the instructions stored in the computer-usable or computer-readable recording medium can also produce a manufactured item that includes instruction means for executing the functions described in each step of the flowchart. Since the computer program instructions can also be loaded onto a computer or other programmable data processing equipment, a series of operation steps are executed on the computer or other programmable data processing equipment to generate a process executed by the computer, and the instructions for executing the computer or other programmable data processing equipment can also provide a number of steps for executing the functions described in each step of the flowchart.

[0054] Also, each step may represent a module, segment, or part of code that includes one or more executable instructions for performing a specified number of logical functions. It should also be noted that in some embodiments, the functions recited in the steps may occur out of order. For example, two steps shown in the following figures may be performed substantially simultaneously, or the steps may sometimes be performed in reverse order depending on the corresponding functions.

[0055] The above description is merely illustrative of the technical idea of ​​the present invention, and various modifications and variations are possible within the scope of the essential quality of the present invention, if one has ordinary knowledge in the technical field to which the present invention belongs. Therefore, the embodiments disclosed in the present invention are for the purpose of explanation, not for the purpose of limiting the technical idea of ​​the present invention, and the scope of the technical idea of ​​the present invention is not limited by such embodiments. The scope of protection of the present invention should be interpreted according to the claims, and all technical ideas within the scope equivalent thereto should be interpreted as being included in the scope of the present invention. [Explanation of symbols]

[0056] 100 Spoken dialogue reconstruction device 110 Input section 120 Processing section 121 Speaker-specific data processing section 122 Block division section 123 Block Alignment Section 124 Block Merger Section 125 Dialogue Reconstruction Department 130 Output section 140 Storage section

Claims

1. In a method for reconstructing a spoken dialogue by a spoken dialogue reconstructing device, acquiring a plurality of speaker-specific speech recognition data including a plurality of speaker-specific first recognition results generated in an EPD (End Point Detection) unit for a speech dialogue in which the speakers are different from each other and a plurality of speaker-specific second recognition results generated at preset time intervals; Dividing the speech recognition data for each of the plurality of speakers into a plurality of blocks using boundaries between tokens according to a preset division criterion; a step of arranging the divided blocks in time order regardless of the speaker for the speech recognition data for the plurality of speakers; a step of merging a plurality of blocks resulting from continuous speech by the same speaker into the plurality of aligned blocks; and reconstructing the plurality of blocks reflecting the result of the merging into a dialogue format by dividing the blocks into the time sequence and the speaker, The second speaker-specific recognition result is generated after a last EPD in which the first speaker-specific recognition result is generated is generated.

2. The step of acquiring speaker-specific speech recognition data includes:

2. The method of claim 1, further comprising the step of generating the speaker-specific speech recognition data by collecting the first speaker-specific recognition result and the second speaker-specific recognition result without overlapping or redundancy.

3. 2. The method of claim 1, wherein the preset division criterion is a silent section or a boundary between words for a certain period of time or more.

4. 2. The method for reconstructing spoken dialogue according to claim 1, wherein the step of merging comprises discriminating between successive utterances by the same speaker based on a silent period of a certain duration or less or on a syntactic characteristic with a previous block.

5. 3. The method for reconstructing a spoken dialogue according to claim 2, further comprising a step of outputting the speech recognition data reconstructed into the dialogue format on a screen, and updating the speaker-specific speech recognition data collectively or updating the data to reflect the first recognition result for each speaker when updating the screen.

6. an input unit for inputting a voice dialogue; a processing unit that processes voice recognition for the voice dialogue input through the input unit, The processing unit includes: a plurality of speech recognition data for each speaker, the plurality of speech recognition data including a first recognition result for each speaker generated in an EPD (End Point Detection) unit and a second recognition result for each speaker generated at a preset time, the plurality of speech recognition data for each speaker is divided into a plurality of blocks using a boundary between tokens according to a preset division criterion for each of the plurality of speech recognition data for each speaker; the plurality of divided blocks are arranged in time order regardless of a speaker; a plurality of blocks resulting from continuous utterances by the same speaker are merged with respect to the arranged blocks; and the plurality of blocks reflecting the result of the merging are reconstructed into a dialogue format by dividing the plurality of blocks into the time order and the speaker; The second speaker-specific recognition result is generated after a final EPD in which the first speaker-specific recognition result is generated is generated.

7. The processing unit includes:

7. The spoken dialogue reconstructing apparatus according to claim 6, wherein the speaker-specific speech recognition data is generated by collecting the speaker-specific first recognition result and the speaker-specific second recognition result without overlapping or redundancy.

8. A computer-readable recording medium storing a computer program, The computer program, when executed by a processor, and dividing the plurality of blocks into a plurality of blocks using inter-token boundaries according to a predetermined division criterion for each of the plurality of speaker-specific speech recognition data. The plurality of blocks include a plurality of blocks including a plurality of first recognition results for each speaker generated in an EPD (End Point Detection) unit for a voice dialogue, the plurality of blocks being generated by different speakers. The plurality of blocks include a plurality of second recognition results for each speaker generated at a preset time interval. The plurality of blocks include a plurality of blocks including a plurality of first recognition results for each speaker generated in an EPD (End Point Detection) unit for a voice dialogue, the plurality of first recognition results for each speaker generated in an EPD (End Point Detection) unit for a voice dialogue, the plurality of second recognition results for each speaker generated in a preset time interval.

9. A computer program stored on a computer-readable recording medium, The computer program, when executed by a processor, and dividing the plurality of blocks into a plurality of blocks using inter-token boundaries according to a predetermined division criterion for each of the plurality of speaker-specific speech recognition data. The plurality of blocks include a plurality of blocks including a plurality of first recognition results for each speaker generated in an EPD (End Point Detection) unit for a voice dialogue, the plurality of blocks being different speakers from each other. The plurality of blocks include a plurality of second recognition results for each speaker generated at a preset time interval. The plurality of blocks include a plurality of blocks including a plurality of first recognition results for each speaker generated in an EPD (End Point Detection) unit for a voice dialogue, the plurality of first recognition results for each speaker generated in a preset time interval.

Citation Information

Patent Citations

  • Intelligent conference support system

    JP2000112931A

  • Discourse summary generation system and discourse summary generation program

    JP2012003701A

  • Compliance check system and compliance check program

    JP2016085697A

  • Convention support device, convention support method, and convention support program

    JP2017161850A

  • Input information support device, input information support method, and input information support program

    JP2017182822A