Conversation support device, conversation support system, conversation support method, and program
The conversation support system addresses real-time content display challenges by integrating real-time sub-section and delayed utterance section text recognition, enhancing user understanding of conversation progress and reliability.
Patent Information
- Application Number
- JP2022018207
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-08
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2042-02-08
AI Technical Summary
Existing conversation support systems struggle to display conversation content in real time for individuals with hearing impairments while ensuring speech recognition accuracy, leading to difficulty in following the conversation progress.
A conversation support system that performs real-time sub-section speech recognition to display partial text information and integrates it with delayed, more accurate utterance section text information, highlighting differences in display mode to ensure reliability.
Enables easier tracking of conversation progress with reliable content display by showing sub-section text in real time and integrating it with more accurate utterance section text, allowing users to notice content changes.
Smart Images

Figure 0007748888000001 
Figure 0007748888000002 
Figure 0007748888000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a conversation support device, a conversation support system, a conversation support method, and a program. [Background technology]
[0002] Conversation support systems have been proposed to support conversations between people with hearing disabilities, such as in meetings. These systems recognize speech uttered during a conversation, convert it into text, and display the text on a screen.
[0003] For example, the conference system described in Patent Document 1 includes a slave unit equipped with a sound pickup unit, a text input unit, and a display unit, and a master unit connected to the slave unit, which creates minutes of a meeting using text information obtained by voice recognition of the voice input from the slave unit or the text information input from the slave unit, and shares the created minutes with the slave unit. In this conference system, when a participant joins a conversation by text, the master unit controls the master unit to wait for other participants to speak, and transmits information to the slave unit to wait for speech. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 2019-179480 Summary of the Invention [Problem to be solved by the invention]
[0005] People with hearing impairments understand the content of a conversation by reading the text displayed on the display. To help them understand the progress of the conversation, it is expected that the text representing the conversation content be displayed sequentially in real time. On the other hand, to ensure speech recognition accuracy, it is preferable to perform speech recognition on the entire content of one utterance all at once, rather than sequentially. In this case, the text to be displayed cannot be determined until the end of one utterance, so the text representing the conversation content cannot be displayed in real time. This has sometimes made it difficult for people with hearing impairments to follow the conversation.
[0006] An object of the present invention is to provide a conversation support device, a conversation support system, a conversation support method, and a program that enable a user to more easily grasp the progress of a conversation while ensuring the reliability of the conversation content. [Means for solving the problem]
[0007] (1) The present invention has been made to solve the above-mentioned problems, and one aspect of the present invention is a system including: a first speech recognition unit that performs speech recognition processing based on a speech signal and determines sub-section text information for each sub-section that is a part of an utterance section; a second speech recognition unit that performs speech recognition processing based on the speech signal and determines utterance section text information for each of the utterance sections; an information integration unit that integrates the sub-section text information with the utterance section text information to generate integrated text information; and an output processing unit that outputs the sub-section text information to a display unit and then outputs the integrated text information to the display unit. The information integration unit integrates a first graph showing candidates for subsection text information for each subsection, obtained by the first speech recognition unit, and their scores and permutations, into a second graph showing candidates for subsection text information for each subsection constituting a part of the utterance section, obtained by the second speech recognition unit, and their scores and permutations, to generate an integrated graph; calculates, using the integrated graph, an utterance section score, which is the score of a candidate for utterance section text information obtained by arranging the candidates, from the scores of the candidates for subsection text information for each subsection; and determines the integrated text information based on the utterance section score. It is a conversation support device.
[0008] (2) Another aspect of the present invention is the conversation support device of (1), wherein the output processing unit may determine a display mode for a differential section, which is a section in the integrated text information where a difference occurs between the integrated text information and the partial section text information, to be different from a display mode for other sections of the integrated text information.
[0010] (3) Another aspect of the present invention is (1)In the conversation support device, the score of the candidate subsection text information may include an acoustic cost and a linguistic cost, and the information integration unit may calculate the speech section score as a weighted average of the sum of the acoustic costs and the sum of the linguistic costs of the candidate subsection text information for each subsection in the speech section.
[0011] (4) Other aspects of the present invention include (1) to (3). (3) In any one of the conversation support devices, the sub-section may be a section corresponding to one or more words.
[0012] (5) Another aspect of the present invention is a method for controlling a computer a first speech recognition step of performing speech recognition processing based on a speech signal to determine subsection text information for each subsection that is a part of an utterance section; a second speech recognition step of performing speech recognition processing based on the speech signal to determine utterance section text information for each of the utterance sections; an information integration step of integrating the subsection text information with the utterance section text information to generate integrated text information; and an output processing step of outputting the subsection text information to a display unit, and then outputting the integrated text information to the display unit, wherein the information integration step integrates a first graph showing candidates for subsection text information for each of the subsections obtained in the first speech recognition step, their scores and permutations, into a second graph showing candidates for subsection text information for each of the subsections that constitute a part of the utterance section, obtained in the second speech recognition step, their scores and permutations, to generate an integrated graph, calculates an utterance section score, which is the score of a candidate for utterance section text information obtained by arranging the candidates, from the scores of the candidates for the subsection text information for each of the subsections, using the integrated graph, and determines the integrated text information based on the utterance section score It may also be a program.
[0013] (6) Other aspects of the present invention include (1) to (3). (4) A conversation support system comprising: any one of the conversation support devices described above; and the display unit.
[0014] (7) In another aspect of the present invention, a conversation support device performs a first speech recognition step of performing speech recognition processing based on a speech signal to determine subsection text information for each subsection that is a part of an utterance section, a second speech recognition step of performing speech recognition processing based on the speech signal to determine utterance section text information for each of the utterance sections, an information integration step of integrating the subsection text information with the utterance section text information to generate integrated text information, and an output processing step of outputting the subsection text information to a display unit and then outputting the integrated text information to the display unit. A conversation support method, wherein the information integration step integrates a first graph showing candidates for subsection text information for each subsection, obtained in the first speech recognition step, and their scores and permutations, into a second graph showing candidates for subsection text information for each subsection constituting a part of the utterance section, obtained in the second speech recognition step, and their scores and permutations, to generate an integrated graph; calculates, using the integrated graph, speech section scores, which are scores of candidates for speech section text information obtained by arranging the candidates, from the scores of the candidates for subsection text information for each subsection; and determines the integrated text information based on the speech section scores. This is a conversation support method. [Effects of the Invention]
[0015] According to the present invention, it is possible to more easily grasp the progress of a conversation while ensuring the reliability of the conversation content. (1) of the present invention, (5)、(6) or (7)According to this aspect, sub-section text information indicating the content of an utterance for each sub-section is sequentially displayed on the display unit, and integrated text information integrated with the speech section text information indicating the content of an utterance for each sub-section is displayed. By displaying the sub-section text information for each sub-section in real time, the user can grasp the progress of the conversation, and thereafter, by displaying the integrated text information for each speech section, the reliability of the content of the conversation can be ensured. Furthermore, in addition to the candidates for speech section text information obtained by the second speech recognition unit, the candidates for partial section text information obtained by the first speech recognition unit can be referenced to improve the reliability of the conversation content.
[0016] According to the second aspect, the difference section is displayed in a different manner from the other sections. This allows the user to easily notice the difference section where a difference has occurred from the partial section text information, thereby avoiding overlooking highly reliable conversation content in the difference section.
[0018] (3) According to this aspect, the speech segment score is obtained by weighting the reliability based on the acoustic features and the reliability based on the linguistic features, thereby making it possible to adjust the contribution of the acoustic features and the linguistic features to the reliability of the conversation content.
[0019] (4) According to this aspect, partial section text information indicating the content of an utterance for each word is sequentially displayed on the display unit 30, and integrated text information integrated with the utterance section text information indicating the content of an utterance for each utterance section is displayed. Displaying partial section text information for each word in real time allows the user to grasp the progress of the conversation, and then displaying integrated text information for each utterance section ensures the reliability of the content of the conversation. [Brief explanation of the drawings]
[0020] [Figure 1] 1 is a schematic block diagram illustrating an example of the configuration of a conversation support system according to an embodiment of the present invention. [Figure 2] FIG. 10 is an explanatory diagram showing a first example of a case where real-time processing is possible. [Figure 3] FIG. 10 is an explanatory diagram showing a second example in which real-time processing is possible. [Figure 4] FIG. 10 is an explanatory diagram showing an example in which real-time processing is not possible. [Figure 5] FIG. 10 is a diagram showing an example of output of subsection text information according to the embodiment. [Figure 6] FIG. 10 is a diagram illustrating an example of a hypothesis lattice. [Figure 7] FIG. 10 is an explanatory diagram illustrating an example of graph integration. [Figure 8] FIG. 4 is an explanatory diagram illustrating an example of timing of the first speech recognition process. [Figure 9] FIG. 10 is an explanatory diagram illustrating an example of timing of the second speech recognition process. [Figure 10] FIG. 10 is an explanatory diagram illustrating an example of output timing of integrated text information. [Figure 11] 10 is a flowchart illustrating an example of a conversation assistance process. DETAILED DESCRIPTION OF THE INVENTION
[0021] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. First, an example configuration of a conversation support system S1 according to this embodiment will be described. FIG. 1 is a schematic block diagram showing an example configuration of the conversation support system S1 according to this embodiment. The conversation support system S1 includes a conversation support device 10, a sound collection unit 20, a display unit 30, and a terminal device 40.
[0022] The conversation support system S1 is used in a conversation involving two or more participants. The participants may include one or more people who have difficulty speaking and / or listening to audio (hereinafter referred to as "people with disabilities"). The people with disabilities may individually operate the terminal device 40 to input text indicating what is being said (hereinafter referred to as "operation text") to the conversation support device 10. People who do not have difficulty speaking and listening to audio (hereinafter referred to as "able-bodied people") may individually use the sound pickup unit 20 or a device equipped with a sound pickup unit (e.g., the terminal device 40) to input their spoken voice to the conversation support device 10. The conversation support device 10 performs known speech recognition processing on the audio data indicating the input voice, and converts it into text indicating what is being said (hereinafter referred to as "speech text"). Each time the conversation support device 10 acquires either the operation text obtained by conversion or the utterance text obtained from the terminal device 40 (hereinafter referred to as "utterance text" to be distinguished from "utterance text"), it displays the acquired utterance text on the display unit 30. A person with a disability can understand the content of the utterance in a conversation by reading the displayed text (hereinafter referred to as "display text").
[0023] The conversation support device 10 determines the start (hereinafter referred to as "start of speech") and the end (hereinafter referred to as "end of speech") of a section in which any participant speaks at once (hereinafter referred to as "speech section") from an audio signal indicating the collected voice. The speech section is identified by this determination. The conversation support device 10 is capable of executing a first speech recognition process and a second speech recognition process in parallel as speech recognition processes for the speech section.
[0024] In the first speech recognition process, the conversation support device 10 determines subsection text information indicating the content of an utterance for each part of an utterance section (hereinafter sometimes referred to as a "subsection") and sequentially outputs the determined subsection text information to the display unit 30. The display unit 30 is capable of displaying the subsection text information indicating the content of an utterance for each subsection as utterance text in real time. That is, the processing result of the first speech recognition process is sequentially reflected in the display text online, and therefore the first speech recognition process is a foreground process. In the following description, this feature of the first speech recognition process may be referred to as online.
[0025] In the second speech recognition process, the conversation support device 10 determines speech section text information indicating the content of the utterance for each speech section in parallel with the first speech recognition process. The conversation support device 10 integrates the determined speech section text information with sub-section text information in the speech section to generate integrated text information. The conversation support device 10 outputs the generated integrated text information to the display unit 30. The display unit 30 displays display text indicating the content of the utterance represented by the integrated text information. A method for integrating the speech section text information and the sub-section text information will be described later.
[0026] Because the end point of an utterance interval is defined when the end of the utterance is detected, output of the utterance interval text information is delayed until the utterance interval is defined at the earliest. The integrated utterance interval text information is displayed after the partial interval text information. Because the processing result of the second speech recognition process is not immediately reflected in the displayed text, the second speech recognition process is background processing. In the following description, this characteristic of the second speech recognition process may be referred to as offline. According to the second speech recognition process, the utterance interval text information is determined taking into account the occurrence probability of utterance content across multiple partial intervals that make up the utterance interval (including the transition probability of utterance content between adjacent partial spaces). Therefore, intervals (hereinafter referred to as "difference intervals") in which the integrated text information differs from the partial interval text information may occur. Furthermore, the utterance interval text information tends to be more reliable than the partial interval text information. Therefore, the conversation support device 10 may display the display text related to the integrated text information in the difference interval in a different manner (e.g., one or a combination of items such as background color, character color, character type, line width, line type, decoration, etc.) from the display text in other intervals. This allows participants who view the displayed text based on the integrated text information to notice changes in the content of the speech from the displayed text based on the provisionally displayed subspace text information, and allows them to access more reliable information.
[0027] The sound collection unit 20 collects sound arriving at the unit itself and outputs sound data indicating the collected sound to the conversation support device 10. The sound collection unit 20 includes a microphone. The number of sound collection units 20 is not limited to one, and may be two or more. The sound collection unit 20 may be, for example, a portable wireless microphone. A wireless microphone mainly collects the speech of each individual holder. The sound collection unit 20 may be a microphone array formed by arranging multiple microphones in different positions. The microphone array as a whole outputs sound data of multiple channels to the conversation support device 10. In the following description, a case where the sound collection unit 20 is a wireless microphone having one microphone will be mainly taken as an example.
[0028] Display unit 30 displays display information, for example, various display screens, based on display data input from conversation support device 10. The display data includes display screen data, which will be described later. Display unit 30 may be any type of display, such as a liquid crystal display (LCD) or an organic electroluminescence display (OLED). Note that the display area of the display constituting display unit 30 may be configured as a single touch panel in which the detection areas of touch sensors are superimposed and integrated.
[0029] The conversation support system S1 may also include an operation unit (not shown). The operation unit receives operations from the user and outputs an operation signal corresponding to the received operation to the conversation support device 10. The operation unit may include a general-purpose input device such as a touch sensor (which may be integrated with the display unit 30), a mouse, or a keyboard, or may include dedicated components such as buttons, knobs, or dials.
[0030] Terminal device 40 includes an operation unit, a display unit, some or all of a sound collection unit, and an input / output unit. In the following description, the operation unit, display unit, sound collection unit, and input / output unit included in terminal device 40 will be referred to as a terminal operation unit, a terminal display unit, a terminal sound collection unit, and a terminal input / output unit, respectively, to distinguish them from the operation unit, display unit, sound collection unit, and input / output unit included in conversation support device 10.
[0031] The terminal input / output unit inputs and outputs various data to and from conversation support device 10. The terminal input / output unit includes, for example, an input / output interface that inputs and outputs data using a predetermined input / output method or communication method. The terminal operation unit receives an operation from the user and outputs an operation signal corresponding to the received operation to conversation support device 10 via the input / output unit. The terminal operation unit includes an input device.
[0032] The terminal display unit displays a display screen based on display screen data input from conversation support device 10 via the input / output unit. The terminal display unit may be integrated with the terminal operation unit and configured as a touch panel. While the display screen is displayed, the terminal operation unit transmits text information indicating text composed of characters specified in response to an operation to conversation support device 10 using the terminal input / output unit (text input).
[0033] The terminal sound collection unit collects sound arriving at the terminal and outputs sound data representing the collected sound to conversation support device 10 using the terminal input / output unit. The terminal sound collection unit includes a microphone. The sound data acquired by the terminal sound collection unit may be subjected to voice recognition processing in conversation support device 10.
[0034] 1 includes one conversation support device 10 and one terminal device 40, but is not limited to this. The number of terminal devices 40 may be two or more, or may be zero. In the example shown in FIG. 1, the conversation support device 10 and the terminal device 40 function as a parent device and a child device, respectively.
[0035] In this application, "conversation" refers to communication between two or more participants, and is not limited to communication using voice, but also includes communication using other types of information media, such as text. Conversation is not limited to spontaneous or voluntary communication between two or more participants, but also includes communication in which a specific participant (e.g., a moderator) controls what other participants say, such as in a conference, presentation, lecture, or ceremony. Furthermore, "utterance" refers to communicating one's intentions using language, and is not limited to communicating one's intentions by uttering voice, but also includes communicating one's intentions using other types of information media, such as text.
[0036] (Conversation support device) Next, an example of the configuration of conversation support device 10 according to this embodiment will be described. Conversation support device 10 includes an input / output unit 110, a control unit 120, and a storage unit 140. The input / output unit 110 can wirelessly or wiredly input and output various data to and from other members or devices using a predetermined input / output method or communication method. The input / output unit 110 can use, for example, a Universal Serial Bus (USB), an input / output method specified in IEEE1394, IEEE802.11, LTE-A (Long Term Evolution Advanced), 5G-NR (5 th Any communication method may be used, such as a communication method defined in the International Standard for Standardization (ISG) for Generation-New Radio (ISNR). The input / output unit 110 is configured to include, for example, one or both of an input / output interface and a communication interface.
[0037] The control unit 120 performs various arithmetic processes to realize and control the functions of the conversation assistance device 10. The control unit 120 may be realized by a dedicated component, or may be realized as a computer including a processor and storage media such as a read-only memory (ROM) and a random access memory (RAM). The processor reads a predetermined program stored in advance in the ROM, expands the read program in the RAM, and uses the storage area of the RAM as a working area. The processor realizes the functions of the control unit 120 by executing processes instructed by various instructions written in the read program. The realized functions may include the functions of each unit described below. In the following description, executing processes instructed by instructions written in a program may be referred to as "executing a program" or "executing a program." The processor is, for example, a central processing unit (CPU).
[0038] The control unit 120 includes an acoustic processing unit 122, a feature calculation unit 124, a first speech recognition unit 126, a second speech recognition unit 128, an information integration unit 130, an output processing unit 132, and an utterance information recording unit 134. The acoustic processing unit 122 receives audio data from the sound collection unit 20 via the input / output unit 110. The acoustic processing unit 122 performs predetermined preprocessing on the received audio data. The preprocessing may include, for example, known noise suppression processing. When audio data from multiple channels is received at once, the preprocessing may include sound source separation processing. The acoustic processing unit 122 may identify speakers by performing known speaker recognition processing on source-specific audio data representing audio separated by the sound source separation processing, and may add speaker identification information indicating the identified speakers to the source-specific audio data. When audio data is received from multiple sound collection units 20, the acoustic processing unit 122 may add a Mic ID to the audio data received from each sound collection unit 20 as identification information identifying the individual sound collection unit 20. The Mic ID may be used as speaker identification information for identifying the speaker who is the exclusive speaker of that sound collection unit 20.
[0039] The acoustic processing unit 122 detects a speech interval from the audio indicated in the preprocessed audio data (which may also include audio data by sound source) obtained by performing preprocessing (speech interval detection). A speech interval refers to a period in which one of the speakers is speaking. A speech interval corresponds to a period in which the audio data contains significant components of the speech audio. A speech interval corresponds to a period that begins when the start of speech is detected and ends when the next time it is determined that speech has ended.
[0040] In voice activity detection, the acoustic processing unit 122 performs a known voice activity detection (VAD) process on the preprocessed voice data to determine whether the frame currently being processed (hereinafter referred to as the "current frame") is a voice activity. For example, the acoustic processing unit 122 calculates the power and the number of zero crossings as features indicating the speech state for each frame of a predetermined length (e.g., 10 to 50 ms) of the acquired voice data. For example, the acoustic processing unit 122 determines, as a voice activity, a frame whose calculated power is greater than the lower limit of the power for the predetermined speech state and whose zero crossing number is within the range for the predetermined speech state (e.g., 300 to 1000 times per second), and determines other frames as non-voice activity.
[0041] When the speech state has been determined as a non-speech interval (hereinafter referred to as a "continuous non-speech interval") for a predetermined number of consecutive frames up to the frame immediately preceding the current frame (hereinafter referred to as a "previous frame"), but the acoustic processing unit 122 determines the speech state of the current frame as a new speech interval, the acoustic processing unit 122 determines the speech state of the current frame as a speech start. In the following description, a frame in which the speech state has been determined as a non-speech interval, which is a non-speech interval, for a predetermined number of consecutive frames up to the current frame, the acoustic processing unit 122 determines the speech state of the frame immediately preceding the continuous non-speech interval as a speech end. In the following description, a frame in which the speech state has been determined as a speech end is referred to as a "speech end frame." The acoustic processing unit 122 can identify the period from the speech start frame to the next speech end frame as a speech interval. The acoustic processing unit 122 sequentially outputs the pre-processed speech data from the speech start frame to the next speech end frame to the feature calculation unit 124 and the speech information recording unit 134 .
[0042] The feature calculation unit 124 calculates acoustic features for each frame of the speech data input from the acoustic processing unit 122. The acoustic features are parameters that indicate the acoustic features of the speech. The feature calculation unit 124 calculates, for example, multidimensional Mel Frequency Cepstrum Coefficients (MFCC). The feature calculation unit 124 outputs the calculated acoustic features to the first speech recognition unit 126 and the second speech recognition unit 128. If speaker identification information is added to the speech data input for each utterance section, the feature calculation unit 124 may add the speaker identification information to the acoustic features in association with the speaker identification information, and output the acoustic features to the output processing unit 132 via the first speech recognition unit 126 and the second speech recognition unit 128.
[0043] The first speech recognition unit 126 performs a first speech recognition process in real time on the acoustic features input from the feature calculation unit 124. As the first speech recognition process, the first speech recognition unit 126 determines subsection text information using a trained first speech recognition model as text information indicating the content of an utterance for each subsection that is part of an utterance section. The first speech recognition unit 126 outputs the determined subsection text information to the information integration unit 130 and the output processing unit 132. The first speech recognition process is an online process for each subsection. However, in order to represent the progress of a conversation in conversation assistance, one subsection is set to have a duration equal to or longer than the time required to pronounce at least a notation unit (e.g., a letter, a number, a symbol, etc.). For example, a period corresponding to one word, a phrase, etc. may be applied as the subsection.
[0044] The first speech recognition unit 126 applies words as subsections. In this case, the first speech recognition unit 126 uses an acoustic model, a context-dependent model, and a word dictionary as mathematical models for the first speech recognition process. The acoustic model is used to estimate context-independent phonemes from a time series including one or more sets of acoustic features. The context-dependent model is used to estimate context-dependent phonemes from context-independent phonemes. The word dictionary is used to estimate words from a phoneme string including one or more context-dependent phonemes. The word dictionary may include word text information indicating the natural language notation of each word.
[0045] The second speech recognition unit 128 performs second speech recognition processing for each speech section on the acoustic features input from the feature calculation unit 124. In other words, the second speech recognition processing is batch processing for each speech section. In the second speech recognition processing, the second speech recognition unit 128 determines speech section text information using a trained second speech recognition model as text information indicating the speech content for each speech section. As mathematical models related to the second speech recognition processing, an acoustic model, a context-dependent model, and a word dictionary, as well as a grammar model indicating the relationship (grammatical rules) between one or more words, are used. The second speech recognition unit 128 outputs the determined speech section text information to the information integration unit 130 and the speech information recording unit 134.
[0046] The information integration unit 130 integrates the speech section text information for each speech section input from the second speech recognition unit with the subsection text information for each subsection input from the first speech recognition unit 126 to generate integrated text information. The information integration unit 130, for example, determines candidates (hypotheses) of speech section text information formed by arranging candidates of subsection text information for each subsection constituting the speech section using the second speech recognition model in that order. The information integration unit 130 calculates a score for each candidate of speech section text information (hereinafter referred to as a "speech section score"). The information integration unit 130 can calculate the sum of scores (hereinafter referred to as a "subsection score") related to each candidate of subsection text information constituting the candidate of speech section text information as the speech section score. The subsection score is a real value indicating the confidence of the candidate of the subsection text information. The confidence indicates the degree of likelihood as a hypothesis. For example, a transition probability can be used as the speech section score. The information integration unit 130 can determine, as the integrated text information, the speech section text information that provides the speech section score indicating the highest reliability. The information integration unit 130 may obtain, from the first speech recognition unit 126, candidates for subsection text information derived as intermediate data when determining the subsection text information for each subsection, and from the second speech recognition unit 128, candidates for subsection text information derived as intermediate data when determining the speech section text information for each subsection. The information integration unit 130 outputs the generated integrated text information to the output processing unit 132.
[0047] The output processing unit 132 generates display screen data that sequentially represents the sub-interval text information for each sub-interval input from the first speech recognition unit 126, and outputs the generated display screen data to the display unit 30 via the input / output unit 110. On the other hand, the output processing unit 132 receives integrated text information for each utterance interval from the information integration unit 130. The input of the integrated text information is delayed compared to the input of the sub-interval text information. The output processing unit 132 updates the display screen data by replacing the sub-interval text information for that utterance interval with the integrated text information. The output processing unit 132 outputs the updated display screen data to the display unit 30. Here, the output processing unit 132 may detect a difference interval in which a difference occurs between the integrated text information for that utterance interval and the sub-interval text information. The output processing unit 132 may temporarily (for example, within a predetermined time period (for example, 2 to 10 seconds) starting from the detection of the difference interval) or permanently set the display mode for the difference interval to a display mode different from that for other intervals. The output processing unit 132 may receive speaker identification information associated with the sub-interval text information and the integrated text information derived from acoustic features for each utterance interval via the first speech recognition unit 126 and the information integration unit 130. The output processing unit 132 may generate display screen data for the speech section including speaker identification information. The speaker identification information may be placed, for example, at the beginning of the subsection text information or the integrated text information, and may be represented by an icon, figure, symbol, or the like for identifying the speaker.
[0048] The speech information recording unit 134 receives input of speech data for each speech section from the acoustic processing unit 122, from the speech start frame to the speech end frame. On the other hand, the speech information recording unit 134 receives input of integrated text information for each speech section from the information integration unit 130. The speech information recording unit 134 associates the input speech data with the integrated text information and records them in the storage unit 140. The storage unit 140 stores integrated text information indicating the speech content for each speech section and minutes data indicating the speech data. Speaker identification information for identifying the speaker may be added to the speech data of the speech section.
[0049] The storage unit 140 temporarily or permanently stores various types of data. The storage unit 140 stores a program describing the processing to be executed by the control unit 120, various types of data used in the processing (including various parameters, initial values, intermediate values, voice recognition models, etc.), and various types of data acquired by the control unit 120. The storage unit 140 is configured to include, for example, the above-mentioned storage media such as ROM and RAM.
[0050] (Real-time processing capability) As described above, the first speech recognition unit 126 determines sub-segment text information for each sub-segment in real time through the first speech recognition process and outputs it to the output processing unit 132. To enable real-time processing, the first speech recognition unit 126 requires that there be no processing step in which the elapsed time from input of input information to output of an output result exceeds the acquisition period for acquiring new input information. Figure 2 illustrates an example of a speech processing procedure that allows real-time processing. In this example, for one frame of speech input, the processing time from the first step and the second step to obtaining an output result is less than one frame.
[0051] Even if the voice input to be processed at one time spans a period of multiple frames, real-time processing is possible if the processing time is one frame or less when the period of the new voice input is one frame. In the example of Figure 3, two frames of voice input are processed at one time, but the voice input for one frame of the two frames is newly acquired, and the voice input for the remaining one frame is processed in the previous processing. Even in such a case, real-time processing is possible because there is no increase in the delay time from the input of unprocessed voice input to the time when processing can begin.
[0052] In contrast, real-time processing is not possible in the example of FIG. 4. In this example, two frames of audio input are processed at a time, and one frame of audio input is newly acquired. However, the processing times for the first and second steps for one frame of audio input are 0.2 and 1.3 frames, respectively. The time when the second step can start for the two frames of audio input up to the second frame is the end of the first step. This time is 0.2 frames after the audio input of the second frame. The time when the second step can start for the two frames of audio input up to the third frame is the end of the immediately preceding second step. This time is 0.5 frames after the audio input of the third frame. The time when the second step can start for the two frames of audio input up to the fourth frame is the end of the immediately preceding second step. This time is 0.8 frames after the audio input of the fourth frame. As such, the delay time until new audio input can be processed increases.
[0053] FIG. 5 shows an example of output of subsection text information according to this embodiment at each time point. The subsection text information, which is the processing result of the first speech recognition process, accumulates roughly over time. In this example, the first speech recognition unit 126 defines subsection text information as a subsection, a period corresponding to one kanji or kana character used in Japanese writing, and repeats the process of outputting the defined subsection text information. In the example of FIG. 5, Japanese text, which is the recognition result, is added one character at a time. "Eh" is displayed as the recognition result at the beginning of the utterance. At the end of the speech section, when the auxiliary verb "desu," which frequently appears at the end of a Japanese declarative sentence, is recognized, it is estimated as the end of the sentence by referring to a word dictionary or grammar dictionary. A period "." indicating the end of the sentence is added, and "desu" is written as the speech content at the end of the speech section.
[0054] If the first speech recognition process is capable of real-time processing, the first speech recognition unit 126 may add newly acquired context-independent phoneme candidates over time to one or more already estimated context-independent pixels to estimate other word candidates with higher reliability. If the processing time from acquiring the speech signal for a new subsection to outputting the utterance text is shorter than the average length of the subsection, real-time display becomes possible. Furthermore, punctuation marks may be added or deleted as the estimated word changes. In the example of Figure 5, the recognition result "eh" in the first line is updated to "rake" in the second line, the recognition result "rake" in the second line is updated to "eh, tree" in the third line, "tree" at the end of the third line is updated to "today" in the fourth line, "spring" at the end of the fifth line is updated to "sunny" in the sixth line, "later" at the end of the ninth line is updated to "rain" and "sama" in the twelfth line is updated to "plan" in the thirteenth line.
[0055] (Hypothetical lattice data) The second speech recognition unit 128 executes the second speech recognition process to determine speech section text information for each speech section and outputs it to the output processing unit 132. As described above, the second speech recognition process includes a process of estimating sub-section text information candidates for each sub-section constituting the speech section, as well as a process of generating speech section text information candidates by concatenating sub-section text information candidates in the order of the sub-sections in the speech section. For each speech section information candidate, the second speech recognition unit 128 calculates, as the speech section score, the sum of the sub-section scores corresponding to the sub-section text information candidates for each sub-section constituting the speech section information candidate. The second speech recognition unit 128 can determine, as the speech section text information, the speech section text information candidate that gives the highest speech section score as the recognition result.
[0056] In the second speech recognition processing, the second speech recognition unit 128 generates hypothesis lattice data indicating a hypothesis lattice using the above mathematical model according to a known method. The hypothesis lattice indicates one or more speech section text candidates arranged in order as hypotheses for each subsection in the speech section. Each subsection text information candidate is associated with its order in the speech section and a subsection score. As illustrated in FIG. 6, the hypothesis lattice is represented as a directed graph having multiple nodes and one or more edges (edges, branches, links) connecting each two nodes. Two of the multiple nodes are associated with a start symbol and an end symbol. The start symbol and end symbol indicate the start and end of the speech, respectively. Each edge is associated with a subsection text information candidate and a subsection score indicating its reliability. Therefore, the candidates for sub-interval text information corresponding to each edge forming each path from the start symbol to the end symbol are arranged in that order to represent the utterance interval text candidates.
[0057] 6, the subintervals are words, and the hypothesis lattice as a whole has the form of a word graph. When the subinterval of interest to be processed appears at the start of an utterance, the second speech recognition unit 128 may apply a start symbol because there is no immediately preceding subinterval. When the subinterval of interest appears at the end of an utterance, the second speech recognition unit 128 may apply an end symbol because there is no immediately succeeding subinterval.
[0058] In the hypothesis lattice, when there are multiple candidate words following an edge, the edge branches into multiple edges at the node at the end of the edge. Each of the multiple candidate words is associated with an individual branched edge. In the example of Figure 7, the edge corresponding to "Ito" branches into two subsequent edges at the node, and each edge is associated with "to" and "mo." When multiple edges share the same candidate word, the edges are merged at the tip of the edge corresponding to the corresponding word. In the example of Figure 7, each of the two edges is associated with the common word "reunion," and the edges following the two edges are merged into one edge via a node and associated with "suru" as the common candidate word.
[0059] The second speech recognition unit 128 can refer to the generated hypothesis lattice data and calculate the confidence score as the sum of the subsection scores assigned to each node for each path from the start symbol to the end symbol. Each subsection score may be a scalar value or a vector value. Each score may be a real value indicating higher confidence as the score increases, or a real value (cost value) indicating higher confidence as the score decreases. The subsection score may be expressed, for example, as a two-dimensional vector including an acoustic cost and a linguistic score (graph cost) as elements. The acoustic cost is an index value indicating the likelihood that a sequence of acoustic features in the subsection is the sequence of acoustic features of the words in the subsection. The acoustic cost is derived from the acoustic features in the subsection using an acoustic model. The linguistic cost is an index value indicating the likelihood of occurrence in the subsection based on linguistic characteristics. The linguistic cost is derived from the acoustic features, context-independent phonemes, and words in the subsection using a context-dependent model, a word dictionary, and a grammar model, respectively. The subsection cost and its elements, the acoustic cost and the linguistic cost, may be expressed as scaled real values so that the speech section score can be calculated efficiently and the recognition accuracy does not decrease.
[0060] When the subsection score includes an acoustic cost and a linguistic cost, the second speech recognition unit 128 can calculate, as the speech section score, a weighted average of the acoustic score, which is the sum of the acoustic costs, and the linguistic score, which is the sum of the linguistic costs. The second speech recognition unit 128 can select the path that minimizes the calculated speech section score, and generate speech section text information by concatenating words corresponding to edges that make up the selected path in that order. In the example of FIG. 6, of three paths starting from the start symbol and ending at the end symbol, the path shown at the top is selected. The words corresponding to each edge that make up the selected path are "family," "with," "reunion," and "to" in that order, and the speech content "to be reunited with family" is estimated.
[0061] The second speech recognition process estimates the content of an utterance by quantitatively evaluating the relationship between candidate subsections (e.g., words) for each utterance section. The length of one utterance section is typically several seconds to several tens of seconds. Because real-time processing for each utterance section is not practical, the process is performed offline. The speech content estimated by the second speech recognition process tends to be more accurately estimated than the speech content estimated by the first speech recognition process, which estimates each subsection, but this is not necessarily the case.
[0062] Therefore, the information integration unit 130 acquires subsection text information candidates, context-dependent phoneme candidates, context-independent phoneme candidates, and acoustic features for each subsection obtained by the first speech recognition processing for that speech section from the first speech recognition unit 126. The information integration unit 130 executes the same procedure as in the second speech recognition processing to arrange the acquired subsection text information candidates in the order of the subsections, and generates data indicating a hypothesis lattice indicating the speech section text candidates as first hypothesis lattice data. When generating the first hypothesis lattice data, the information integration unit 130 uses an acoustic model, a context-dependent model, and a word dictionary, and does not necessarily use a grammar model.
[0063] The information integration unit 130 acquires hypothesis lattice data (hereinafter referred to as "second hypothesis lattice data") generated in the second speech recognition process from the second speech recognition unit 128. The information integration unit 130 combines, for each speech section, a first hypothesis lattice (hereinafter referred to as "first graph") represented by the first hypothesis lattice data and a second hypothesis lattice (hereinafter referred to as "second graph") represented by the second hypothesis lattice data, and defines the obtained graph as a combined graph (graph integration).
[0064] In graph integration, the information integration unit 130 incorporates a unique edge across the first graph and the second graph, as well as the subinterval text information and subinterval score corresponding to that edge, into the elements of the combined graph. If there is an overlapping edge between the first graph and the second graph, the information integration unit 130 integrates the edges into a single edge and determines a combined value obtained by combining the subinterval scores of the edges (for example, if the individual subinterval scores are transition probabilities, the sum of those values) as a new subinterval score. The information integration unit 130 incorporates the integrated edge, the subinterval text information and new subinterval score corresponding to that edge, into the elements of the combined graph. An edge that overlaps with an edge of interest to be processed means that the candidate subinterval text information corresponding to the edge of interest is common, and there are no edges immediately before or after the edge of interest that correspond to the candidate subinterval text information. However, if the edge of interest is the edge at the beginning of an utterance interval, the immediately preceding edge is not referenced, and if the edge of interest is the edge at the end of an utterance interval, the immediately succeeding edge is not referenced. One end of the edge at the beginning of an utterance interval is associated with the start symbol, and one end of the edge at the end of an utterance interval is associated with the end symbol, thereby distinguishing them from other types of edges. Therefore, in the combined graph, the unique paths in the first and second graphs before combining are parallel, and the common paths are consolidated into one.
[0065] Figure 7 shows an example of the first graph and the second graph in the upper left and lower left, respectively, and an example of the combined graph on the right. For example, at the beginning of the first graph, there are edges corresponding to "Phantom Thief," "Ito," and "Ito." At the beginning of the second graph, there are edges corresponding to "Kato," "Phantom Thief," and "Dividend." The edges corresponding to "Ito," "Ito," "Kato," and "Dividend" are unique between the first graph and the second graph, so they are maintained. The edge corresponding to "Kato" is common to both the first graph and the second graph, so it is merged into one of them. Then, the sum of the transition probabilities, which are the subinterval scores of those edges in the first graph and the second graph, is assigned to the merged edge of the new subinterval score. The first graph has an edge associated with "recently," but the second graph does not. On the other hand, the second graph has an edge associated with "management," but the first graph does not. Therefore, both the edge associated with "recently" and the edge associated with "management" are adopted. The edge associated with "reunion" and the subsequent edge associated with "do" are merged. This is because such an edge exists in both the first graph and the second graph. The sum of the weight values associated with each edge before the merger is assigned as a new weight value to the merged edge.
[0066] The information integration unit 130 uses the connection graph to calculate an utterance section score for each path based on the sub-segment scores corresponding to the edges that make up each path, and selects (re-evaluates) the path that gives the largest utterance section score. The information integration unit 130 can generate integrated text information indicating the content of the utterance in the utterance section by concatenating, in that order, the words corresponding to the edges that make up the selected path. In the example of Figure 7, of the paths from the start symbol to the end symbol, paths that include edges corresponding to "phantom thief," "with," "reunion," and "do" are selected. Integrated text information indicating the content of the utterance as "reunited with the phantom thief" is generated. In addition, when the information integration unit 130 generates integrated text information using a connection graph, it is sufficient to generate second hypothesis lattice data indicating candidates for speech section text information in the second speech recognition process, and there is no need to determine a single piece of speech section text information that will be the final processing result.
[0067] The generation of hypothesis lattices and speech recognition using hypothesis lattices are described in detail in the following documents. These methods can be applied to this embodiment. Daniel Povey, Mirko Hannermann, et al: “GENERATING EXACT LATTICES IN THE WFST FRAMEWORK”, Proceedings of International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2012, 25-30 March, 2012 “Lattices in Kaldi”, [online], Kaldi Project,<URL: https: / / www.kaldi-asr.org / doc / lattices.html>
[0068] (Processing timing) In the first speech recognition process, the period to be processed at one time is limited to a subsection, thereby enabling online real-time processing. In the example of Fig. 8, the first speech recognition process for a certain subsection is completed by the time the acoustic features for the next subsection are acquired from the feature calculation unit 124. The output processing unit 132 acquires subsection text information indicating the recognition results for each subsection from the first speech recognition unit 126, and can display the display text indicated by the subsection text information on the display unit 30 in real time.
[0069] In the second speech recognition process, the period to be processed at one time is considered to be a speech segment, making online real-time processing unrealistic. The length of one speech segment is typically several seconds to several tens of seconds, and the second speech recognition process evaluates the relevance of multiple subsegments. In the example of FIG. 9, the second speech recognition process for a certain subsegment cannot be completed even when acoustic features for the subsequent subsegment are acquired. Therefore, each time new acoustic features are acquired, the delay time from the acquisition of the new acoustic features to the start of the second speech recognition process increases. In this embodiment, the second speech recognition process is performed offline, and speech segment text information, which is the processing result for each speech segment, is acquired.
[0070] In graph integration, a combined graph is generated by combining the first graph with the second graph. A speech section score is calculated for each path on the generated combined graph, and the path with the largest speech section score is selected. Graph integration is performed on the premise that the information integration unit 130 can obtain candidates for subsection text information for each subsection in the speech section that will become elements of the first graph, and the second graph associated with that speech section. As illustrated in FIG. 10 , graph integration begins after the first speech recognition process and the second speech recognition process are completed. In re-evaluation, a speech section score is calculated for each path using the combined graph obtained by graph integration and used to determine integrated text information. Then, display text based on the integrated text information is displayed on the display unit 30 later than display text based on the subsection text information. In the example of FIG. 10 , immediately after the first speech recognition process for the speech data in the speech section is completed, display text indicating the recognition result for that speech section is displayed, and the second speech recognition process can then be started. In this embodiment, the second speech recognition process may be started after the first speech recognition process has started, before the first speech recognition process has finished, and may be executed in parallel with part or all of the processing period of the first speech recognition process for that speech section. This shortens the processing period from the start of the first speech recognition process to the output of integrated text information, which is the processing result.
[0071] (Conversation support processing) Next, an example of the conversation support process according to this embodiment will be described below. Fig. 11 is a flowchart showing an example of the conversation support process according to this embodiment. (Step S102) The sound processing unit 122 performs preprocessing on the sound data input from the sound collection unit 20. (Step S104) The acoustic processing unit 122 performs a speech detection process on the preprocessed audio data and determines whether or not speech has started based on the detected speech state. If it is determined that speech has started (step S104: YES), the process proceeds to step S106. If it is not determined that speech has started (step S104: NO), the process returns to step S102.
[0072] (Step S106) The feature calculation unit 124 calculates an acoustic feature for each frame of the preprocessed speech data. (Step S108) The first speech recognition unit 126 performs a first speech recognition process on the calculated acoustic signal, and determines subsection text information indicating the content of the speech for each subsection that is a part of the speech section. (Step S110) The output processing unit 132 generates display screen data indicating sub-section text information for each sub-section, and outputs the generated display screen data to the display unit 30. The display unit 30 displays display text indicating the speech content for each sub-section in real time. (Step S112) The acoustic processing unit 122 performs a voice detection process on the preprocessed voice data and determines whether the speech has ended based on the detected speech state. If it is determined that the speech has ended (step S112 YES), the process proceeds to step S114. The period from the start of speech to the end of speech corresponds to the speech period. If it is not determined that the speech has ended (step S112 NO), the process returns to step S102.
[0073] (Step S114) The second speech recognition unit 128 performs a second speech recognition process on the calculated acoustic signal and determines speech section text information indicating the content of the utterance for each speech section. During the second speech recognition process, a second graph indicating a path consisting of a permutation of candidates for subsection text information for each subsection belonging to the utterance section is determined. (Step S116) The information integration unit 130 constructs a first graph indicating a path consisting of a sequence of candidates for subsection text information within the utterance section obtained in the first speech recognition process. The information integration unit 130 combines the second graph and the first graph to generate a combined graph (graph integration). (Step S118) The information integration unit 130 calculates (re-evaluates) the speech section score for each path shown in the integrated graph, and selects a path based on the calculated speech section score. The information integration unit 130 determines, as integrated text information, a permutation of candidates for subsection text information corresponding to each edge constituting the selected path. (Step S120) The output processing unit 132 replaces the sub-section text information in the utterance section with the integrated text information to update the display screen data, and outputs the updated display screen data to the display unit 30. Thus, the content of the utterance in the utterance section is updated to that indicated by the integrated text information.
[0074] As described above, the conversation support device 10 according to this embodiment includes a first speech recognition unit 126 that performs speech recognition processing (e.g., first speech recognition processing) based on a speech signal to determine sub-segment text information for each sub-segment that is a part of an utterance segment, and a second speech recognition unit 128 that performs speech recognition processing (e.g., second speech recognition processing) based on the speech signal to determine speech segment text information for each utterance segment. The conversation support device 10 also includes an information integration unit 130 that integrates the sub-segment text information with the utterance segment text information to generate integrated text information, and an output processing unit 132 that outputs the sub-segment text information (e.g., by including it in display screen data) to the display unit 30, and then outputs the integrated text information (e.g., by including it in the display screen data) to the display unit 30. According to this configuration, partial section text information indicating the speech content of each partial section is sequentially displayed on the display unit 30, and integrated text information integrated with the speech section text information indicating the speech content of each speech section is displayed. Displaying the partial section text information for each partial section in real time allows the user to understand the progress of the conversation, and then displaying the integrated text information for each speech section ensures the reliability of the conversation content.
[0075] Furthermore, the output processing unit 132 may determine the display mode of a difference section, which is a section in the integrated text information where a difference occurs between the integrated text information and the partial section text information, to be different from the display mode of other sections of the integrated text information. With this configuration, the difference section is displayed in a different display mode from the other sections, and the user can easily notice the difference section where a difference has occurred from the partial section text information, thereby avoiding overlooking highly reliable conversation content in the difference section.
[0076] Furthermore, the information integration unit 130 generates an integrated graph by integrating a first graph showing candidates for subsection text information for each subsection, obtained by the first speech recognition unit 126, and their scores and permutations, into a second graph showing candidates for subsection text information for each subsection constituting a part of an utterance section, obtained by the second speech recognition unit 128, and their scores and permutations. The information integration unit 130 may use the integrated graph to calculate an utterance section score, which is the score of the candidate for utterance section text information obtained by arranging the candidates, from the scores of the candidates for subsection text information for each subsection, and determine integrated text information based on the utterance section score. According to this configuration, in addition to the candidates for speech section text information obtained by the second speech recognition unit 128, the candidates for partial section text information obtained by the first speech recognition unit 126 can be referenced to improve the reliability of the conversation content.
[0077] In addition, the score of the candidate subsection text information may include an acoustic cost and a linguistic cost, and the information integration unit 130 may calculate the speech section score as a weighted average of the sum of the acoustic costs and the sum of the linguistic costs of the candidate subsection text information for each subsection in the speech section. According to this configuration, the speech segment score is obtained by weighting the reliability based on acoustic features and the reliability based on linguistic features, thereby making it possible to adjust the contribution of acoustic features and linguistic features to the reliability of the conversation content.
[0078] A subinterval may also be an interval corresponding to one or more words. According to this configuration, partial section text information indicating the speech content of each word is sequentially displayed on the display unit 30, and integrated text information integrated with the speech section text information indicating the speech content of each speech section is displayed. Displaying partial section text information for each word in real time allows the user to understand the progress of the conversation, and then displaying integrated text information for each speech section ensures the reliability of the conversation content.
[0079] One embodiment of the present invention has been described in detail above with reference to the drawings, but the specific configuration is not limited to that described above, and various design changes and the like are possible within the scope that does not deviate from the gist of the present invention.
[0080] For example, sound collection unit 20 and display unit 30 do not necessarily have to be integrated with conversation support device 10, and any one or a combination of them may be separate from conversation support device 10 as long as they can be connected wirelessly or wired to be able to send and receive various types of data. Utterance information recording unit 134 may be omitted. In the above description, the partial intervals are mainly words, but this is not limiting. The partial intervals may be units other than words, such as phrases or characters.
[0081] The information integration unit 130 does not necessarily have to perform graph integration when integrating speech section text information and subsection text information to generate integrated text information. The information integration unit 130 may replace subsection text information obtained by the first speech recognition process in a certain speech section with speech section text information obtained by the second speech recognition process in that subsection and adopt this as the integrated text information. If the speech section text information obtained by the second speech recognition process includes a subsection for which subsection text information that will be the recognition result cannot be identified, the information integration unit 130 may include the subsection text information obtained by the first speech recognition process for that subsection in the integrated text information rather than discarding it. [Explanation of symbols]
[0082] S1...conversation support system, 10...conversation support device, 110...input / output unit, 120...control unit, 122...acoustic processing unit, 124...feature amount calculation unit, 126...first speech recognition unit, 128...second speech recognition unit, 130...information integration unit, 132...output processing unit, 134...utterance information recording unit, 140...storage unit
Claims
1. a first speech recognition unit that performs speech recognition processing based on a speech signal and determines sub-segment text information for each sub-segment that is a part of an utterance segment; a second speech recognition unit that performs speech recognition processing based on the speech signal and determines speech section text information for each of the speech sections; an information integration unit that integrates the speech section text information with the subsection text information to generate integrated text information; After outputting the partial section text information to a display unit, an output processing unit that outputs the integrated text information to the display unit, The information integration unit a first graph showing candidates for subsection text information for each subsection obtained by the first speech recognition unit, and scores and permutations of the candidates; generating an integrated graph by integrating candidates for subsection text information for each subsection constituting a part of the utterance section obtained by the second speech recognition unit into a second graph showing the scores and permutations of the candidates; calculating, from the scores of the subsection text information candidates for the subsections using the integrated graph, an utterance section score, which is a score of the utterance section text information candidate obtained by arranging the candidates; determining the integrated text information based on the speech segment score; Conversation support device.
2. The output processing unit The conversation support device according to claim 1 , wherein a display mode of a difference section, which is a section in the integrated text information where a difference occurs between the integrated text information and the partial section text information, is determined to be different from a display mode of other sections of the integrated text information.
3. the score of the candidate subsection text information includes an acoustic cost and a linguistic cost; The information integration unit A weighted average of the sum of the acoustic costs and the sum of the linguistic costs of the candidates for the subsection text information for each subsection in the utterance section is calculated as the utterance section score. The conversation support device according to claim 1 .
4. The subinterval is an interval corresponding to one or more words. The conversation support device according to any one of claims 1 to 3.
5. To the computer a first speech recognition step of performing speech recognition processing based on a speech signal and determining sub-segment text information for each sub-segment that is a part of an utterance segment; a second speech recognition step of performing speech recognition processing based on the speech signal to determine speech section text information for each of the speech sections; an information integration step of integrating the sub-section text information into the utterance section text information to generate integrated text information; After outputting the partial section text information to a display unit, an output processing step of outputting the integrated text information to the display unit, The information integration step includes: a first graph showing candidates for subsection text information for each subsection obtained in the first speech recognition step, and scores and permutations of the candidates; generating an integrated graph by integrating candidates for subsection text information for each subsection constituting a part of the utterance section obtained in the second speech recognition step into a second graph showing the scores and permutations of the candidates; calculating, from the scores of the subsection text information candidates for the subsections using the integrated graph, an utterance section score, which is a score of the utterance section text information candidate obtained by arranging the candidates; determining the integrated text information based on the speech segment score; program.
6. The conversation support device according to any one of claims 1 to 4; The display unit Conversation support system.
7. The conversation support device a first speech recognition step of performing speech recognition processing based on a speech signal and determining sub-segment text information for each sub-segment that is a part of an utterance segment; a second speech recognition step of performing speech recognition processing based on the speech signal to determine speech section text information for each of the speech sections; an information integration step of integrating the sub-section text information into the utterance section text information to generate integrated text information; After outputting the partial section text information to a display unit, an output processing step of outputting the integrated text information to the display unit, The information integration step includes: a first graph showing candidates for subsection text information for each subsection obtained in the first speech recognition step, and scores and permutations of the candidates; generating an integrated graph by integrating candidates for subsection text information for each subsection constituting a part of the utterance section obtained in the second speech recognition step into a second graph showing the scores and permutations of the candidates; calculating, from the scores of the subsection text information candidates for the subsections using the integrated graph, an utterance section score, which is a score of the utterance section text information candidate obtained by arranging the candidates; determining the integrated text information based on the speech segment score; Conversation support methods.
Citation Information
Patent Citations
Voice recognition device and speech recognition program
JP2007017911A
Voice recognition system, voice recognition method, and voice recognition program
JP2007033671A
Speech recognition device and vehicle system using the same
JP2009265307A
Conference system, control method therefor, and program
JP2019179480A
Text correction device and text correction method
JP2020197592A