Conversation content output device, conversation content output method, and conversation content output system
The conversation content output device enhances speaker identification and transcription accuracy by analyzing audio data to detect speech segments, calculate feature quantities, and exclude unimportant sentences, addressing the limitations of existing dialogue summarization systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-04-09
AI Technical Summary
Existing dialogue summarization systems face accuracy issues in speaker identification due to the reliance on textified dialogue data and short audio utterances, leading to decreased speaker identification accuracy.
A conversation content output device and system that utilizes audio data to detect speech segments, calculate feature quantities, and perform speaker identification, followed by speech recognition and analysis to generate a transcript, excluding unnecessary sentences based on speech length, likelihood, and analysis results.
Improves the accuracy of selecting and transcribing relevant conversations between multiple users by accurately identifying speakers and excluding unimportant sentences, resulting in a more precise transcription.
Smart Images

Figure 2026062026000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a conversation content output device, a conversation content output method, and a conversation content output system.
Background Art
[0002] Patent Document 1 discloses a dialogue summarization system that extracts one or more important sentences from dialogue content and generates summary data consisting of the important sentences. The dialogue summarization system is based on dialogue structure data having information on each statement in the dialogue content, information on a score indicating the importance of each statement, and information on blocks in units of consecutive statements for each speaker. Until a predetermined summary condition is satisfied, the statement with the highest score is extracted from the dialogue structure data as an important sentence, a predetermined score is assigned to the first block in which the important sentence is extracted and the second block in the vicinity thereof, and further, a predetermined score is assigned to and added to the scores of each statement included in the first and second blocks according to a predetermined condition.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in Patent Document 1 described above, speaker discrimination is performed by analyzing dialogue data obtained by textifying dialogue content. Therefore, the summary data generated by the dialogue summarization system has a problem from the viewpoint of the accuracy of the person regarded as the speaker of each uttered sentence when used as evidence of the dialogue.
[0005] Furthermore, conventionally, a method for identifying speakers in a dialogue has been to extract the speaker's voiceprint from the audio data of each utterance and perform speaker matching, thereby accurately identifying the speaker of each utterance. However, audio data sometimes includes utterances that are very short. Such audio data contains a small amount of voiceprint, which leads to a decrease in the accuracy of speaker identification.
[0006] This disclosure was devised in view of the conventional circumstances described above, and aims to provide a conversation content output device, a conversation content output method, and a conversation content output system that more appropriately select conversations to be transcribed from among conversations between multiple users included in audio data. [Means for solving the problem]
[0007] This disclosure provides a conversation content output device that generates and outputs a transcript of a conversation between multiple speakers, comprising: an acquisition unit that acquires audio data from which the conversation has been recorded; a detection unit that detects a speech section spoken by any of the speakers from the audio data and calculates the length of the speech section; an identification unit that calculates a feature quantity from the audio data corresponding to the speech section, calculates a likelihood indicating which speaker the calculated feature quantity belongs to, and identifies the speaker of the speech section based on the likelihood; an analysis unit that performs speech recognition on the audio data to generate a plurality of conversation sentences that represent the conversation in text, and generates an analysis result by analyzing the plurality of conversation sentences in chronological order; and an output unit that determines unnecessary sentences to be excluded from the transcript of the conversation based on at least two of the length of the speech section corresponding to each of the plurality of conversation sentences, the likelihood, and the analysis result, and outputs a transcript of the conversation sentences with the unnecessary sentences excluded from the plurality of conversation sentences, and outputs the speaker information of the speech section corresponding to the conversation sentences in association with each other.
[0008] Furthermore, this disclosure provides a conversation content output method performed by at least one processor that generates a transcript of a conversation between multiple speakers, the method comprising: acquiring audio data from which the conversation has been recorded; detecting a speech segment spoken by any of the speakers from the audio data; calculating the length of the speech segment; calculating a feature quantity from the audio data corresponding to the speech segment; calculating a likelihood that the calculated feature quantity belongs to any of the speakers; identifying the speaker of the speech segment based on the likelihood; performing speech recognition on the audio data to generate a plurality of conversation sentences that represent the conversation; generating an analysis result by analyzing the plurality of conversation sentences in chronological order; determining unnecessary sentences to be excluded from the conversation transcript based on at least two of the length of the speech segment corresponding to each of the plurality of conversation sentences, the likelihood, and the analysis result; and outputting a transcript of the conversation sentences with the unnecessary sentences excluded from the plurality of conversation sentences, along with information on the speaker of the speech segment corresponding to the conversation sentence.
[0009] Furthermore, this disclosure relates to a conversation content output system comprising: a device for generating transcripts of conversations of multiple speakers; and a display device that can communicate with the device and displays the transcripts, wherein the device acquires audio data from which the conversation has been recorded, detects a speech segment from the audio data in which any of the speakers are speaking, calculates the length of the speech segment, calculates a feature quantity from the audio data corresponding to the speech segment, calculates a likelihood indicating which of the speakers the calculated feature quantity belongs to, identifies the speaker of the speech segment based on the likelihood, and the audio The system provides a conversation content output system that performs speech recognition on voice data to generate multiple conversation sentences by transcribing the conversation into text, generates analysis results by analyzing the multiple conversation sentences in chronological order, determines unnecessary sentences to be excluded from the conversation transcript based on at least two of the length of the utterance section corresponding to each of the multiple conversation sentences, the likelihood, and the analysis results, and transmits the transcript of the conversation sentences with the unnecessary sentences excluded from the multiple conversation sentences, along with the speaker information of the utterance section corresponding to the conversation sentences, to the display device for display. [Effects of the Invention]
[0010] According to this disclosure, it is possible to more appropriately select the conversations to be transcribed from among the conversations of multiple users included in the audio data. [Brief explanation of the drawing]
[0011] [Figure 1] A diagram showing an example of a use case for the conversation content output system according to the embodiment. [Figure 2] This figure shows an example of the data stored in the registered speaker database and feature database in the embodiment. [Figure 3] Block diagram showing an example of the internal configuration of the authentication analysis device in the embodiment. [Figure 4] A flowchart illustrating an example of the operation procedure of the authentication analysis device in the embodiment. [Figure 5] A flowchart illustrating an example of the operation procedure of the authentication analysis device in the embodiment. [Figure 6] A flowchart illustrating an example of the operation procedure of the authentication analysis device in the embodiment. [Figure 7] A flowchart illustrating an example of the operation procedure of the authentication analysis device in the embodiment. [Figure 8] Diagram explaining transcription example 1 [Figure 9] Diagram explaining transcription example 2 [Modes for carrying out the invention]
[0012] Hereinafter, embodiments specifically disclosing the conversation content output device, conversation content output method, and conversation content output system according to this disclosure will be described in detail with reference to the drawings as appropriate. However, unnecessarily detailed explanations may be omitted. For example, detailed explanations of already well-known matters and redundant explanations of substantially identical configurations may be omitted. This is to avoid the following explanation becoming unnecessarily verbose and to facilitate understanding by those skilled in the art. The accompanying drawings and the following explanation are provided to enable those skilled in the art to fully understand this disclosure and are not intended to limit the subject matter described in the claims.
[0013] First, the use cases of the conversation content output system 100 according to the embodiment will be described with reference to Figures 1 and 2, respectively. Figure 1 is a diagram showing an example of a use case of the conversation content output system 100 according to the embodiment. Figure 2 is a diagram showing an example of data stored in the registered speaker database DB1 and the feature database DB2 in the embodiment.
[0014] The conversation content output system 100 collects audio data, such as group calls, conversations, or web conferences, which capture the utterances of multiple users, and selects the utterances from the audio data that are to be transcribed. The conversation content output system 100 determines which utterances should be excluded from the conversation between multiple users (e.g., interjections or fillers) and which utterances should not be excluded as evidence of agreement, approval, or confirmation between multiple users. Based on the determination results, the conversation content output system 100 outputs the transcribed conversation content (i.e., written conversation) from the conversation between multiple users in association with the speaker identification results (speaker information), thereby generating a transcription result that visualizes the conversation content and the speaker of each utterance.
[0015] The conversation content output system 100 includes a processing device P1, an identification and analysis device S1, a registered speaker database DB1, and a feature database DB2. Note that the conversation content output system 100 shown in FIG. 1 is an example and is not limited thereto. For example, the processing executed by the identification and analysis device S1 may be executable by the processing device P1. In such a case, the processing device P1 and the identification and analysis device S1 are integrally configured. Also, for example, in the case of a web conference where multiple users speak from different locations, the processing device P1 may be not one but multiple. Further, the processing device P1 may be composed of multiple devices. For example, it may be composed of a device capable of collecting the user's spoken voice (e.g., a microphone, etc.) and a device capable of displaying the transcription result (e.g., a monitor, etc.).
[0016] The processing device P1 is realized by, for example, a microphone and a monitor, a Personal Computer (hereinafter referred to as "PC"), a notebook PC, a tablet terminal, a smartphone, or a telephone. The processing device P1 includes a sound collection device capable of collecting the user's spoken voice and a monitor capable of outputting (displaying) a screen SC1 including the transcription result.
[0017] The processing device P1 is connected to be capable of wired or wireless communication with the identification and analysis device S1 and executes data transmission and reception. Here, the wireless communication mentioned here is communication via a wireless Local Area Network (LAN) such as Wi-Fi (registered trademark).
[0018] Note that the screen SC1 mentioned here is a screen in which the conversation content along the time series and the speaker information corresponding to each utterance in the conversation content are written in association with each other and visualized. The screen SC1 shown in FIG. 1 is an example in which the conversation content of users "Mr. A" and "Mr. B" is written in the time series, and includes the conversation text of "Mr. A", "It's like this here.", the conversation text of "Mr. B", "And this is the result.", and the conversation text of "Mr. A", "I see. Then I can feel at ease.",
[0019] The identification and analysis device S1 is connected so as to be able to transmit and receive data to and from the processing device P1, the registered speaker database DB1, and the feature quantity database DB2, respectively.
[0020] The identification and analysis device S1 acquires voice data, which is the conversation voices of a plurality of users transmitted from the processing device P1. The identification and analysis device S1 uses the acquired voice data to execute a speaking time calculation process for measuring the speaking time of each speaking section, a speaker identification process for identifying the speaker of each speaking section, and a text data (conversation text) analysis process for analyzing the conversation content written down from the voice data. The identification and analysis device S1 calculates the importance of each conversation text based on the speaking time, the speaker identification result, and the text analysis result, selects the conversation texts to be excluded from the writing target, and generates output data obtained by excluding the selected conversation texts from all the conversation texts and outputs the output data to the processing device P1.
[0021] The registered speaker database DB1 is a so-called storage, and is configured using a storage medium such as a flash memory, a Hard Disk Drive (hereinafter referred to as "HDD"), or a Solid State Drive (hereinafter referred to as "SSD"). The registered speaker database DB1 stores, for each user, information that enables identification of the user (for example, user ID or user information) and a reference ID for referring to the feature quantity of the voice of each of a plurality of users stored (registered) in the feature quantity database DB2 in association with each other. Further, the registered speaker database DB1 may be configured integrally with the identification and analysis device S1.
[0022] In the present disclosure, as an example, the registered speaker database DB1 stores, for each user, a user ID assigned to each user, user information (in the present disclosure, the surname and name of the user, respectively), and a reference ID in association with each other. Note that the user information is information regarding the user, and is, for example, a user name, a user ID, or identification information assigned to each user. Note that the user ID may be a serial number, or may be a unique number, symbol, or character string that does not overlap.
[0023] The registered speaker database DB1 shown in Figure 2 stores user information corresponding to user ID "rg12345", including the user's last name "Tanaka", first name "〇〇", and reference IDs "referencerf000" and "referencerf001" for referencing the user's features corresponding to user ID "rg12345". It also stores user information corresponding to user ID "rg77777", including the user's last name "Matsushita", first name "□□", and reference IDs "referencerf003" and "referencerf004" for referencing the user's features corresponding to user ID "rg77777". Furthermore, it stores user information corresponding to user ID "rg32109", including the user's last name "Sato", first name "△△", and reference ID "referencerf005" for referencing the user's features corresponding to user ID "rg32109".
[0024] The feature database DB2 is a type of storage, configured using storage media such as HDDs or SSDs. The feature database DB2 stores (registers) each user's feature data (binary data) in association with a reference ID.
[0025] As an example, the feature database DB2 shown in Figure 2 stores the feature quantities of any user's voice registered in the registered speaker database DB1, along with the reference ID assigned to each feature quantity, for each feature quantity.
[0026] The feature database DB2 shown in Figure 2 stores the user voice feature vectors "001011..." corresponding to reference ID "rf000", the user voice feature vectors "100111..." corresponding to reference ID "rf001", and the user voice feature vectors "110101..." corresponding to reference ID "rf002".
[0027] Referring to Figure 3, an example of the internal configuration of the identification and analysis device S1 will be described. Figure 3 is a block diagram showing an example of the internal configuration of the identification and analysis device S1 in this embodiment.
[0028] The identification and analysis device S1 includes a communication unit 10, a processor 11, and a memory 12.
[0029] The communication unit 10 is connected to the processing unit P1 via wired or wireless connection and performs data transmission and reception. The communication unit 10 outputs various data transmitted from the processing unit P1 to the processor 11. The communication unit 10 also transmits various data transmitted from the processor 11 to the processing unit P1, the registered speaker database DB1, or the feature database DB2.
[0030] The processor 11 is configured using, for example, a Central Processing Unit (CPU) or a Field Programmable Gate Array (FPGA), and works in cooperation with the memory 12 to perform various processing and control. Specifically, the processor 11 refers to the programs and data held in the memory 12 and executes those programs to realize various functions such as the voice input unit 111, the speech interval detection unit 112, the feature quantity calculation unit 113, the feature quantity registration unit 114, the comparison score calculation unit 115, the speech time calculation unit 116, the speech recognition unit 117, the text analysis unit 118, the output correction unit 119, and the transcription result output unit 120.
[0031] The voice input unit 111 receives voice data transmitted from the processing unit P1. The voice data may be data recorded from real-time conversations between multiple users, or it may be pre-recorded data. The voice input unit 111 outputs the input voice data to the speech interval detection unit 112 and the speech recognition unit 117, respectively.
[0032] The speech segment detection unit 112 detects a speech segment in which at least one user is speaking from the audio data output from the audio input unit 111. The speech segment detection unit 112 outputs the audio data corresponding to the detected speech segment to the feature calculation unit 113 and the speech time calculation unit 116.
[0033] The feature calculation unit 113 calculates features from the speech data corresponding to the speech interval output from the speech interval detection unit 112. The feature calculation unit 113 outputs the calculated feature data to the comparison score calculation unit 115. When the feature calculation unit 113 performs the user feature registration process, it outputs the calculated features to the feature registration unit 114.
[0034] The feature registration unit 114 reads various data (for example, user information and features) stored in the registered speaker database DB1 and the feature database DB2, respectively. The feature registration unit 114 outputs the read data to the comparison score calculation unit 115.
[0035] The comparison score calculation unit 115 compares the features output from the feature calculation unit 113 with multiple features read by the feature registration unit 114 for each user. The comparison score calculation unit 115 calculates a score (hereinafter referred to as the "comparison score") that indicates the similarity between the features calculated from the audio data and the features of each user, and outputs it to the output correction unit 119.
[0036] The speech time calculation unit 116 counts and calculates the speech time, which is the length of the speech segment, based on the speech segment audio data output from the speech segment detection unit 112. The speech time calculation unit 116 outputs the calculated speech time information to the output correction unit 119.
[0037] In this disclosure, as an example, the processor 11 counts the speech time using the speech time calculation unit 116. However, instead of counting the speech time, the processor 11 may count the number of phonemes contained in the speech data of the speech interval. In this case, the processor 11 outputs the counted number of phonemes to the output correction unit 119. Alternatively, the processor 11 may count both the speech time and the number of phonemes of the speech interval. In this case, the processor 11 outputs both the counted speech time and the number of phonemes to the output correction unit 119.
[0038] The speech recognition unit 117 performs speech recognition processing on the audio data output from the speech input unit 111, generates text data representing the recognized speech content, and outputs it to the text analysis unit 118 and the transcription result output unit 120.
[0039] The text analysis unit 118 performs natural language processing on the text data output from the speech recognition unit 117 to analyze the conversational text (spoken content). The text analysis unit 118 outputs the analysis results of the conversational text (spoken content) to the output correction unit 119.
[0040] The output correction unit 119 selects the conversation text (utterance content) to be transcribed as a transcription result, that is, to be output to the processing unit P1, based on the comparison score calculated for each user, the utterance time corresponding to the utterance interval, and the text analysis results. The output correction unit 119 outputs the selected utterance content information to the transcription result output unit 120. The output correction unit 119 also performs data mapping from various units to select the utterance content based on the audio data acquisition time used to calculate the feature quantities, the audio data acquisition time corresponding to the utterance interval used to calculate the utterance time, and the audio data acquisition time corresponding to the text analysis results.
[0041] Furthermore, if the output correction unit 119 obtains phoneme count information instead of speech duration information, or if it obtains both speech duration and phoneme count information, it selects the conversational text to be transcribed as the transcription result based on the comparison score, the speech duration or phoneme count, and the text analysis result.
[0042] The transcription result output unit 120 transcribes the conversation content of multiple users based on the text data representing the conversation text (utterance content) output from the speech recognition unit 117 and the conversation text information output from the output correction unit 119. It then generates a screen SC1 containing the transcribed conversation content and the corresponding speaker information, sends it to the processing unit P1, and outputs (displays) it.
[0043] Memory 12 includes, for example, Random Access Memory (RAM) used as work memory when executing each process of the processor 11, and Read Only Memory (ROM) which stores programs and data that define the operation of the processor 11. Data or information generated or acquired by the processor 11 is temporarily stored in RAM. Programs that define the operation of the processor 11 are written to ROM.
[0044] <Processing by the speech interval detection unit and the speech time calculation unit> Next, referring to Figure 4, the processing of the speech interval detection unit 112 and the speech time calculation unit 116 of the identification analysis device S1 will be described. Figure 4 is a flowchart illustrating an example of the operation procedure of the identification analysis device S1 in this embodiment.
[0045] The speech segment detection unit 112 receives the input of audio data output from the audio input unit 111 (i.e., audio input) (St11). The speech segment detection unit 112 determines whether the target segment, which is a segment of the input audio data, contains sound, that is, whether or not it contains a sound that resembles a person's voice (St12).
[0046] The target section may be the section of the input audio data itself when acquiring audio data in real time at predetermined intervals (e.g., 0.5 seconds, 1 second, or 2 seconds), or it may be a section obtained by dividing the input audio data into predetermined intervals (e.g., 1 second, 2 seconds, etc.) when acquiring pre-recorded audio data. When acquiring audio data in real time at predetermined intervals (e.g., 0.5 seconds, 1 second, or 2 seconds, etc.), the processor 11 repeatedly executes the flow shown in Figure 4 at predetermined intervals. When acquiring pre-recorded audio data, the processor 11 repeatedly executes the processes of steps St12 to St19 shown in Figure 4 for all sections obtained by dividing the total length of the audio data into predetermined intervals.
[0047] If the speech segment detection unit 112 determines in step St12 that the target segment has sound (St12, YES), it determines in step St13 whether the target segment immediately preceding the current target segment is silent, that is, whether it does not contain any sound resembling a person's voice. If the target segment is the first segment, the processing in step St13 may be omitted, and the process may proceed to step St14.
[0048] If the speech interval detection unit 112 determines in step St13 that the previous target interval was silent (St13, YES), it places a speech start point mark on the target interval to indicate that a user has started speaking (St14).
[0049] On the other hand, in step St13, the speech interval detection unit 112 determines that if the previous target interval is not silent (St13, NO), then it determines that the user's speech is continuing. The speech time calculation unit 116 increments the speech time counter, which indicates the length of the speech (St15).
[0050] If the speech segment detection unit 112 determines in step St12 that the target segment is not audible (St12, NO), it determines whether the previous target segment is audible (St16). If the target segment is the first segment, the processing in step St16 may be omitted, and the process may proceed to step St17.
[0051] If the speech segment detection unit 112 determines in step St16 that the previous target segment was audible (St16, YES), it places a speech end point mark on the target segment to indicate that any user utterance has ended (St17).
[0052] The speech time calculation unit 116 calculates the speech time, which is the length of the speech interval, based on the length of the audio data from the section marked with a speech start point to the section marked with a speech end point, and records it in the memory 12 (St18). The speech time calculation unit 116 may record the speech time in association with the speech start time and speech end time information.
[0053] The speech time calculation unit 116 resets the current speech time counter value to 0 (zero) (St19).
[0054] <Processing by the Feature Calculation Unit and Comparison Score Calculation Unit> Next, with reference to Figure 5, the processing of the feature quantity calculation unit 113 and the comparison score calculation unit 115 of the discrimination analysis device S1 will be described. Figure 5 is a flowchart illustrating an example of the operation procedure of the discrimination analysis device S1 in the embodiment. The flow shown in Figure 5 is the processing that is performed on the audio data corresponding to the speech segment after the speech segment detection unit 112 has detected one speech segment in the flow shown in Figure 4, that is, after the speech start point mark and speech end point mark have been assigned.
[0055] The feature calculation unit 113 acquires speech data for the speech interval from the interval marked with a speech start point mark to the interval marked with a speech end point mark, based on the speech interval detection unit 112 (St21). The feature calculation unit 113 converts the speech data into features (St22).
[0056] The comparison score calculation unit 115 refers to both the registered speaker database DB1 and the feature database DB2, compares all the features registered in the feature database DB2 with the transformed features, and calculates a comparison score (St23) that indicates the similarity of the transformed features to each user registered in the registered speaker database DB1. The calculation of the comparison score may be performed using known techniques such as Fast Fourier Transform (FFT) or Deep Neural Network (DNN).
[0057] The comparison score calculation unit 115 obtains the value of the largest comparison score among the multiple comparison scores calculated, and the user registration information (e.g., user ID or user information, etc.) corresponding to the feature for which the largest comparison score was calculated, and outputs it to the output correction unit 119 as the result of speaker identification (St24).
[0058] As described above, the identification and analysis device S1 in this embodiment can calculate feature quantities from the utterances spoken by the user, thereby enabling it to calculate more feature quantities more efficiently from a single utterance, and thus improving the accuracy of user identification. In other words, the identification and analysis device S1 can identify with higher accuracy which user uttered each conversational sentence (utterance), and can generate transcription results that can be used as evidence of interactions between multiple users, meeting summaries, or meeting minutes.
[0059] <Processing by the text analysis unit> Next, referring to Figure 6, the processing of the text analysis unit 118 of the identification and analysis device S1 will be described. Figure 6 is a flowchart illustrating an example of the operation procedure of the identification and analysis device S1 in the embodiment.
[0060] The speech recognition unit 117 performs speech recognition on the audio data output from the speech input unit 111 and generates text data representing the spoken content. The speech recognition unit 117 separates the strings representing the spoken content contained in the generated text data with periods or commas and generates conversational sentences (text data) one by one. The speech recognition unit 117 outputs the generated conversational sentence data to the text analysis unit 118 one sentence at a time.
[0061] The text analysis unit 118 receives input of a single sentence of conversation output from the speech recognition unit 117 (St31). The text analysis unit 118 performs natural language processing on the input conversation sentence and analyzes the conversation sentence. Based on the analysis results of the conversation sentence, the text analysis unit 118 determines whether the conversation sentence is a simple exchange of words (St32). A simple exchange of words here refers to responses, reactions, confirmations, agreements, interjections, or fillers such as "yes" or "uh, that's right."
[0062] If the text analysis unit 118 determines in step St32 that the conversation text is a simple exchange of words (St32, YES), it determines whether the analysis result of the conversation text immediately preceding the current conversation text is a question (St33).
[0063] On the other hand, if the text analysis unit 118 determines in step St32 that the conversation text is not a simple exchange of words (St32, NO), it determines that the importance of the current conversation text is "high" (St34).
[0064] Furthermore, if the text analysis unit 118 determines in step St33 that the analysis result of the previous conversation text is a question (St33, YES), it determines that the importance of the current conversation text is "high" (St34).
[0065] On the other hand, if the text analysis unit 118 determines in step St33 that the analysis result of the previous conversation text is not a question (St33, NO), it determines that the importance of the current conversation text is "low" (St35).
[0066] Although not shown in Figure 6, if the speech is unclear and cannot be recognized, as in the case of the speech COM24 shown in Figure 9 later, the text analysis unit 118 may determine that the importance of the corresponding conversational text is "low".
[0067] As described above, the identification and analysis device S1 in the embodiment can determine that a conversational sentence is of high importance if, based on the preceding conversational sentence, it determines that this simple response is a conversational sentence expressing the user's intention (for example, agreement, approval, or confirmation). Conversely, if the identification and analysis device S1 determines that a conversational sentence is of low importance if, based on the preceding conversational sentence, it determines that this simple response is not a conversational sentence expressing the user's intention.
[0068] <Processing of the output correction unit> Next, the processing of the output correction unit 119 of the discrimination analysis device S1 will be described with reference to Figure 7. Figure 7 is a flowchart illustrating an example of the operation procedure of the discrimination analysis device S1 in the embodiment. In Figure 7, an example is shown in which speech time is used as a condition for determining whether or not to exclude conversational text from the output result (transcription result), but the number of phonemes may also be used, or both speech time and the number of phonemes may be used.
[0069] The output correction unit 119 acquires the authentication score and registered speaker information output from the comparison score calculation unit 115 (St41). The output correction unit 119 acquires the speech time information of the voice data used to calculate the authentication score output from the speech time calculation unit 116 and the importance information of the conversation text output from the text analysis unit 118 (St42).
[0070] The output correction unit 119 determines whether the speech duration of the audio data corresponding to the conversation text is less than the threshold A0 (for example, 3 seconds, 5 seconds, etc.) (St43). The threshold A0 may be set to any number of seconds.
[0071] In step St43, if the output correction unit 119 determines that the utterance time of the audio data corresponding to the conversation text is less than threshold A0 (St43, YES), it determines whether the value of the largest comparison score calculated using the audio data corresponding to the conversation text is less than threshold A1 (St44).
[0072] On the other hand, in the processing of step St43, the output correction unit 119 determines that if the utterance time of the audio data corresponding to the conversation text is not less than the threshold A0 (St43, NO), it will not exclude this conversation text from the output result (transcription result).
[0073] In step St44, the output correction unit 119 determines whether the importance of the conversation text is "low" if it determines that the value of the largest comparison score is less than the threshold A1 (St44, YES).
[0074] On the other hand, in the processing of step St44, if the output correction unit 119 determines that the value of the largest comparison score is not less than the threshold A1 (St44, NO), it determines not to exclude this conversation text from the output result (transcription result).
[0075] In step St45, if the output correction unit 119 determines that the importance of the conversation text is "low" (St45, YES), it decides to exclude the conversation text from the output result (transcription result) (St46).
[0076] On the other hand, in the processing of step St45, if the output correction unit 119 determines that the importance of the conversation text is not "low" (St45, NO), it determines not to exclude this conversation text from the output result (transcription result).
[0077] Furthermore, the output correction unit 119 recalculates the comparison score (St47) based on clustering of feature quantities calculated from the audio data corresponding to each of the multiple conversation sentences including this conversation sentence, or based on the feature quantities of the audio data including this conversation sentence and the preceding and succeeding conversation sentences that are consecutive to it. Note that the processing in step St47 is not mandatory and may be omitted.
[0078] Here, the feature clustering process is described. Based on the utterance time corresponding to each conversation sentence, the output correction unit 119 acquires features corresponding to a predetermined number of conversation sentences (e.g., 4 sentences, 6 sentences, etc.) such that the sum of the utterance time of the conversation sentences that are determined not to be excluded from the output result and the utterance time of at least one conversation sentence that is consecutive with this conversation sentence is a predetermined time (e.g., 1 minute, 2 minutes, etc.). The output correction unit 119 performs clustering based on the acquired features. Based on the clustering results, the output correction unit 119 determines at least one user who uttered the predetermined number of conversation sentences. The output correction unit 119 updates the speaker information identified based on the authentication score with the speaker information determination result based on the clustering results. As a result, the identification analysis device S1 can identify speakers of conversation sentences with short utterance times and few features with higher accuracy.
[0079] For example, based on the clustering results, if the users who uttered a predetermined number of conversational sentences are "Person A" and "Person B", the output correction unit 119 overwrites the speaker information of conversational sentences clustered to "Person A" with "Person A", and the speaker information of conversational sentences clustered to "Person B" with "Person B", respectively.
[0080] Next, the process of recalculating the comparison score will be explained. The output correction unit 119 acquires feature quantities corresponding to conversation sentences that have been determined not to be excluded from the output result, and feature quantities corresponding to at least one conversation sentence that is consecutive to this conversation sentence in the time series before and after. The output correction unit 119 recalculates the comparison score based on a comparison between the feature quantities of the multiple conversation sentences and the registered feature quantities. Note that the recalculation of the comparison score may be performed by the comparison score calculation unit 115. The output correction unit 119 compares the recalculated comparison score with the comparison score acquired from the comparison score calculation unit 115, and determines the registered speaker corresponding to the comparison score with the larger value as the speaker information for the conversation sentence. As a result, the discrimination analysis device S1 can identify speakers of conversation sentences with fewer feature quantities with higher accuracy.
[0081] For example, if the recalculated comparison score is "500" and indicates registered speaker "Mr. A", and the comparison score obtained from the comparison score calculation unit 115 is "450" and indicates registered speaker "Mr. B", the output correction unit 119 determines that the speaker of this conversation is "Mr. A", and overwrites the speaker information from "Mr. B" to "Mr. A".
[0082] In addition, in steps St43 and St44 described above, the interval (time period) of the audio data in which the conversation text was generated may not match the utterance interval (time period) in which the comparison score was calculated. In such cases, the output correction unit 119 may, in step St43, recalculate the utterance time from the start time to the end time of the audio data in which the conversation text was generated, and determine whether the recalculated utterance time is less than threshold A0. Furthermore, in step St44, the output correction unit 119 may, based on the start time and end time of the audio data in which the conversation text was generated, calculate the average value of the comparison score calculated from the audio data in the interval (time period) corresponding to these start and end times, and determine whether the calculated average value of the comparison score is less than threshold A1.
[0083] As described above, the identification and analysis device S1 in this embodiment can accurately identify the user who uttered the voice corresponding to each conversational sentence in a conversation between multiple users, based on the features, and can also select the conversational sentences to be transcribed using natural language processing. As a result, the identification and analysis device S1 can not only exclude simple response sentences (e.g., replies, reactions, confirmations, agreements, interjections, or fillers), i.e., unimportant conversational sentences, from the transcription results, but can also include simple response sentences in the transcription results if they are determined to be highly important based on the flow of the conversation.
[0084] Furthermore, the identification and analysis device S1 in this embodiment outputs the speaker identification result and the transcription result in association. Therefore, the identification and analysis device S1 can generate a transcription result that can be used as evidence exchanged between multiple users.
[0085] Furthermore, even when the comparison score is less than threshold A1 and the speaker identification accuracy is low, the identification analysis device S1 in this embodiment can identify speakers of utterances with short speaking times (less than threshold A0) with higher accuracy by executing the process in step St47.
[0086] Next, an example of the transcription result will be explained with reference to Figure 8. Figure 8 is a diagram illustrating transcription example 1. Note that transcription example 1 shown in Figure 8 is just one example and is not limited to this.
[0087] In the example shown in Figure 8, two users, "Person A" and "Person B," are having a conversation. The conversation in Figure 8 proceeds in the following order: Person A's utterance COM11 "This is how it is," Person B's utterance COM12 "Okay," Person A's utterance COM13 "And this is the result," and Person B's utterance COM14 "I see, that's reassuring."
[0088] Table TB11 stores the processing results of each part of the processor 11, including the following items: "Speaker (Speaker)" which indicates the speaker identification result, "Comparison Score" which indicates the calculation result of the comparison score, "Length" which indicates the calculation result of the utterance time, "Recognition Result" which indicates the speech recognition result, and "Importance" which indicates the importance of the conversation text. Although Table TB11 is illustrated for clarity, its generation is not mandatory and may be omitted.
[0089] The processor 11 calculates a comparison score based on the features derived from the audio data of the utterance COM11, "This is how it is." The processor 11 then determines the comparison score "600" which has the highest value among the calculated comparison scores, and identifies speaker "Mr. A" corresponding to the feature for which this comparison score "600" was derived as the speaker of the utterance COM11.
[0090] The processor 11 calculates a speech duration of "2 seconds" based on the speech start time mark and speech end time mark of the speech data of the speech COM 11.
[0091] Processor 11 performs speech recognition on the speech data of the utterance COM 11 to obtain the recognition result "This is how it is...". Processor 11 performs natural language processing on the recognized conversational sentence "This is how it is...", determines that this conversational sentence is not a simple exchange, and that the preceding conversational sentence is not a question, and determines that its importance is "high".
[0092] Furthermore, the processor 11 calculates a comparison score based on the features calculated from the audio data of the utterance COM12 "Yes." The processor 11 determines the comparison score "400" which has the highest value among the calculated comparison scores, and determines that speaker "Mr. B" who corresponds to the feature for which this comparison score "400" was calculated is the speaker of the utterance COM12.
[0093] The processor 11 calculates a speech duration of "0.5 seconds" based on the speech start time mark and speech end time mark of the speech data of the speech COM 12.
[0094] Processor 11 performs speech recognition on the speech data of the utterance COM 12 and obtains the recognition result "yes". Processor 11 performs natural language processing on the conversational sentence with the recognition result "yes", determines that this conversational sentence is a simple response and that the preceding conversational sentence (utterance COM 11) is not a question, and determines that its importance is "low".
[0095] The processor 11 repeatedly performs the above-described process and determines whether to exclude each conversational sentence from the transcription results based on the speech duration, comparison score, and importance. Based on the conversational sentences that are determined not to be excluded from the transcription results, the processor 11 generates a screen SC11 containing the transcription results and sends it to the processing unit P1 for output.
[0096] The screen SC11 shown in Figure 8 excludes conversational sentences corresponding to utterances COM12 that were determined to have a "low" importance, and includes "Person A: This is how it is," "Person A: And this is the result," and "Person B: I see, that's reassuring."
[0097] Next, an example of the transcription result will be explained with reference to Figure 9. Figure 9 is a diagram illustrating transcription example 2. Note that transcription example 2 shown in Figure 9 is just one example and is not limited to this.
[0098] In the example shown in Figure 9, two users, "Person A" and "Person B," are having a conversation. The conversation in Figure 9 proceeds in the following order: Person A's utterance COM21 "That concludes the explanation. Are you satisfied so far?", Person B's utterance COM22 "Yes.", Person A's utterance COM23 "Now let's move on to the next explanation.", Person B's utterance COM24 "~~~~~".
[0099] Table TB21 stores the processing results of each part of the processor 11, including the following items: "Speaker (Speaker)" which indicates the speaker identification result, "Comparison Score" which indicates the calculation result of the comparison score, "Length" which indicates the calculation result of the utterance time, "Recognition Result" which indicates the speech recognition result, and "Importance" which indicates the importance of the conversation text. Although Table TB21 is illustrated for clarity, its generation is not mandatory and may be omitted.
[0100] Processor 11 calculates a comparison score based on the features derived from the speech data of the utterance COM21, "That concludes the explanation. Are you satisfied so far?". Processor 11 then determines the comparison score "700" which has the highest value among the calculated comparison scores, and identifies speaker "Mr. A" corresponding to the feature from which this comparison score "700" was derived as the speaker of the utterance COM21.
[0101] The processor 11 calculates a speech duration of "4 seconds" based on the speech start time mark and speech end time mark of the speech data of the speech COM 21.
[0102] Processor 11 performs speech recognition on the speech data of the utterance COM21 and obtains the recognition result "That concludes the explanation...". Processor 11 performs natural language processing on the recognized conversational sentence "That concludes the explanation...", determines that this conversational sentence is not a simple exchange, and that the preceding conversational sentence is not a question, and determines that its importance is "high".
[0103] Furthermore, the processor 11 calculates a comparison score based on the features calculated from the audio data of the utterance COM22 "Yes." The processor 11 determines the comparison score "400" which has the highest value among the calculated comparison scores, and determines that speaker "Mr. B" who corresponds to the feature for which this comparison score "400" was calculated is the speaker of the utterance COM22.
[0104] The processor 11 calculates a speech duration of "0.5 seconds" based on the speech start time mark and speech end time mark of the speech data of the speech COM 22.
[0105] Processor 11 obtains the recognition result "yes" by performing speech recognition on the voice data of the utterance COM22. Processor 11 performs natural language processing on the conversational text of the recognition result "yes" and determines that this conversational text is a simple response and that the preceding conversational text (utterance COM21) is a question, and determines that its importance is "high". As a result, the identification and analysis device S1 can output utterances that serve as evidence of agreement, approval, etc., made between multiple users.
[0106] The processor 11 repeatedly performs the above-described process and determines whether to exclude each conversational sentence from the transcription results based on the speech duration, comparison score, and importance. Based on the conversational sentences that are determined not to be excluded from the transcription results, the processor 11 generates the transcription results. The processor 11 generates a screen SC21 containing the transcription results and sends it to the processing unit P1 for output.
[0107] Furthermore, the processor 11 calculates a comparison score based on the features calculated from the speech data of the utterance COM24 "~~~~~". The processor 11 determines the comparison score "680" which has the highest value among the calculated comparison scores, and determines that speaker "Mr. B" who corresponds to the feature for which this comparison score "680" was calculated is the speaker of the utterance COM24.
[0108] The processor 11 calculates a speech duration of "2.8 seconds" based on the speech start time mark and speech end time mark of the speech COM24's voice data.
[0109] The processor 11 performs speech recognition on the speech data of the utterance COM24 and obtains the recognition result "~~~~~". The character "~" shown here indicates that the utterance is unclear and cannot be recognized. Based on the recognition result "~~~~~", the processor 11 determines that the content of the utterance is unclear and that its importance is "low".
[0110] The processor 11 repeatedly performs the above-described process and determines whether to exclude each conversational sentence from the transcription results based on the speech duration, comparison score, and importance. Based on the conversational sentences that are determined not to be excluded from the transcription results, the processor 11 generates the transcription results. The processor 11 generates a screen SC21 containing the transcription results and sends it to the processing unit P1 for output.
[0111] The screen SC21 shown in Figure 9 excludes the conversational text corresponding to the utterance COM24, which was determined to have a "low" importance, and includes "Person A: That concludes the explanation. Are you satisfied so far?", "Person B: Yes.", and "Person A: Now let's move on to the next explanation."
[0112] (Other embodiments) In the above-described embodiment, the identification and analysis device S1 is described as an example in which speaker identification processing is performed for each of the speech segments, but it is not limited to this. Speaker identification processing may be performed only on conversational text or speech segments that meet a predetermined quality from the viewpoint of speaker identification accuracy. For example, if the identification and analysis device S1 determines that the speech data corresponding to the conversational text or speech segment does not have enough features to perform speaker identification, that is, the speech duration or number of phonemes results in low speaker identification accuracy and an uncertain speaker identification result, it may set the importance of the conversational text to "low" or exclude it from the output result (transcription result). In this way, the identification and analysis device S1 can suppress the output of conversational texts where the speaker is uncertain and effectively suppress the generation of records (transcription results) that are contrary to the user's wishes.
[0113] (Note) The following technologies are disclosed based on the above description of embodiments.
[0114] (Technology 1) A conversation content output device (identification and analysis device S1) that generates and outputs a transcript of a conversation between multiple speakers (users), The acquisition unit (communication unit 10) acquires audio data from the aforementioned conversation, A detection unit (speech segment detection unit 112) detects a speech segment spoken by any of the speakers from the aforementioned audio data and calculates the length of the speech segment (speech time), An identification unit (comparison score calculation unit 115) calculates feature quantities from the speech data corresponding to the utterance interval, calculates a likelihood (comparison score) indicating which speaker the calculated feature quantities belong to, and identifies the speaker of the utterance interval based on the likelihood, An analysis unit (text analysis unit 118) performs speech recognition on the aforementioned audio data to generate multiple conversational texts that transcribe the conversation into written form, and generates analysis results (importance) by analyzing the multiple conversational texts in chronological order, The system includes an output unit (output correction unit 119 and transcription result output unit 120) that determines unnecessary sentences to be excluded from the transcript of the conversation based on at least two of the length of the speech interval corresponding to each of the plurality of conversation sentences, the likelihood, and the analysis result (importance information), and outputs a transcript of the conversation sentences with the unnecessary sentences excluded from the plurality of conversation sentences, and outputs the speaker information of the speech interval corresponding to the conversation sentences in association with each other. Conversation content output device (identification and analysis device S1). As a result, the discrimination analysis device S1 can calculate features from the utterances spoken by the user, thereby more efficiently calculating more features from a single utterance and improving the accuracy of user identification. In other words, the discrimination analysis device S1 can more accurately identify which user uttered each conversational sentence (utterance) corresponding to an utterance, and can generate transcription results that can be used as evidence of interactions between multiple users, meeting summaries, or meeting minutes. Furthermore, the discrimination analysis device S1 can remove conversational sentences from multiple users that do not need to be transcribed based on utterance time, comparison score, or importance.
[0115] (Technology 2) If the analysis unit (text analysis unit 118) determines that the conversation text is a text showing a response, it generates an analysis result (importance information) corresponding to the determination result of whether the conversation text immediately preceding the conversation text is a question. The conversation content output device (identification and analysis device S1) described in (Technology 1). As a result, even if the conversational text is a simple exchange, the identification and analysis device S1 can analyze whether or not this simple exchange expresses the user's intention (for example, agreement, approval, or confirmation) based on the preceding (i.e., the previous) conversational text.
[0116] (Technology 3) Based on the analysis results, the output unit determines that the previous conversation sentence is not a question, and if so, it determines that the conversation sentence is an unnecessary sentence. The conversation content output device (identification and analysis device S1) described in (Technology 2). As a result, even if the conversation text is a simple exchange, if the preceding conversation text (i.e., the one before) is a question, the identification and analysis device S1 does not remove this conversation text from the transcription result, and can output a transcription result that includes the user's response to the question.
[0117] (Technology 4) The identification unit (comparison score calculation unit 115) calculates a likelihood (comparison score) indicating which registered speaker the calculated feature belongs to, based on the feature quantities of a plurality of registered speakers that have been registered in advance, and identifies which registered speaker the speaker in the utterance interval belongs to, based on the likelihood (comparison score). A conversation content output device (identification and analysis device S1) described in any one of (Technology 1) to (Technology 3). As a result, the identification and analysis device S1 can improve the accuracy of speaker identification by incorporating features that indicate the individuality of the user's voice into the speaker identification process.
[0118] (Technology 5) The output unit (output correction unit 119 and transcription result output unit 120) determines that the length of the utterance section corresponding to the conversation text is less than a first threshold (threshold A0), and performs clustering based on the feature quantities of the utterance section corresponding to the conversation text and the feature quantities of the utterance section corresponding to at least one other conversation text that is time-series consecutive with the conversation text, and identifies the speaker of the utterance section corresponding to the conversation text. A conversation content output device (identification and analysis device S1) described in any one of (Technology 1) to (Technology 4). As a result, the discrimination analysis device S1 can identify speakers in conversational texts with short speech durations and few features with higher accuracy.
[0119] (Technology 6) The output unit (output correction unit 119 and transcription result output unit 120) determines that the likelihood (comparison score) corresponding to the conversation text is less than the second threshold (threshold A1), calculates a likelihood (comparison score) indicating which speaker the feature quantities of the utterance section corresponding to the conversation text and the feature quantities of the utterance section corresponding to at least one other conversation text that is time-series consecutive with the conversation text belong to, and identifies the speaker of the utterance section corresponding to the conversation text. A conversation content output device (identification and analysis device S1) described in any one of (Technology 1) to (Technology 5). As a result, the discrimination analysis device S1 can identify speakers in conversational texts with fewer features with higher accuracy.
[0120] (Technology 7) The detection unit (speech segment detection unit 112) counts the number of phonemes included in the speech segment, The output unit (output correction unit 119 and transcription result output unit 120) determines unnecessary sentences to be excluded from the transcription of the conversation based on the length of the speech interval or the number of phonemes corresponding to each of the plurality of conversation sentences, or at least two of the likelihood and the analysis result (importance). A conversation content output device (identification and analysis device S1) described in any one of (Technology 1) to (Technology 6). This allows the discrimination and analysis device S1 to further use the number of phonemes to determine which conversational sentences should be excluded from the transcription results.
[0121] (Technology 8) A method for outputting conversation content performed by at least one processor 11 that generates a transcript of a conversation between multiple speakers (users), The audio data of the aforementioned conversation is acquired, From the aforementioned audio data, a speech segment spoken by any of the speakers is detected, and the length of the speech segment is calculated. Feature quantities are calculated from the speech data corresponding to the aforementioned speech interval, and the likelihood (comparison score) indicating which speaker the calculated feature quantities belong to is calculated. Based on the likelihood described above, the speaker of the utterance segment is identified. The audio data is subjected to speech recognition to generate multiple conversational sentences that represent the conversation, and the multiple conversational sentences are analyzed in chronological order to generate analysis results (importance information). Based on at least two of the following: the length of the speech interval corresponding to each of the plurality of conversation sentences, the likelihood (comparison score), and the analysis result (importance information), unnecessary sentences to be excluded from the conversation transcript are determined, and the transcript of the conversation sentence from which the unnecessary sentences have been excluded is output in association with the speaker information of the speech interval corresponding to the conversation sentence. How to output conversation content. As a result, the processor 11 of the identification and analysis device S1 can calculate features from the utterances spoken by the user, thereby more efficiently calculating more features from a single utterance and improving the accuracy of user identification. In other words, the processor 11 can more accurately identify which user uttered each conversational sentence (utterance) corresponding to an utterance, and can generate transcription results that can be used as evidence of interactions between multiple users, meeting summaries, or meeting minutes. Furthermore, the processor 11 can remove conversational sentences from multiple users that do not need to be transcribed based on utterance time, comparison score, or importance.
[0122] (Technology 9) A device (identification and analysis device S1) that generates a transcript of a conversation between multiple speakers (users), A conversation content output system 100 comprising a display device (processing device P1) that can communicate with the aforementioned device (identification and analysis device S1) and displays the transcript, The aforementioned device (identification and analysis device S1) The audio data of the aforementioned conversation is acquired, From the aforementioned audio data, a speech segment spoken by any of the speakers is detected, and the length of the speech segment is calculated. Feature quantities are calculated from the speech data corresponding to the aforementioned speech interval, and the likelihood (comparison score) indicating which speaker the calculated feature quantities belong to is calculated. Based on the likelihood described above, the speaker of the utterance segment is identified. The audio data is subjected to speech recognition to generate multiple conversational sentences that represent the conversation, and the multiple conversational sentences are analyzed in chronological order to generate analysis results (importance information). Based on at least two of the following: the length of the speech interval corresponding to each of the plurality of conversation sentences, the likelihood (comparison score), and the analysis result (importance information), unnecessary sentences to be excluded from the conversation transcript are determined, and the transcript of the conversation sentence from which the unnecessary sentences have been excluded, along with the speaker information of the speech interval corresponding to the conversation sentence, is transmitted to the display device (processing device P1) for display. Conversation content output system 100. As a result, the conversation content output system 100 can calculate features from the utterances spoken by the user, thereby more efficiently calculating more features from a single utterance and improving the accuracy of user identification. In other words, the conversation content output system 100 can more accurately identify which user spoke each conversation sentence (utterance) corresponding to an utterance segment, and can generate transcripts that can be used as evidence of interactions between multiple users, meeting summaries, or meeting minutes. Furthermore, the conversation content output system 100 can remove conversation sentences from multiple users that do not need to be transcribed based on utterance time, comparison score, or importance.
[0123] Although various embodiments have been described above with reference to the drawings, it goes without saying that this disclosure is not limited to such examples. It is clear to those skilled in the art that various modifications, alterations, substitutions, additions, deletions, and equivalents can be conceived within the scope of the claims, and these are also understood to fall within the technical scope of this disclosure. Furthermore, the components of the various embodiments described above can be combined arbitrarily without departing from the spirit of the invention. [Industrial applicability]
[0124] This disclosure is useful as a conversation content output device, conversation content output method, and conversation content output system for more appropriately selecting conversations to be transcribed from among conversations between multiple users included in audio data. [Explanation of Symbols]
[0125] 10 Communications Department 11 processors 12 memory 100 Conversation Content Output System 111 Voice Input Section 112 Speech interval detection unit 113 Feature Calculation Unit 114 Feature Registration Unit 115 Comparison Score Calculation Unit 116 Speech time calculation unit 117 Voice Recognition Unit 118 Sentence Analysis Department 119 Output Correction Section 120 Result Output Section DB1 Registered Speaker Database DB2 Feature Database P1 Processing Unit S1 identification analysis device SC1, SC11, SC21 screens
Claims
1. A conversation content output device that generates and outputs a transcript of a conversation between multiple speakers, An acquisition unit that acquires audio data from the aforementioned conversation, A detection unit that detects a speech segment spoken by any of the speakers from the aforementioned audio data and calculates the length of the speech segment, An identification unit that calculates feature quantities from the speech data corresponding to the utterance interval, calculates a likelihood indicating which speaker the calculated feature quantities belong to, and identifies the speaker of the utterance interval based on the likelihood, An analysis unit that performs speech recognition on the aforementioned audio data to generate multiple conversational sentences by transcribing the conversation into text, and generates analysis results by analyzing the multiple conversational sentences in chronological order, The system includes an output unit that determines unnecessary sentences to be excluded from the transcript of the conversation based on at least two of the length of the speech interval corresponding to each of the plurality of conversation sentences, the likelihood, and the analysis results, and outputs a transcript of the conversation sentences with the unnecessary sentences excluded from the plurality of conversation sentences, along with the speaker information of the speech interval corresponding to the conversation sentences. A device for outputting conversation content.
2. If the analysis unit determines that the conversation text is a text indicating a response, it generates an analysis result corresponding to the determination result of whether the conversation text immediately preceding the conversation text is a question. The conversation content output device according to claim 1.
3. Based on the analysis results, the output unit determines that the previous conversation sentence is not a question, and if so, it determines that the conversation sentence is an unnecessary sentence. The conversation content output device according to claim 2.
4. The identification unit calculates a likelihood that the calculated feature belongs to which registered speaker based on the feature quantities of a plurality of registered speakers that have been registered in advance, and identifies which registered speaker the speaker in the utterance interval belongs to based on the likelihood. The conversation content output device according to claim 1.
5. If the output unit determines that the length of the utterance segment corresponding to the conversation text is less than a first threshold, it performs clustering based on the feature quantities of the utterance segment corresponding to the conversation text and the feature quantities of the utterance segment corresponding to at least one other conversation text that is time-series consecutive with the conversation text, and identifies the speaker of the utterance segment corresponding to the conversation text. The conversation content output device according to claim 1.
6. If the output unit determines that the likelihood corresponding to the conversation text is less than a second threshold, it calculates the likelihood that the feature quantities of the utterance section corresponding to the conversation text and the feature quantities of the utterance section corresponding to at least one other conversation text that is sequentially continuous with the conversation text belong to which speaker, and identifies the speaker of the utterance section corresponding to the conversation text. The conversation content output device according to claim 1.
7. The detection unit counts the number of phonemes included in the speech interval, The output unit determines which unnecessary sentences to be excluded from the transcript of the conversation based on the length of the speech interval corresponding to each of the plurality of conversation sentences, or at least two of the number of phonemes, the likelihood, and the analysis result. The conversation content output device according to claim 1.
8. A method for outputting conversation content performed by at least one processor that generates a transcript of a conversation between multiple speakers, The audio data of the aforementioned conversation is acquired, From the aforementioned audio data, a speech segment spoken by any of the speakers is detected, and the length of the speech segment is calculated. Feature quantities are calculated from the speech data corresponding to the aforementioned utterance interval, and the likelihood that the calculated feature quantities belong to which speaker is calculated. Based on the likelihood described above, the speaker of the utterance segment is identified. The process involves performing speech recognition on the aforementioned audio data to generate multiple conversational sentences by transcribing the conversation into text, and then generating analysis results by analyzing the multiple conversational sentences in chronological order. Based on the length of the speech interval corresponding to each of the plurality of conversation sentences, the likelihood, and at least two of the analysis results, unnecessary sentences to be excluded from the transcript of the conversation are determined, and the transcript of the conversation sentence from which the unnecessary sentences have been excluded is output in association with the speaker information of the speech interval corresponding to the conversation sentence. How to output conversation content.
9. A device that generates transcripts of conversations between multiple speakers, A conversation content output system comprising a display device that can communicate with the aforementioned device and displays the transcript, The aforementioned device is The audio data of the aforementioned conversation is acquired, From the aforementioned audio data, a speech segment spoken by any of the speakers is detected, and the length of the speech segment is calculated. Feature quantities are calculated from the speech data corresponding to the aforementioned utterance interval, and the likelihood that the calculated feature quantities belong to which speaker is calculated. Based on the likelihood described above, the speaker of the utterance segment is identified. The process involves performing speech recognition on the aforementioned audio data to generate multiple conversational sentences by transcribing the conversation into text, and then generating analysis results by analyzing the multiple conversational sentences in chronological order. Based on the length of the speech segment corresponding to each of the plurality of conversation sentences, the likelihood, and at least two of the analysis results, unnecessary sentences to be excluded from the transcript of the conversation are determined, and the transcript of the conversation sentence from which the unnecessary sentences have been excluded is transmitted to the display device in association with the speaker information of the speech segment corresponding to the conversation sentence, and displayed. A system for outputting conversation content.
Citation Information
Patent Citations
Dialogue summarization system and dialogue summarization program
JP2013120514A