Information processing device and generation method

WO2026203370A1PCT designated stage Publication Date: 2026-10-01NT T INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/013004
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2026-10-01

Smart Images

  • Figure JP2025013004_01102026_PF_FP_ABST
    Figure JP2025013004_01102026_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device (10) comprises: a speech recognition unit that executes speech recognition on an utterance log; a summarization unit that summarizes utterance text obtained by the speech recognition; and a generation unit that generates speech synthesis data for reading out summary text in which the utterance text is summarized.
Need to check novelty before this filing date? Find Prior Art

Description

Information Processing Apparatus and Generation Method

[0001] The present invention relates to an information processing apparatus and a generation method.

[0002] From the perspective of supporting meetings, functions such as a transcription function that converts utterance logs in meetings into text and a summarization function that summarizes the text have been provided.

[0003] Kohei Matsuura, Takanori Ashihara, Takafumi Moriya, Tomohiro Tanaka, Marc Delcroix, Atsunori Ogawa, Ryo Masumura, "LEVERAGING LARGE TEXT CORPORA FOR END-TO-END SPEECH SUMMARIZATION", In Proc. International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2023

[0004] However, in the conventional technologies represented by the above-mentioned transcription function and summarization function, it is necessary to perform operations related to text display such as scrolling, zooming in and out when checking an utterance log, or to keep focusing on the displayed text, which requires time and effort. Therefore, there is room for improvement in terms of improving the efficiency of grasping the content of an utterance log.

[0005] Although the above problem is described by taking an utterance log in a meeting as an example, the same problem can also occur in scenes where an utterance log such as a lecture or a presentation is checked.

[0006] In one aspect, an object of the present invention is to provide an information processing apparatus and a generation method that can achieve improved efficiency in grasping the content of an utterance log.

[0007] According to one aspect, an information processing apparatus includes: a speech recognition unit that performs speech recognition on an utterance log; a summarization unit that summarizes an utterance text obtained by the speech recognition; and a generation unit that generates speech synthesis data for reading out a summary text obtained by summarizing the utterance text.

[0008] According to one embodiment, it is possible to achieve improved efficiency in grasping the content of an utterance log.

[0009] Figure 1 is a block diagram showing an example of the functional configuration of an information processing device. Figure 2 is a diagram showing an example of speech recognition results. Figure 3 is a diagram showing an example of the sorting results of spoken text. Figure 4 is a schematic diagram showing an example of the input / output configuration of LLM. Figure 5 is a diagram showing an example of a prompt. Figure 6 is a diagram showing an example of summary text. Figure 7 is a flowchart showing the procedure for DB construction processing. Figure 8 is a flowchart showing the procedure for generation processing. Figure 9 is a flowchart showing the procedure for speech synthesis processing. Figure 10 is a diagram showing an example of hardware configuration.

[0010] The following describes embodiments for implementing the information processing device and generation method relating to this disclosure (hereinafter referred to as "Embodiments") with reference to the attached drawings. It should be noted that these embodiments represent only one example or aspect, and the following description does not limit the structure, operation, function, properties, characteristics, methods, and applications relating to this disclosure.

[0011] <Overall Configuration> Figure 1 is a block diagram showing an example of the functional configuration of the information processing device 10. For example, Figure 1 shows an information processing device 10 that provides a generation function to generate speech synthesis data that reads aloud a summary text which is a summary of the speech text obtained from the speech log.

[0012] The following is merely one example of a usage scenario in which speech logs are collected, using a conferencing system 3 that provides communication functions for sharing video, audio, documents, etc., between multiple locations.

[0013] In one embodiment, the information processing device 10 may be implemented by a server device. For example, the information processing device 10 can provide the above generation function as a cloud service by running a PaaS (Platform as a Service) type middleware or a SaaS (Software as a Service) type application.

[0014] As shown in Figure 1, the information processing device 10 can be connected to the conference system 3 and the user terminal 30 via a network NW in a communicative manner. For example, the network NW may be implemented by any type of communication network, such as the Internet or a LAN (Local Area Network), whether wired or wireless.

[0015] Conference System 3 is a system that provides communication functions for sharing video, audio, documents, and other materials between multiple locations. For example, Conference System 3 may be implemented using any type of web conferencing system, such as on-premise, cloud-based, or browser-based. However, it is not limited to these, and Conference System 3 may also be implemented using a video conferencing system.

[0016] The user terminal 30 is a terminal device used by a user who receives the above-described generation function. As merely an example, the user terminal 30 may be used not only by participants in a meeting provided by the conference system 3, but also by all related parties. For example, the user terminal 30 may be implemented using any computer, including personal computers, smartphones, tablet devices, and wearable devices.

[0017] While the above-mentioned prediction function is provided as a cloud service, it is not limited to this. For example, the above-mentioned prediction function may be provided on-premises. Alternatively, the above-mentioned prediction function may be packaged as a feature of a service or application provided by a service provider offering services related to remote conferencing.

[0018] Furthermore, while the above-mentioned prediction function is implemented as an example in a client-server system, it is not limited to this. For example, the above-mentioned prediction function may be provided as a standalone function by having an application running on a user terminal participating in a meeting conducted by the conference system 3 execute processing corresponding to the above-mentioned prediction function on the user terminal.

[0019] <One aspect of the challenge> As explained in the background technology section above, when reviewing speech logs, it is necessary to perform operations related to the display of text, such as scrolling or zooming, or to focus on the displayed text. Therefore, there is room for improvement in terms of making the understanding of the content of speech logs more efficient.

[0020] <One aspect of the problem-solving approach> Therefore, the generation function according to this embodiment generates speech synthesis data that reads aloud a summary text which is a summary of the speech text obtained from the speech log. Hereinafter, the speech synthesis data that reads aloud the summary text may be referred to as "summary voice".

[0021] One aspect of this is that, according to the summarized voice, when reviewing the speech log, users can understand the content of the speech log without having to perform operations related to the display of text, such as scrolling or zooming in and out, thus enabling hands-free operation. Another aspect is that, according to the summarized voice, when reviewing the speech log, users can understand the content of the speech log without having to intently look at the displayed text, thus enabling eyes-free operation.

[0022] Therefore, the generation function according to this embodiment makes it possible to improve the efficiency of understanding the content of speech logs.

[0023] <Configuration of Information Processing Device 10> Next, the functional configuration of the information processing device 10 that provides the above generation function will be described. Figure 1 schematically shows the blocks related to the generation function of the information processing device 10. As shown in Figure 1, the information processing device 10 has a communication control unit 11, a storage unit 13, and a control unit 15. Note that Figure 1 only shows a selection of the functional units related to the above generation function, and the information processing device 10 may also have functional units other than those shown.

[0024] The communication control unit 11 is a functional unit that controls communication with other devices such as the conference system 3. In one embodiment, the communication control unit 11 can be implemented by a network interface card such as a LAN card. In one aspect, the communication control unit 11 receives audio data input during a conference conducted by the conference system 3, or receives generation requests from the user terminal 30 requesting the generation of the summary audio. In another aspect, the communication control unit 11 outputs a response to the generation request, such as summary text or summary audio, to the user terminal 30.

[0025] The storage unit 13 is a functional unit that stores various types of data. In one embodiment, the storage unit 13 may be implemented by internal, external, or auxiliary storage of the information processing device 10. For example, the storage unit 13 stores a voice database (DB: Database) 13A. The voice database 13A will be described later in conjunction with the scenes in which registration or retrieval of the voice database 13A is performed.

[0026] The control unit 15 is a functional unit that performs overall control of the information processing device 10. For example, the control unit 15 can be implemented by a hardware processor. As shown in Figure 1, the control unit 15 includes a user communication unit 15A, a speaker management unit 15B, a speech recognition unit 15C, a summarization unit 15D, and a generation unit 15E. The control unit 15 may also be implemented by hardwired logic or the like.

[0027] The user communication unit 15A has a function to control the transmission of information between the information processing device 10 and the user terminal 30. In one embodiment, the user communication unit 15A can receive a request to generate a summary voice via the user terminal 30. When receiving such a generation request, the user communication unit 15A can receive the user terminal 30 to specify meeting identification information as an example of a speech log to be used for generating the summary voice. Furthermore, the user communication unit 15A can also receive the user terminal 30 to specify a time limit for the summary voice. The "time limit" here may refer to the upper limit imposed on the playback time of the speech synthesis data that reads the summary text aloud.

[0028] As just one example, the user communication unit 15A can receive user input from the user terminal 30 such as "Summarize this meeting in about one minute." Such user input may be voice input, text written in natural language, or instructions input via operations on GUI (Graphical User Interface) components.

[0029] As will be explained in detail later using Figures 2 to 6, the user communication unit 15A performs the following exchanges with other functional units when generating summarized audio and summarized text. One aspect is that the user communication unit 15A exchanges information such as the utterance text and speaking speed for each speaker, as well as audio data corresponding to the utterances, with the speaker management unit 15B. Another aspect is that the user communication unit 15A instructs the summarization unit 15D to summarize the entire utterance text to fit within a specified time limit, and exchanges information on the estimation of important utterances and the summary text. Further aspects include that, depending on the content of the summary text, the user communication unit 15A requests the generation unit 15E to create synthesized speech that resembles the speaker's voice quality, and to synthesize speech from parts of the summary text that are not specific to a particular speaker, and then integrates the synthesized speech with the audio data of utterances estimated to be of high importance, depending on the content of the summary text.

[0030] The speaker management unit 15B has the function of managing speakers in utterances such as those in meetings. One aspect of this is that when the speaker management unit 15B collects audio data from the conference system 3, it collects audio data for each channel corresponding to the user terminal 30 used by the speaker. In other words, the speaker management unit 15B handles audio data from multiple channels corresponding to each speaker and controls the speech recognition unit 15C to obtain speech recognition results for each speaker, such as utterance text and utterance segments for the audio data. This creates and manages an audio database 13A in which a set of data in which utterance text and utterance segments are associated for each speaker is registered. For example, an utterance segment can be identified by combining one or more of the utterance start time, utterance end time, and utterance segment length obtained as speech recognition results, thereby identifying the audio data corresponding to the utterance segment. Such an audio database 13A is constructed with a schema that can extract audio data corresponding to the speaker and the utterance text obtained as a result of speech recognition. In other respects, the speaker management unit 15B may also have a function to calculate, for each speaker, the number of characters spoken per unit time, for example, one minute, i.e., the speaking speed, from the total number of characters in the utterance text and the total speaking time. In further respects, the speaker management unit 15B may have a function to extract speech data corresponding to the utterance text with the largest number of characters from the set of utterance texts as a sample speech for speech synthesis, in order to synthesize a speech that is close to the speaker's voice quality in speech synthesis.

[0031] The speech recognition unit 15C has the function of performing the task of recognizing information contained in speech, so-called speech recognition. One aspect is that the speech recognition unit 15C can perform tasks such as estimating the speech interval and transcribing the speech text from the speech data input from the speaker management unit 15B. Here, as an example, the speech start time and speech end time are given from the input speech data, but the speech start time or speech end time and the speech interval length may also be obtained. Another aspect is that the speech recognition unit 15C may have the function of extracting speech data of speech intervals along with speech recognition. Further aspect is that the speech recognition unit 15C may have a speaker separation function. Here, an example is given in which each speaker is connected by an independent channel, but if the speaker separation function is present, the speech of multiple speakers may be included in one channel. Furthermore, since actual speech often contains fillers and rephrasing, the speech recognition unit 16 may be implemented by a system that performs summarization along with speech recognition, as described in the technology cited in Non-Patent Literature 1 above. In this case as well, information that can identify the utterance interval, such as the utterance start time, utterance end time, or utterance interval length, can also be obtained.

[0032] The summarization unit 15D has the function of summarizing one or more spoken texts. Such summarization of spoken texts may be implemented by a large language model, or LLM (Large Language Model), as an example. Here, as an example, we give an example in which an LLM operates on the information processing device 10, but the LLM does not necessarily have to operate on the information processing device 10. For example, the summarization unit 15D can also call the process of summarizing spoken texts by sending an API (Application Programming Interface) to an MLaaS (Machine Learning as a Service) or the like that provides a summarization service.

[0033] One aspect of this is that the summarization unit 15D receives the target speech recognition result and speaker information from the user communication unit 15A and summarizes it in a number of characters that fits within the time limit specified in the generation request. At this time, the summarization unit 15D can also embed instructions to input into the LLM that present important utterances and their speakers. Although an example using LLM has been given here, summarization may also be achieved using tools or services that perform equivalent functions.

[0034] The generation unit 15E has the function of generating speech synthesis data that reads a summary text aloud. In one aspect, the generation unit 15E receives audio data and speech synthesis text from the user communication unit 15A for speech synthesis that mimics the speaker's voice quality, and creates speech synthesis data based on the speaker's voice quality contained in the audio data and according to the content described in the speech synthesis text. In another aspect, the generation unit 15E can also create speech data using synthesized speech that is not based on a specific speaker. For example, by applying the speech synthesis engine described in Reference 1 listed below, as well as the technologies described in References 2 and 3 listed below, it is possible to achieve speech synthesis that reflects the speaker's voice quality based on the meeting speech log.

[0035] Reference 1: “OpenAI Voice Engine” [searched on March 25, 2020], Internet <URL: https: / / openai.com / index / navigating-the-challenges-and-opportunities-of-synthetic-voices / > Reference 2: International Publication No. 2023 / 007306 Reference 3: International Publication No. 2023 / 036399

[0036] <Application Examples> Next, we will list examples of usage scenarios in which the above generation function is applied and explain the processing content corresponding to each example.

[0037] As merely an example, let's consider a case where the above generation function is applied to the speech log of a remote meeting with four participants: speaker 1, speaker 2, speaker 3, and speaker 4. For example, each participant's user terminal 30 in the remote meeting has an independent channel, and the audio data input from each participant's user terminal 30 is input to the information processing device 10 separately for each channel.

[0038] Under this input system, the speaker management unit 15B manages audio data for the number of participants, i.e., the number of channels. For example, the speaker management unit 15B can obtain the spoken text and utterance segments obtained in the transcription task by having the speech recognition unit 15C perform speech recognition for each participant, i.e., speaker.

[0039] Figure 2 shows an example of speech recognition results. Figure 2 shows speech recognition results for four channels, Ch1 to Ch4. For example, in the example of channel Ch1 shown in Figure 2, two speech texts corresponding to two speech segments are extracted and illustrated from the speech recognition results for the audio data input from channel Ch1. Specifically, the speech text "Um, the main topic is the removal of playground equipment in the park," and timestamps of speech start time "29.735" and speech end time "34.515" are obtained. Furthermore, the speech text "Subtopic, um, I'd like to start by talking about the necessity of playground equipment," and timestamps of speech start time "35.518" and speech end time "41.045" are obtained. Speech recognition results can be obtained in a similar format for the other channels Ch2 to Ch4, although the values ​​of the speech text and speech segments may differ. In this context, the speaker's label can be any information; the speaker may be identified by their attribute information, such as their name or affiliation.

[0040] The speaker management unit 15B constructs the speech database 13A by integrating the utterance text, utterance start time and utterance end time into the speech data for each speaker based on the utterance text and utterance section obtained from the speech recognition unit 15C. When constructing the speech database in this manner, the speaker management unit 15B sorts the set of utterance texts for all speakers 1 to 4 in the order of utterances of the entire meeting based on the utterance start time of the utterance of each speaker. Note that when sorting utterances in order, if consecutive utterances are from the same speaker, the consecutive utterances can be collectively merged into one data entry in the speech database 13A.

[0041] For example, when the utterance texts of channels Ch1 to Ch4 shown in Fig. 2 are sorted in utterance order, the result shown in Fig. 3 is obtained. Fig. 3 is a diagram showing an example of a sorted result T1 of utterance texts. For example, in the example of the sorted result T1 of utterance texts shown in Fig. 3, this means that utterances are made in the order of speaker 1, speaker 1, speaker 2, speaker 3, speaker 4, ... At this time, since the speakers of the first utterance and the second utterance are both speaker 1, the first utterance text and the second utterance text can be merged into one data entry as: "Um, regarding the main topic, I am talking about the removal of playground equipment in the park. Regarding the sub-topic, um, I would like to start a little talk from the point of the necessity of the playground equipment." At this time, the speech database 13A is constructed such that information of speech data (such as file information and time information) can be obtained by performing full-text search on each utterance or the merged utterance text. Reference 4 listed below can be applied to constructing such a speech database 13A.

[0042] Reference 4: "Executing Japanese full-text search using pg_bigm", [retrieved on March 25, 2024], Internet <URL: https: / / jpn.nec.com / postgresql / technical_info / pg_bigm_v2.html>

[0043] Further, the speaker management unit 15B calculates, for each speaker, the total number of characters and the total speech duration of M utterance texts produced by the speaker. For example, the speaker management unit 15B can calculate the total number of characters by counting all characters included in the M utterance texts, or may calculate the total number of characters by excluding characters such as blanks and punctuation marks from counting. Further, the speaker management unit 15B can calculate the total speech duration by summing the lengths of utterance sections of the M utterance texts, for example, the differences between the utterance start times and the utterance end times. Then, based on the total number of characters and the total speech duration, the speaker management unit 15B can calculate the number of characters uttered per unit time as the speech rate Sc_n of speaker n. For example, the speech rate Sc_n can be calculated according to the following equation (1). Further, the speaker management unit 15B calculates a statistical value of Sc_n, for example an average value or a median value, as the speech rate Sc for the entire conference.

[0044] Sc_n: Speech rate of speaker n (number of characters / minute) = N: Total number of characters in utterance texts of speaker n / T: Total speech duration of speaker n (Σ length of each utterance section: in units of minutes) ... (1)

[0045] Further, the speaker management unit 15B extracts, as sample speech for speech synthesis, for each speaker, audio data corresponding to the utterance text having the maximum number of characters among the M utterance texts of said speaker.

[0046] The above preprocessing is executed in real time concurrently with the progress of the remote conference, or after the end of the remote conference when audio data for each channel is input. As a result, the speech database 13A is constructed by integrating data in which speech recognition results such as utterance texts and utterance sections, the speech rate of each speaker, the speech rate of the entire conference, and sample speech for speech synthesis are associated with the audio data of each speaker.

[0047] After constructing such a speech database 13A, a user specifies the time limit for a summarized speech, such as "summarize the content in about 30 seconds", for the audio data of the remote conference via voice, keyboard input, screen operation or the like to the user communication unit 15A, whereby a summarized speech in which the content of the utterance log of the conference is summarized within the specified time limit can be obtained.

[0048] In this case, the user communication unit 15A obtains all the utterances from the remote meeting from the speaker management unit 15B in a format sorted in utterance order, i.e., chronological order, so that the large-scale language model can grasp the flow of the remote meeting and perform summarization. For example, the speaker management unit 15B sorts all the utterances from the speech database 13A in chronological order based on the utterance start time, and then outputs an utterance log, for example, the sorted result T1 of the utterances shown in Figure 3, with each sorted utterance labeled with one of the speakers from 1 to 4.

[0049] Furthermore, the user communication unit 15A obtains the overall speaking speed Sc from the speaker management unit 15B and calculates the number of characters (summary character count) when speaking within the time limit of "30 seconds" instructed via the user terminal 30. At this time, a margin may be set by multiplying the summary character count by a coefficient greater than 1. For example, suppose the speaking speed Sc is 360 characters / minute. In this case, since the time limit is 30 seconds, the calculation result of Sc * 0.5, which is 180 characters, is used as the basis, and an arbitrary margin is allowed to be factored in to set the upper and lower limits. Here, as an example, a margin of ±5% is set, resulting in a lower limit of 171 characters and an upper limit of 189 characters. Of course, a margin does not necessarily have to be set for the summary character count. Also, although an example using the speaking speed Sc is given here to calculate the summary character count, the speaking speed Sc_n of individual speakers can also be used. For example, when receiving a request to generate summary audio, a specific speaker n can be input via user input. P This can also be applied to use cases that require the generation of summarized speech by focusing on specific speaker n. P It is possible to narrow down the utterances to specific speakers and embed instruction sentences that quote those utterances and instruct them to summarize. Furthermore, when calculating the number of characters in the summary, instead of the overall speaking speed (Sc) of the meeting, the number of characters of a specific speaker specified by user input (n) is used. P Speech speed Sc_n P You can also use this to calculate the total number of characters.

[0050] The summarization unit 15D provides the speech log, illustrated in Figure 3, and the number of characters in the summary to a large-scale language model to obtain a summary result. At this time, it is instructed to also obtain utterances that are important for the summary.

[0051] Figure 4 is a schematic diagram showing an example of the input / output configuration of the LLM5. As shown in Figure 4, the summarization unit 15D inputs a prompt 20 to the LLM5 that contains the chronological sorting result of the speech log acquired by the speaker management unit 15B, such as the speech text, the speech rate Sc acquired from the voice database 13A, and the number of summary characters set based on the time limit specified via the user terminal 30.

[0052] Figure 5 shows an example of prompt 20. As shown in Figure 5, prompt 20 has embedded the sorted result T1 of the utterance text exemplified in Figure 3 in chronological order as an example of an utterance log. Furthermore, prompt 20 has embedded a range specification of 171 characters at the lower limit and 189 characters at the upper limit as an example of the number of characters to summarize. In addition, prompt 20 has embedded instructions to summarize the utterance log within the summarization character range and to quote important utterances in a way that the speaker can recognize.

[0053] When such a prompt 20 is input to the LLM5, it outputs a summary text 40 as illustrated in Figure 6. Figure 6 is a diagram showing an example of the summary text 40. As shown in Figure 6, the summary text 40 includes a quoted section in which important utterances from each speaker (1 to 4) are quoted in conversational format, and an explanatory section.

[0054] The user communication unit 15A then generates a summary speech based on the summary text 40 output by the LLM 5. For example, in the case of the summary text 40 shown in Figure 6, the quoted portions, such as those enclosed in quotation marks, are treated as summaries of the utterances of speakers 1 to 4. These quoted portions of speakers 1 to 4 can be synthesized into speech according to two modes, Mode 1 and Mode 2, as illustrated below.

[0055] (1) Mode 1 In Mode 1, the realism and persuasiveness of the summarized audio can be improved by using the speaker's audio data corresponding to the quoted portion as is. In this case, a search for audio data is performed in the audio database 13A. For example, the speaker management unit 15B can obtain audio data corresponding to the quoted portion by setting important words such as nouns and verb stems as search keywords from the morphemes obtained as a result of performing morphological analysis on the quoted portion, i.e., the text within the quotation marks. In this case, if an entry contains multiple audio data, the speaker management unit 15B can also perform filtering to exclude audio data that does not contain the keyword.

[0056] (2) Mode 2 In Mode 2, speech synthesis is performed based on the quoted portion to produce a voice similar to that of the speaker, thereby improving the realism of the output speech of the quoted portion. In this case, the speaker management unit 15B obtains a sample voice of the speaker corresponding to the quoted portion from the speech DB 13A, and the generation unit 15E generates the synthesized speech of the quoted portion, i.e., the text within the quotation marks, using the sample voice.

[0057] In addition, by combining Mode 1 and Mode 2 described above, it is possible to generate audio corresponding to the quoted portion of the summarized text.

[0058] Subsequently, the user communication unit 15A calculates the playback time, i.e., the speaking time, from the audio data and speech synthesis data of the quoted portion for each speaker. Then, it calculates the remaining speaking time to be allocated to the playback of the explanatory portion of the summary text other than the quoted portion, i.e., the text other than the quotation marks. Finally, it generates speech synthesis data to read the explanatory portion in a voice quality that is not related to each of the speakers 1 to 4, so that it fits within the calculated remaining speaking time range.

[0059] Then, the user communication unit 15A concatenates the audio data and speech synthesis data of the quoted portions from speakers 1 to 4 with the speech synthesis data of the explanatory portion in the order they appear in the concatenated text, and outputs the concatenated speech synthesis data as a concatenated audio to the user terminal 30. This allows the user terminal 30 to play the concatenated audio.

[0060] <Processing Flow> Next, the processing flow of the information processing device 10 according to this embodiment will be described. Here, the processes executed by the information processing device 10 will be described in the following order: (1) DB registration process, (2) generation process, and (3) speech synthesis process.

[0061] (1) DB Registration Process Figure 7 is a flowchart showing the procedure for the DB registration process. As shown in Figure 7, the speaker management unit 15B acquires channel-specific audio data of the number of participants in a meeting conducted by the conference system 3 (step S101).

[0062] Subsequently, the speaker management unit 15B executes a loop process 1, which repeats the processes from step S102 to step S106 below a number of times corresponding to the number of channels, i.e., the total number of speakers N.

[0063] In other words, the speaker management unit 15B causes the speech recognition unit 15C to perform speech recognition on the voice of the nth speaker (step S102). Then, the speaker management unit 15B acquires the utterance text and utterance segment obtained as a result of the speech recognition in step S102 (step S103).

[0064] Next, the speaker management unit 15B calculates the total number of characters and total speaking time of the M utterance texts obtained as the speech recognition results of the nth speaker (step S104). Then, based on the total number of characters and total speaking time obtained in step S104, the speaker management unit 15B calculates the number of characters spoken per unit time as the speech speed Sc_n of the nth speaker (step S105).

[0065] Furthermore, the speaker management unit 15B extracts audio data corresponding to the speech text with the maximum number of characters among the M speech texts of the nth speaker as sample audio for speech synthesis (step S106).

[0066] As this loop process 1 is repeated, speech recognition results, speech rate, and sample speech for speech synthesis are obtained for each of the N speakers.

[0067] Subsequently, the speaker management unit 15B calculates statistical values ​​of Sc_n for the N speakers, such as the average or median, as the speaking speed Sc for the entire meeting (step S107).

[0068] Furthermore, the speaker management unit 15B sorts the set of speech texts for all N speakers in the order of speech throughout the meeting, based on the start time of each speaker's utterance (step S108).

[0069] Finally, the speaker management unit 15B saves data associated with each speaker's voice data, including speech recognition results such as utterance text and utterance segments, each speaker's speaking speed, the overall speaking speed of the meeting, and sample voices for speech synthesis, to the voice database 13A (step S109), and then terminates the process.

[0070] (2) Generation Process Figure 8 is a flowchart showing the steps of the generation process. This process can be started as an example when a request to generate a summary voice is received from the user terminal 30.

[0071] As shown in Figure 8, the user communication unit 15A receives meeting identification information and a specified time limit for the summary audio from the user terminal 30 (step S301). Subsequently, the user communication unit 15A obtains the overall speaking speed Sc of the meeting from the speaker management unit 15B (step S302).

[0072] Then, the user communication unit 15A calculates the number of characters to be summarized based on the time limit specified in step S301 and the speech rate Sc obtained in step S302 (step S303).

[0073] Subsequently, the summarization unit 15D embeds the speech log corresponding to the meeting identification information specified in step S301 and the summary character count calculated in step S303 into the prompt (step S304).

[0074] Then, the user communication unit 15A inputs the prompt obtained as a result of the embedding in step S304 to the LLM5 (step S305). The user communication unit 15A then obtains the summary text output by the LLM5 that received the prompt 20 (step S306).

[0075] Subsequently, the generation unit 15E performs a "speech synthesis process" to generate speech synthesis data that reads aloud the summary text acquired in step S306 (step S307).

[0076] Finally, the user communication unit 15A outputs the speech synthesis data obtained in step S307 as a compiled voice to the user terminal 30 (step S308), and terminates the process.

[0077] (3) Speech Synthesis Processing Figure 9 is a flowchart showing the procedure for speech synthesis processing. This process is merely an example and corresponds to the process in step S307 shown in Figure 8. As shown in Figure 9, the generation unit 15E executes loop processing 1, which repeats the processes from step S501 to step S502 below a number of times corresponding to the total number of speakers N.

[0078] Specifically, the generation unit 15E searches the audio database 13A for audio data corresponding to the quoted portion of speaker n's speech from the summary text acquired in step S306 (step S501). Then, using the audio data obtained in step S501, the generation unit 15E generates speech synthesis data that reads aloud the quoted portion of speaker n's speech (step S502).

[0079] As this loop process 1 is repeated, speech synthesis data of the quoted portion is generated for each of the N speakers.

[0080] Subsequently, the user communication unit 15A calculates the playback time, i.e., the total speaking time, from the speech synthesis data of the quoted portions for N speakers (step S503). Then, the user communication unit 15A calculates the remaining speaking time to be allocated to the playback of the explanatory portion by subtracting the total speaking time of the quoted portions calculated in step S503 from the time limit of the summarized audio (step S504).

[0081] Next, the user communication unit 15A calculates the speaking speed for reading the explanatory text using the remaining speaking time calculated in step S504 (step S505).

[0082] Then, the user communication unit 15A generates speech synthesis data that reads the explanatory portion at the speaking speed calculated in step S505, using a voice quality that is not related to each speaker (step S506).

[0083] Then, the user communication unit 15A generates a combined audio (step S507) by concatenating the speech synthesis data of the quoted portion generated in loop processing 1 and the speech synthesis data of the explanatory portion generated in step S506 in the order described in the combined text, and then processes it.

[0084] <Summary> As described above, the information processing device 10 according to this embodiment generates speech synthesis data that reads aloud a summary text which is a summary of the speech text obtained from the speech log. Therefore, the information processing device 10 according to this embodiment can improve the efficiency of understanding the content of the speech log.

[0085] <Embodiment 2> Now, although Embodiment 1 of the present disclosure has been described, various applications are possible, and furthermore, it may be implemented in various different forms other than Embodiment 1 described above.

[0086] <Exercise of Creative Ability> The details described in Embodiment 1 above, such as speech logs and prompts, are merely examples and can be changed. Furthermore, the flowchart described in Embodiment 1 above can also be modified in terms of processing order, as long as it is consistent with the original design.

[0087] <System> Unless otherwise specified, the processing procedures, control procedures, specific names, and information including various data and parameters shown in the above document and drawings may be changed at will. For example, one or more of the user communication unit 15A, speaker management unit 15B, speech recognition unit 15C, summarization unit 15D, and generation unit 15E of the information processing device 10 may be composed of separate devices.

[0088] Furthermore, the components of each illustrated device are functionally conceptual and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown. That is, all or part of them can be functionally or physically distributed and integrated in any units according to various loads and usage conditions. Note that each configuration may also be a physical configuration.

[0089] Furthermore, the processing performed by the illustrated apparatus can be implemented, in whole or in part, by a program executed by a hardware processor such as an MPU (Micro-Processing Unit) or CPU (Central Processing Unit), or by hardware using wired logic.

[0090] <Hardware> Next, an example of the hardware configuration of the information processing device 10 described in this embodiment will be explained. For example, it can be implemented by installing a program that realizes the functions of the information processing device 10 on a computer. For example, by having the computer run the above program, which is provided as packaged software or online software, the computer can be made to function as the information processing device 10. The computer referred to here includes desktop or notebook personal computers, rack-mounted server computers, etc. In addition, the computer category also includes smartphones, mobile phones and PHS (Personal Handyphone System) and other mobile communication terminals, as well as PDAs (Personal Digital Assistants). Furthermore, the functions of the information processing device 10 may be implemented on a cloud server.

[0091] An example of a computer that executes the above program (generation program) will be explained using Figure 10. As shown in Figure 10, the computer 1000 has, for example, a memory 1010, a CPU 1020, a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0092] Memory 1010 includes ROM (Read Only Memory) 1011 and RAM (Random Access Memory) 1012. ROM 1011 stores, for example, a boot program such as BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to the hard disk drive 1090. The disk drive interface 1040 is connected to the disk drive 1100. The disk drive 1100 is used to insert a removable storage medium, such as a magnetic disk or an optical disk. The serial port interface 1050 is used to connect, for example, a mouse 1110 and a keyboard 1120. The video adapter 1060 is used to connect, for example, a display 1130.

[0093] Here, as shown in Figure 10, the hard disk drive 1090 stores, for example, the OS 1091, the application program 1092, the program module 1093, and the program data 1094. The storage unit 13 described in the above embodiment is equipped, for example, in the hard disk drive 1090 or the memory 1010.

[0094] Then, the CPU 1020 reads the program module 1093 and program data 1094 stored in the hard disk drive 1090 into the RAM 1012 as needed and executes the above-described procedures.

[0095] Furthermore, the program module 1093 and program data 1094 related to the above-mentioned generation program are not limited to being stored in the hard disk drive 1090, but may also be stored in a removable storage medium and read by the CPU 1020 via a disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 related to the above-mentioned program may be stored in another computer connected via a network such as a LAN or WAN (Wide Area Network) and read by the CPU 1020 via a network interface 1070.

[0096] 3 Conference System 10 Information Processing Device 11 Communication Control Unit 13 Storage Unit 13A Voice Database 15 Control Unit 15A User Communication Unit 15B Speaker Management Unit 15C Speech Recognition Unit 15D Summarization Unit 15E Generation Unit 30 User Terminal

Claims

1. An information processing device comprising: a speech recognition unit that performs speech recognition on a speech log; a summarization unit that summarizes the speech text obtained by the speech recognition; and a generation unit that generates speech synthesis data that reads aloud the summarized text obtained from the speech text.

2. The information processing apparatus according to claim 1, characterized in that the summarization unit causes the large-scale language model to generate the summary text by inputting a prompt into the large-scale language model in which the utterance text and a limited number of characters set based on the speech rate calculated from the utterance log and the time limit specified by user input are embedded.

3. The information processing apparatus according to claim 1, characterized in that the generation unit generates audio that reads aloud the quoted portion of the summary text from audio data corresponding to the quoted portion, and generates audio that reads aloud the explanatory portion of the summary text other than the quoted portion using synthesized speech.

4. A generation method performed by an information processing device, comprising: a speech recognition step of performing speech recognition on a speech log; a summarization step of summarizing the speech text obtained by the speech recognition; and a generation step of generating speech synthesis data that reads aloud the summarized text obtained by summarizing the speech text.