Information processing system, information processing method, and information processing program

The information processing system addresses the issue of perspective mismatch in summary generation by using multiple generative models and feedback mechanisms to deliver contextually relevant summaries.

JP7733853B1Active Publication Date: 2025-09-08MYNAVI CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025084883
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-09-08
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

Conventional summary generation processes using language models fail to provide summaries that reflect the user's intended perspective based on the specific context of their utterance, leading to unsatisfactory results.

Method used

An information processing system that utilizes multiple generative models to generate summaries by considering different perspectives and user feedback, adjusting prompts based on evaluation, and allowing users to select the most appropriate summary result.

Benefits of technology

The system provides summaries that accurately reflect the user's intended perspective by integrating multiple viewpoints and adapting to user feedback, resulting in more relevant and accurate summary results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007733853000001_ABST
    Figure 0007733853000001_ABST
Patent Text Reader

Abstract

An information processing system, an information processing method, and an information processing program are provided that output a summary result that reflects a viewpoint required according to a user's speech scene. [Solution] In the summary generation system 10, the processor of the server 20 acquires voice data indicating the user's speech content uploaded by a general user to the summary service via a general user terminal 60, inputs the voice data and a prompt including a corresponding instruction sentence corresponding to the speech scene identified based on the voice data from among multiple summary instruction sentences indicating summary instructions for each scene registered in correspondence with multiple scenes into a generation model, and outputs a summary result of the acquired speech content as the output result of the generation model.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing system, an information processing method, and an information processing program. [Background technology]

[0002] Patent Document 1 discloses a technique for improving the performance of a language model and providing a more versatile summary generation function. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2024-167144 Summary of the Invention [Problem to be solved by the invention]

[0004] As described in Patent Document 1, a summary generation process using a language model is provided. However, conventional summary generation processes often perform summarization uniformly with a single prompt, and there is a problem in that a summary result that meets the needs of a user who wants to summarize utterance content cannot be obtained, even though the perspective desired by the user varies depending on the purpose of each utterance scene.

[0005] Therefore, an object of the present disclosure is to provide an information processing system, an information processing method, and an information processing program that can provide a summary result that reflects a viewpoint required according to a user's speech scene. [Means for solving the problem]

[0006] The information processing system of the first aspect includes a processor, and the processor acquires voice data indicating the content of a user's utterance; A first prompt including the voice data and a predetermined transcription instruction is generated and input to a first generative model; text data corresponding to the voice data is obtained as an output result of the first generative model; a second prompt including the text data, user information of a user who executed a generation process to generate a summary result corresponding to the voice data, and a predetermined search instruction is generated and input to a second generative model; a speech scene corresponding to the voice data and a corresponding instruction sentence associated with the speech scene are obtained as an output result of the second generative model; generating a third prompt including an instruction to extract summary information related to a first perspective and inputting the third prompt into a third generative model; obtaining summary information related to the first perspective as an output result of the third generative model; generating a fourth prompt including the audio data, the identified speech scene, the summary information related to the first perspective, and an instruction to extract summary information related to a predetermined second perspective and inputting the fourth prompt into a fourth generative model; obtaining summary information related to the second perspective as an output result of the fourth generative model; generating a fifth prompt including the audio data, the identified speech scene, the corresponding instruction sentence, and the summary information related to the first perspective and the second perspective and inputting the fifth prompt into a fourth generative model; Input to the generative model, before Note Fifth A summary of the utterance content obtained as an output result of the generative model is output.

[0007] With the above configuration, the information processing system of the first aspect can provide a summary result that reflects a viewpoint required according to a user's speech scene.

[0009] With the above configuration, one According to the information processing system of this aspect, it is possible to summarize the content of an utterance by taking into consideration multiple pieces of summary information relating to different perspectives in an integrated manner, thereby providing a highly practical summary result that reflects more of the perspectives required depending on the speech scene.

[0010] No. two The information processing system of the first aspect of In an information processing system according to an embodiment, the processor acquires evaluation information indicating a user's evaluation of the output summary result of the utterance content, and if the evaluation information indicates a negative evaluation, receives input of feedback information indicating a reason for the negative evaluation from the user, and based on the feedback information, calculates a summary result that has received the negative evaluation. Fifth Adjust the prompt.

[0011] With the above configuration, two According to the information processing system of this aspect, by adaptively adjusting the prompt taking into account specific feedback information indicating the reason for the negative evaluation, it is possible to improve the generation accuracy so as to obtain summary results that more appropriately reflect the perspective required depending on the speech scene.

[0012] No. three The information processing system of the first aspect or No. Second In an information processing system according to an embodiment, the processor includes a plurality of the Fifth The generative model Fifth A prompt is input to generate a plurality of summary results of the utterance content, and a user is allowed to select a summary result to be output from among the plurality of generated summary results.

[0013] With the above configuration, three According to the information processing system of this aspect, the user can select from summary results obtained by multiple generative models, thereby actively obtaining the summary result that best suits their intentions depending on the speech scene.

[0014] No. four The information processing method of the aspect includes acquiring voice data indicating the content of a user's utterance, A first prompt including the voice data and a predetermined transcription instruction is generated and input to a first generative model; text data corresponding to the voice data is obtained as an output result of the first generative model; a second prompt including the text data, user information of a user who executed a generation process to generate a summary result corresponding to the voice data, and a predetermined search instruction is generated and input to a second generative model; a speech scene corresponding to the voice data and a corresponding instruction sentence associated with the speech scene are obtained as an output result of the second generative model; generating a third prompt including an instruction to extract summary information related to a first perspective and inputting the third prompt into a third generative model; obtaining summary information related to the first perspective as an output result of the third generative model; generating a fourth prompt including the audio data, the identified speech scene, the summary information related to the first perspective, and an instruction to extract summary information related to a predetermined second perspective and inputting the fourth prompt into a fourth generative model; obtaining summary information related to the second perspective as an output result of the fourth generative model; generating a fifth prompt including the audio data, the identified speech scene, the corresponding instruction sentence, and the summary information related to the first perspective and the second perspective and inputting the fifth prompt into a fourth generative model; Input to the generative model, before Note Fifth The computer executes a process of outputting a summary result of the utterance content obtained as an output result of the generative model.

[0015] With the above configuration, four According to the information processing method of the above aspect, it is possible to provide a summary result that reflects a viewpoint required according to a speech scene of a user.

[0016] No. Five The information processing program of the aspect acquires voice data indicating the content of a user's utterance, A first prompt including the voice data and a predetermined transcription instruction is generated and input to a first generative model; text data corresponding to the voice data is obtained as an output result of the first generative model; a second prompt including the text data, user information of a user who executed a generation process to generate a summary result corresponding to the voice data, and a predetermined search instruction is generated and input to a second generative model; a speech scene corresponding to the voice data and a corresponding instruction sentence associated with the speech scene are obtained as an output result of the second generative model; generating a third prompt including an instruction to extract summary information related to a first perspective and inputting the third prompt into a third generative model; obtaining summary information related to the first perspective as an output result of the third generative model; generating a fourth prompt including the audio data, the identified speech scene, the summary information related to the first perspective, and an instruction to extract summary information related to a predetermined second perspective and inputting the fourth prompt into a fourth generative model; obtaining summary information related to the second perspective as an output result of the fourth generative model; generating a fifth prompt including the audio data, the identified speech scene, the corresponding instruction sentence, and the summary information related to the first perspective and the second perspective and inputting the fifth prompt into a fourth generative model; Input to the generative model, before Note Fifth The computer is caused to execute a process of outputting a summary result of the utterance content acquired as an output result of the generative model.

[0017] With the above configuration, Five According to the information processing program of the aspect, it is possible to provide a summary result that reflects a viewpoint required according to a speech scene of a user. [Effects of the Invention]

[0018] As described above, the information processing system, information processing method, and information processing program according to the present disclosure can provide a summary result that reflects a viewpoint required according to a user's speech scene. [Brief explanation of the drawings]

[0019] [Figure 1] FIG. 1 is a diagram illustrating an example of a schematic configuration of a summary generation system. [Figure 2] FIG. 2 is a block diagram showing the hardware configuration of a server, an administrative user terminal, and a general user terminal. [Figure 3] FIG. 2 is a block diagram showing an example of various information stored in the storage of the server. [Figure 4] FIG. 10 is a diagram illustrating an example of a registration screen displayed on a display unit of the administrative user terminal. [Figure 5] 10 is a flowchart showing the flow of a generation process for generating a summary result corresponding to audio data. [Figure 6] 10 is a flowchart showing the flow of an adjustment process for adjusting the prompts used to generate a summary result. [Figure 7] FIG. 10 is a diagram showing an example of a display screen of a summary result displayed on a display unit of a general user terminal. [Figure 8] FIG. 10 is a diagram showing an example of a feedback input screen displayed on a display unit of a general user terminal. [Figure 9] FIG. 10 is a diagram showing an example of a display screen of a plurality of summary results displayed on a display unit of a general user terminal. DETAILED DESCRIPTION OF THE INVENTION

[0020] The following describes the summary generation system 10 according to this embodiment. The summary generation system 10 is a system that has the function of generating and outputting a summary result based on audio data.

[0021] (First embodiment) First, a first embodiment of the summary generation system 10 will be described.

[0022] FIG. 1 is a diagram showing an example of a schematic configuration of a summary generation system 10. As shown in FIG. 1, the summary generation system 10 includes a server 20, an administrative user terminal 40, and a general user terminal 60. The server 20, the administrative user terminal 40, and the general user terminal 60 are connected to each other in a state where they can communicate with each other via a network N. The network N may be, for example, the Internet, a LAN (Local Area Network), or a WAN (Wide Area Network).

[0023] The server 20 is a server computer that executes various processes related to the summary generation system 10. The summary generation system 10 provides a summarization service that outputs a summary result according to the content of audio data, and users can use the summarization function by accessing the summarization service. The server 20 may be configured as a single server computer, or may have a distributed configuration in which multiple server computers are linked together.

[0024] The administrative user terminal 40 is a terminal used by an administrative user who is a user of the summary service and can register summary instruction sentences indicating summary instructions for each scene and obtain summary results. On the other hand, the general user terminal 60 is a terminal used by a general user who is a user of the summary service and can only obtain summary results. For example, the administrative user is expected to be a user such as a manager who oversees multiple general users in an organization such as a company.

[0025] In this embodiment, the administrative user terminal 40 and the general user terminal 60 are both configured as PCs (Personal Computers), but are not limited to this. Although the figure shows one administrative user terminal 40 and one general user terminal 60, in reality, there may be multiple administrative user terminals 40 and multiple general user terminals 60 depending on the number of users of the summary service.

[0026] 2 is a block diagram showing the hardware configuration of the server 20, the administrative user terminal 40, and the general user terminal 60. Note that since the server 20, the administrative user terminal 40, and the general user terminal 60 basically have a general computer configuration, the server 20 will be described as a representative.

[0027] 2, the server 20 includes a CPU (Central Processing Unit) 21, a ROM (Read Only Memory) 22, a RAM (Random Access Memory) 23, a storage 24, an input unit 25, a display unit 26, and a communication unit 27. Each component is connected to each other via a bus 28 so as to be able to communicate with each other.

[0028] The CPU 21 is a central processing unit that executes various programs and controls each part. That is, the CPU 21 reads programs from the ROM 22 or the storage 24 and executes the programs using the RAM 23 as a work area. The CPU 21 controls each of the above components and performs various arithmetic processing in accordance with the programs stored in the ROM 22 or the storage 24. The CPU 21 is an example of a "processor" in the present disclosure.

[0029] The ROM 22 stores various programs and various data. The RAM 23 serves as a working area for temporarily storing programs or data.

[0030] The storage 24 is configured by a storage device such as a hard disk drive (HDD), a solid state drive (SSD), or a flash memory, and stores various programs and various data.

[0031] The input unit 25 includes, for example, a keyboard, a mouse, various buttons, a microphone, a camera, and the like, and is used to perform various inputs.

[0032] The display unit 26 is, for example, a liquid crystal display, and displays various information. The display unit 26 may function as the input unit 25 by adopting a touch panel system.

[0033] The communication unit 27 is an interface for communicating with other devices. For this communication, a wired communication standard such as Ethernet (registered trademark) or FDDI, or a wireless communication standard such as 4G, 5G, or Wi-Fi (registered trademark) is used.

[0034] In addition, the functions of the CPU 41, ROM 42, RAM 43, storage 44, input unit 45, display unit 46, communication unit 47, and bus 48 of the administrative user terminal 40, and the CPU 61, ROM 62, RAM 63, storage 64, input unit 65, display unit 66, communication unit 67, and bus 68 of the general user terminal 60 have the same functional configuration as the CPU 21, ROM 22, RAM 23, storage 24, input unit 25, display unit 26, communication unit 27, and bus 28 of the server 20 described above.

[0035] 3 is a block diagram showing an example of various information stored in the storage 24 of the server 20. As shown in FIG. 3, the storage 24 stores an information processing program 30, a generative model 31, and a database 32.

[0036] The information processing program 30 is a program for causing the CPU 21 to execute various processes described below. When executing the information processing program 30, the server 20 executes processing based on the information processing program 30 using the hardware resources shown in Fig. 2. The information processing program 30 is an example of an "information processing program" in the present disclosure.

[0037] The generative model 31 is a so-called generative AI. An example of the generative model 31 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ) and Gemini (Internet Search <url: https: gemini.google.com ?hl="ja">) and other generative AIs. The generative model 31 is configured to realize processing that can be executed by various known generative AIs. The generative model 31 is obtained by causing a neural network to perform deep learning. An instruction statement is input to the generative model 31, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The generative model 31 performs inference on the input inference data in accordance with the instruction indicated by the instruction statement, and outputs the inference result in a data format such as voice data, graph data, table data, image data, and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The generative model 31 is an example of a "generative model" in the present disclosure.

[0038] In this embodiment, five types of generative models 31 are provided: a first generative model 31A, a second generative model 31B, a third generative model 31C, a fourth generative model 31D, and a fifth generative model 31E.

[0039] The first generative model 31A performs a transcription process to convert the voice data into text data.

[0040] The second generative model 31B selects a corresponding instruction sentence according to the speech scene corresponding to the text data from among a plurality of summary instruction sentences registered in association with a plurality of scenes, based on the text data acquired by the first generative model 31A. Here, the "plurality of scenes" includes, for example, "sales negotiation," "internal meeting," "interview," and "recruitment interview."

[0041] The third generative model 31C, the fourth generative model 31D, and the fifth generative model 31E execute summarization processing using the corresponding directive statement selected by the second generative model 31B. The specific processing contents of the third generative model 31C, the fourth generative model 31D, and the fifth generative model 31E will be described later.

[0042] Various types of information related to the summarization service are stored in the database 32. For example, the database 32 stores user information that allows users to be identified, prompt information that indicates summarization instructions registered for each administrative user, and result information that indicates the summarization results corresponding to the voice data.

[0043] The user information includes, for example, the user ID, name, organization, authority classification, and login history. The prompt information includes the content of the summary instruction, scene information associated with the summary instruction, registrant information, and registration date and time. The result information includes the audio data, the generated summary result, the summary execution date and time, identification information of the prompt used, and user evaluation information.

[0044] 4 is a diagram showing an example of a registration screen displayed on the display unit 46 of the administrative user terminal 40. The registration screen is a screen for an administrative user to register and manage summary instruction sentences that can be used by the administrative user and multiple general users that the administrative user supervises in the summary service.

[0045] The registration screen shown in FIG. 4 displays list information 50 and an edit button 55. The list information 50 is information in a table format showing a list of summary instruction statements registered by the administrative user, and includes a plurality of records 51, 52, and 53.

[0046] Each record includes the ID, title, and content of the summary instruction. For example, record 51 has an ID of "1," a title of "First Interview," and a content of "This is a prompt for the first interview...," record 52 has an ID of "2," a title of "Second Interview," and a content of "This is a prompt for the second interview...," and record 53 has an ID of "3," a title of "Interview Preparation," and a content of "This is a prompt for interview preparation...."

[0047] The edit button 55 is an operation button for editing the information of the summary instruction statement shown in the list information 50. When the edit button 55 is pressed, an editing screen is displayed, and the administrative user can register a new summary instruction statement by adding a new record, or modify the contents of an existing record.

[0048] 5 is a flowchart showing the flow of a generation process for generating a summary corresponding to speech data using a generation model 31. The generation process is performed by the CPU 21 reading the information processing program 30 from the storage 24, expanding it in the RAM 23, and executing it. As an example, the generation process is started when a user logs in to the summary service and performs a predetermined operation to start the generation process. The following explanation assumes that the user is a "general user."

[0049] In step S10 shown in FIG. 5, the CPU 21 acquires voice data indicating the user's speech content, which the general user uploaded to the summarization service via the general user terminal 60. Here, the speech content is not limited to a presentation or explanation by a single user, but also includes dialogue or conversation between multiple users. Therefore, the voice data can cover speech content in a variety of situations, such as presentations, meetings, interviews, and personal meetings. The voice data may be data recorded in real time using a microphone built into or connected to the general user terminal 60, or may be a voice file recorded in advance by an external recording device that is imported into the general user terminal 60 and then uploaded to the summarization service. The CPU 21 then proceeds to step S11.

[0050] In step S11, the CPU 21 generates a first prompt including the voice data acquired in step S10 and a predetermined transcription instruction, and then the CPU 21 proceeds to step S12.

[0051] Here, the "transcription instruction" is, for example, an instruction such as "Please convert the content of this audio into text as accurately and chronologically as possible," and is guide information that indicates the content, format, perspective, etc. of the output required of the first generative model 31A. As a result, the CPU 21 generates, for example, a first prompt such as the following:

[0052] <First prompt> "Audio data" File name: meeting_20240501.wav Content: Recorded audio of a sales meeting attended by a general user "Instruction text" Please transcribe the entire contents of this audio data. If there are changes in speaker, please distinguish between each speaker. Please write in a natural style that maintains the chronological order as much as possible, and adjust colloquial expressions as appropriate.

[0053] In step S12, CPU 21 inputs the first prompt generated in step S11 to first generative model 31A, and obtains text data corresponding to the voice data as an output result of first generative model 31A in response to the first prompt. Then, CPU 21 proceeds to step S13.

[0054] In step S13, the CPU 21 generates a second prompt including the text data acquired in step S12, the user information of the general user who executed the generation process, and a predetermined search instruction, and then proceeds to step S14.

[0055] Here, the "search instruction" is, for example, an instruction such as "Please identify the speech scene in which this utterance content occurred and select a corresponding instruction sentence corresponding to the speech scene from the database 32," and is guide information indicating the content, format, viewpoint, etc. of the output required of the second generation model 31B. Furthermore, the second generation model 31B selects a corresponding instruction sentence by searching a group of summary instruction sentences associated with a general user based on user information, specifically, a group of summary instruction sentences registered by an administrative user who oversees the general user. In this way, the CPU 21 generates, for example, a second prompt such as the following:

[0056] <Second prompt> "Text data" Thank you for your time today. First, I'd like to ask about the challenges your company faces. "User information" User ID:user_1023 Organization: Sales Group 2 "Search Instructions" Please select one of the following speech scenes in which the above utterance content occurred: "Sales Negotiation," "Internal Meeting," "Interview," or "Recruitment Interview." Also, depending on the identified speech scene, please select a corresponding instruction sentence from Database 32 based on the above user information.

[0057] In step S14, the CPU 21 inputs the second prompt generated in step S13 to the second generation model 31B, and acquires a corresponding instruction sentence associated with the utterance scene as an output result of the second generation model 31B in response to the second prompt. Then, the CPU 21 proceeds to step S15.

[0058] In step S15, the CPU 21 generates a third prompt including the voice data acquired in step S10, the speech scene identified in step S14, and an instruction to extract summary information related to the predetermined first viewpoint, and then the CPU 21 proceeds to step S16.

[0059] Here, the "summary information extraction instruction related to the first viewpoint" is, for example, an instruction such as "Please extract the other person's concerns from the content of this utterance," and is guide information indicating the content, format, viewpoint, etc. of the output required of the third generative model 31C. As a result, the CPU 21 generates, for example, a third prompt such as the following:

[0060] <Third prompt> "Audio data" File name: meeting_20240501.wav Content: Recorded audio of a sales meeting attended by a general user "Speech scene" Sales negotiations "First perspective (extraction instructions)" From this audio, please extract the other party's concerns (cost, delivery time, specifications, etc.). Please briefly summarize each concern and quote what the speaker said when they expressed their concerns.

[0061] In step S16, CPU 21 inputs the third prompt generated in step S15 into the third generative model 31C and obtains summary information related to the first perspective as an output result of the third generative model 31C in response to the third prompt. Here, "summary information related to the first perspective" refers to text information that extracts, for example, information indicating the other party's concerns (such as concerns about costs, concerns about delivery dates, or specification requirements) and classifies and organizes them by concern. If necessary, the summary information may also include a portion of the typical utterance content when the concern is expressed. Then, CPU 21 proceeds to step S17.

[0062] In step S17, CPU 21 generates a fourth prompt including the voice data acquired in step S10, the speech scene identified in step S14, the summary information related to the first perspective acquired in step S16, and an instruction to extract summary information related to a predetermined second perspective. The second perspective is a perspective different from the first perspective, such as "extracting the other person's interests" or "confirming the next action." Then, CPU 21 proceeds to step S18.

[0063] Here, the "summary information extraction instruction related to the second perspective" is, for example, an instruction such as "Please extract the item that the other party showed the most interest in from the content of this utterance," and is guide information indicating the content, format, perspective, etc. of the output required of the fourth generation model 31D. As a result, the CPU 21 generates, for example, the following fourth prompt.

[0064] <Fourth prompt> "Audio data" File name: meeting_20240501.wav Content: Recorded audio of a sales meeting attended by a general user "Speech scene" Sales negotiations "Summary information relating to the first aspect" Explicit concerns about costs (e.g., "I'm worried about staying within budget") Concerns about delivery dates (e.g., "It's going to be tough if it's not within three weeks.") "Second Perspective (Extraction Instructions)" From this conversation, please extract the items that the other party was particularly interested in (e.g., solution content, implementation history, comparison with other companies). If possible, please summarize the specific content of the conversation.

[0065] In step S18, CPU 21 inputs the fourth prompt generated in step S17 into the fourth generative model 31D and obtains summary information related to the second perspective as the output result of the fourth generative model 31D in response to the fourth prompt. Here, "summary information related to the second perspective" refers to, for example, extracted items in which the other party showed particular interest (e.g., product features, price conditions, comparison with other companies), and text information summarized for each item of interest. If necessary, quotations of specific statements expressing interest may also be included. CPU 21 then proceeds to step S19.

[0066] In step S19, CPU 21 generates a fifth prompt including the voice data acquired in step S10, the speech scene identified in step S14 and the corresponding instruction sentence selected, the summary information related to the first perspective acquired in step S16, and the summary information related to the second perspective acquired in step S18. Then, CPU 21 proceeds to step S20.

[0067] Here, the "response instruction" is a summary instruction associated with the sales negotiation, such as "Please briefly summarize this entire conversation, summarizing the customer's main concerns and interests in one sentence each. In addition, please clearly indicate any comments that may affect sales decisions, such as price, delivery date, and implementation timing, or any content related to comparisons with competing products." Based on this, the CPU 21 generates a fifth prompt, for example, as follows:

[0068] <Fifth prompt> "Audio data" File name: meeting_20240501.wav Content: Recorded audio of a sales meeting attended by a general user "Speech scene" Sales negotiations "Summary information relating to the first aspect" Explicit concerns about costs (e.g., "I'm worried about staying within budget") Concerns about delivery dates (e.g., "It's going to be tough if it's not within three weeks.") "Summary information relating to the second viewpoint" Customers were very interested in "product adoption cases" (e.g., "Have you seen any success stories at other companies?") Questions also focused on "after-sales support" (e.g., "What kind of follow-up system is in place after implementation?") "Response Instructions" Summarize the entire conversation. Include a brief, one-sentence summary of the customer's main concerns and areas of interest. Also, if the conversation included any sales decisions, such as price, delivery time, or implementation schedule, be sure to highlight those points. Also, include any comparisons or evaluations of competing products.

[0069] In step S20, CPU 21 inputs the fifth prompt generated in step S19 into fifth generation model 31E and obtains a summary of the utterance content as an output result of fifth generation model 31E in response to the fifth prompt. The "summary result" here is a summary generated in an integrated manner based on multiple pieces of summary information, namely, the first perspective obtained in step S16 and the second perspective obtained in step S18. Therefore, the summary result clearly reflects, for example, the other party's concerns extracted from the first perspective and the interests extracted from the second perspective. Then, CPU 21 proceeds to step S21.

[0070] In step S21, the CPU 21 outputs the summary result acquired in step S20 to the general user terminal 60. In the general user terminal 60, the CPU 61 causes the display unit 66 to display the summary result. The CPU 21 also generates result information that associates the summary result with at least the fifth prompt used to generate the summary result, and stores this in the database 32. The CPU 21 then terminates the generation process.

[0071] 6 is a flowchart showing the flow of an adjustment process for adjusting a prompt, specifically a fifth prompt, used in generating a summary result using the generative model 31. The adjustment process is performed by the CPU 21 reading the information processing program 30 from the storage 24, expanding it into the RAM 23, and executing it. As an example, the adjustment process is started when the generation process is completed.

[0072] In step S30 shown in Fig. 6, the CPU 21 acquires evaluation information indicating a user's evaluation of the summary result of the utterance content output in the generation process. For example, when a general user selects the GOOD button 73 or the BAD button 74 (see Fig. 7 for both) displayed on the display unit 66 of the general user terminal 60, the CPU 21 can acquire evaluation information indicating a positive or negative evaluation. In this way, the evaluation information is information indicating what kind of evaluation the user made of the generated summary result, and is configured to clearly indicate whether it was positive or negative. Then, the CPU 21 proceeds to step S31.

[0073] In step S31, the CPU 21 determines whether the evaluation information acquired in step S30 indicates a negative evaluation. If the CPU 21 determines that the evaluation information indicates a negative evaluation (step S31: YES), the process proceeds to step S32. On the other hand, if the CPU 21 determines that the evaluation information indicates a positive evaluation (step S31: NO), the process proceeds to step S34. As an example, the CPU 21 determines that a positive evaluation is indicated when the GOOD button 73 is operated, and determines that a negative evaluation is indicated when the BAD button 74 is operated.

[0074] In step S32, the CPU 21 receives input of feedback information from the user indicating the reason for the negative evaluation. For example, a general user inputs text into an input field 81 displayed on the display unit 66 of the general user terminal 60 and then operates the complete button 82 (see FIG. 8 for both). This allows the CPU 21 to obtain feedback information including the content of the text. In this way, the feedback information is information indicating the reason for the negative evaluation of the generated summary result, and includes specific requests for improvement or problems, such as "the summary is inaccurate," "the main point of the story is missing," or "the points of interest are not reflected." The CPU 21 then proceeds to step S33.

[0075] In step S33, CPU 21 adjusts the fifth prompt used to generate the negatively evaluated summary result based on the feedback information acquired in step S32. For example, CPU 21 can update the content of the summarization instruction for the fifth prompt stored in database 32 by extracting keywords included in the feedback content, categorizing issues using natural language processing, or comparing the content with a preset correction template. As a result, for example, if the original summarization instruction was "Please summarize this conversation briefly," and the feedback content is "The customer's request is missing," the summarization instruction can be adjusted to "Please summarize this conversation briefly and clearly state the customer's request." Then, CPU 21 proceeds to step S34.

[0076] In step S34, the CPU 21 updates the result information generated in the generation process based on the processing results of the adjustment process. For example, if the evaluation information acquired in step S30 indicates a positive evaluation, the CPU 21 adds evaluation information indicating a positive evaluation to the result information corresponding to the summary result that received the positive evaluation and records it. On the other hand, if the evaluation information indicates a negative evaluation, the CPU 21 adds evaluation information indicating a negative evaluation to the result information corresponding to the summary result that received the negative evaluation, and may also associate and store the fifth prompt before adjustment with the fifth prompt after adjustment. This enables output control that references the evaluation trend in subsequent processing and also enables the prompt update history to be used in the future. The CPU 21 then terminates the adjustment process.

[0077] FIG. 7 is a diagram showing an example of a display screen of the summary result displayed on the display unit 66 of the general user terminal 60. As shown in FIG.

[0078] The display screen shown in FIG. 7 displays target text information 70, summary prompt information 71, summary result information 72, a GOOD button 73, and a BAD button 74.

[0079] The target text information 70 is a portion that displays text data corresponding to the voice data, that is, text that is a transcription of the spoken content shown in the voice data.

[0080] Summary prompt information 71 is a section that displays the fifth prompt used to generate the summary result.

[0081] The summary result information 72 is a section that displays, in text, the summary result generated based on the fifth prompt indicated in the summary prompt information 71. For ease of explanation, illustrations of the specific text displayed in the target text information 70, the summary prompt information 71, and the summary result information 72 are omitted.

[0082] The GOOD button 73 is a GUI (Graphical User Interface) button that allows a general user to input a positive evaluation of the summary result shown in the summary result information 72.

[0083] The BAD button 74 is a GUI button that allows a general user to input a negative evaluation of the summary result shown in the summary result information 72. When the BAD button 74 is selected by a general user, a feedback input screen shown in Fig. 8 is displayed, and the reason for the negative evaluation is accepted.

[0084] FIG. 8 is a diagram showing an example of a feedback input screen displayed on the display unit 66 of the general user terminal 60. As shown in FIG.

[0085] The feedback input screen shown in FIG. 8 displays message information 80, an input field 81, and a complete button 82.

[0086] The message information 80 is a section that displays a guidance message that prompts general users to input information. In Fig. 8, the message information 80 displays a guidance message saying "Please enter the reason for your negative evaluation," and guides general users to input feedback information.

[0087] The input field 81 is a text input area where a general user can freely enter the reason for a negative evaluation via the input unit 65. The general user can write specific complaints or requests for improvement regarding the summary result.

[0088] When a general user enters the reason in the input field 81 and presses the complete button 82, feedback information is sent to the summary service. In the adjustment process shown in FIG. 6, the CPU 21 of the server 20 adjusts the fifth prompt used to generate the negatively evaluated summary result based on the feedback information.

[0089] As described above, in the summary generation system 10, the CPU 21 inputs a fifth prompt to the fifth generation model 31E, the fifth prompt including voice data indicating the content of a user's utterance and a corresponding instruction sentence corresponding to the utterance scene identified based on the voice data from among multiple summarization instruction sentences registered in association with multiple scenes. The CPU 21 then outputs a summary result of the utterance content acquired as an output result of the fifth generation model 31E. As a result, the summary generation system 10 generates a summary result of the utterance content based on an appropriate summarization instruction corresponding to the user's utterance scene, even when different perspectives are required for each utterance scene, such as a sales negotiation or a job interview. Therefore, the summary generation system 10 can provide a summary result that reflects the perspective required for the user's utterance scene.

[0090] Furthermore, in the summary generation system 10, when acquiring a summary result of the utterance content, the CPU 21 causes the third generation model 31C to generate summary information related to a first perspective and the fourth generation model 31D to generate summary information related to a second perspective based on the audio data. Then, the CPU 21 causes the fifth generation model 31E to summarize the utterance content based on multiple pieces of summary information including the summary information related to the first and second perspectives. As a result, the summary generation system 10 can summarize the utterance content by comprehensively considering multiple pieces of summary information related to different perspectives, thereby providing a highly practical summary result that reflects more of the perspectives required depending on the utterance scene.

[0091] Furthermore, in summary generation system 10, when the evaluation information for the output summary result of the utterance content indicates a negative evaluation, CPU 21 accepts input of feedback information from the user indicating the reason for the negative evaluation. Then, based on the feedback information, CPU 21 adjusts the fifth prompt used to generate the summary result that received the negative evaluation. As a result, summary generation system 10 adaptively adjusts the fifth prompt in consideration of the specific feedback information indicating the reason for the negative evaluation, thereby improving generation accuracy so as to obtain a summary result that more appropriately reflects the perspective required for the utterance scene.

[0092] (Second embodiment) Next, a second embodiment of the summary generation system 10 will be described while omitting or simplifying parts that overlap with the above embodiment.

[0093] The summary generation system 10 according to the second embodiment differs from the first embodiment in that it provides multiple fifth generation models 31E (e.g., fifth generation models 31E-1, 31E-2, etc.) with different learning characteristics, and inputs a fifth prompt to each fifth generation model 31E to generate multiple summary results of the spoken content.

[0094] Here, "multiple fifth generative models 31E with different learning characteristics" refers to, for example, models trained at different learning rates using the same training data, or generative models optimized for different types of summarization perspectives (e.g., fact-focused type, key point extraction type, natural tone type). As a result, the summarization results output from each fifth generative model 31E differ in expression style, tendency to extract information, level of abstraction of description, etc.

[0095] 5, the CPU 21 inputs the fifth prompt generated in step S19 into a plurality of fifth generation models 31E to generate a plurality of summary results of the utterance content. Then, in step S21, the CPU 21 outputs the plurality of summary results to the general user terminal 60 and displays the plurality of summary results on the display unit 66. Then, the CPU 21 accepts a user's selection of a summary result to be output from among the plurality of summary results.

[0096] FIG. 9 is a diagram showing an example of a display screen of a plurality of summary results displayed on the display unit 66 of the general user terminal 60. As shown in FIG.

[0097] The display screen shown in Fig. 9 displays target text information 70, summary prompt information 71, summary result information 90, summary result information 91, selection button 92, and selection button 93. Note that the target text information 70 and summary prompt information 71 have the same configuration as in Fig. 7, and therefore their explanation will be omitted.

[0098] The summary result information 90 is a portion that displays the summary result A generated by inputting the fifth prompt to the fifth generative model 31E-1.

[0099] The summary result information 91 is a part that displays the summary result B that is generated by inputting a fifth prompt into the fifth generation model 31E-2, which has learning characteristics different from those of the fifth generation model 31E-1. For ease of explanation, illustrations of the specific text displayed in the summary result information 90 and the summary result information 91 are omitted.

[0100] The selection button 92 is a GUI button for selecting summary result A as the output target from among the multiple displayed summary results.

[0101] The selection button 93 is a GUI button for selecting summary result B as the output target from among the multiple displayed summary results.

[0102] A general user can specify any summary result as the output target by operating any of the selection buttons. In the second embodiment, the summary result selected as the output target becomes the processing target of the adjustment process shown in FIG. 6. For example, when selection button 92 is operated to select summary result A as the output target, a pop-up screen for inputting an evaluation, including a GOOD button 73 and a BAD button 74, is displayed to acquire evaluation information for summary result A. Thereafter, in accordance with the acquired evaluation information, CPU 21 sequentially executes the processes corresponding to steps S31 to S34 of the adjustment process shown in FIG. 6.

[0103] As described above, in the summary generation system 10, the CPU 21 inputs a fifth prompt to each of multiple fifth generation models 31E having different learning characteristics, causing multiple summary results of the utterance content to be generated. The CPU 21 then accepts a user's selection of a summary result to be output from among the multiple generated summary results. As a result, the summary generation system 10 allows the user to select from among the summary results obtained by the multiple fifth generation models 31E, thereby enabling the user to actively obtain a summary result that best suits their intentions depending on the speech scene.

[0104] (others) Although the embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, the technical scope of the present disclosure is not limited to such examples. It is clear that a person skilled in the art of the present disclosure can conceive of various modifications or alterations within the scope of the technical idea described in the claims, and it is understood that these modifications or alterations also naturally fall within the technical scope of the present disclosure.

[0105] Furthermore, the effects described in the above embodiments are explanatory or exemplary and are not limited to those described in the above embodiments. In other words, the technology according to the present disclosure may achieve other effects that are obvious to a person skilled in the art of the present disclosure from the description in the above embodiments, in addition to or instead of the effects described in the above embodiments.

[0106] The processes described in the above embodiments can also be realized by dedicated hardware circuits, in which case they may be executed by a single piece of hardware or by multiple pieces of hardware.

[0107] In the above embodiment, the fifth generation model 31E summarizes the utterance content based on a plurality of pieces of summary information including summary information relating to the first perspective and the second perspective, but this is not limitative. For example, the plurality of pieces of summary information may include summary information relating to three or more types of perspectives, not limited to the first perspective and the second perspective.

[0108] In the above embodiment, digressive or self-referential content included in an utterance may be excluded from the summary as an exclusion target. Specifically, the CPU 21 detects an utterance including a specific phrase such as "I'm digressing, but..." as an exclusion target. This detection may be performed by referring to a dictionary of terms to be excluded stored in the database 32 of the storage 24. The CPU 21 then excludes the utterance determined to be an exclusion target from the summary.

[0109] In the above embodiment, the degree of emphasis of the summary may be controlled according to the importance of the utterance. Specifically, the CPU 21 may combine intonation information indicating the intonation of the utterance content based on the voice data with external sensor information indicating the user's gaze information, etc., to score the importance of each utterance. The CPU 21 may then add summary information corresponding to the utterance content whose score is equal to or greater than a predetermined value to the fifth prompt.

[0110] In the above embodiment, the summary result may be configured to reflect the audience's reaction to the utterance. Specifically, the CPU 21 may analyze nodding sounds, backchannels, and facial expression changes acquired from a microphone, camera, or the like provided in the input unit of each user terminal, and calculate a reaction intensity score based on the reaction. The CPU 21 may then add summary information corresponding to the utterance content whose reaction intensity score is equal to or greater than a predetermined value to the fifth prompt.

[0111] In the above embodiment, a configuration may be adopted in which standardized speech portions such as greetings are automatically abbreviated in the summary result. Specifically, the CPU 21 may compare the standardized speech portions included in the speech content with a dictionary of standard expressions stored in the database 32 of the storage 24, replace the standardized speech portions with abbreviations such as "opening greetings," and include the replacement results in the summary result.

[0112] In the above embodiment, each fifth prompt may be individually optimized based on the past feedback tendency of each user. Specifically, CPU 21 may chronologically analyze the prompt information and result information stored in storage 24 to identify a group of users with similar feedback tendencies. Then, CPU 21 may adjust the fifth prompt available to the group of users based on the feedback tendency of the group of users.

[0113] In the above embodiment, a configuration may be adopted in which a summary instruction sentence for a scene corresponding to an upcoming schedule is generated based on schedule information indicating the schedules of multiple general users managed by an administrative user. Specifically, the CPU 21 may acquire the schedule information indicating the schedules of the multiple general users from the corresponding general user terminals 60. The CPU 21 may then identify a scene that will be required in the future based on the schedule information, generate a summary instruction sentence corresponding to the scene, and transmit the summary instruction sentence to the administrative user terminal 40. This reduces the burden on the administrative user in registering summary instruction sentences.

[0114] In the above embodiment, as part of the summarization process by the third generative model 31C, the fourth generative model 31D, and the fifth generative model 31E, an extraction process is executed to extract a first perspective (e.g., the other party's concerns) and a second perspective (e.g., the other party's interests) corresponding to a speech scene from the speech content indicated in the audio data. However, the extraction process is not limited to being executed as part of the summarization process, and can also be executed as an independent processing function separate from the summarization process. In this case, for example, the results of the extraction process based on the speech content in a sales negotiation (e.g., "budget," "desired implementation date," and "department to which the decision maker belongs") are stored in the storage 24 as structured data (e.g., JSON format) and can be used for various purposes. Specifically, the results of the extraction process may be used for summarization as in the above embodiment, or may be linked to an external business system such as a CRM (Customer Relationship Management) system and used as data for customer analysis and sales support.

[0115] In the present disclosure, the term "information processing system" is a concept that encompasses both a system configured with a single device and a system configured with a combination of multiple devices. For example, the information processing system of the present disclosure may be configured with a single server 20, or may be configured with a combination of the server 20 and an administrative user terminal 40, or the server 20 and a general user terminal 60, etc.

[0116] In this disclosure, the term "processor" refers to a processor in a broad sense, including general-purpose processors (e.g., CPU: Central Processing Unit, etc.) and dedicated processors (e.g., GPU: Graphics Processing Unit, ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array, programmable logic device, etc.).

[0117] The operations of the processor in the present disclosure may be performed not only by a single processor but also by multiple processors located in physically separate locations working together. Furthermore, the order of the operations of the processor is not limited to the order described in the present disclosure and may be changed as appropriate.

[0118] In the above embodiment, the information processing program 30 is pre-stored (installed) in the storage 24, but the present invention is not limited to this. The information processing program 30 may be provided in a form recorded on a recording medium such as a CD-ROM (Compact Disk Read Only Memory), a DVD-ROM (Digital Versatile Disk Read Only Memory), or a USB (Universal Serial Bus) memory. The information processing program 30 may also be downloaded from an external device via a network N. The technology disclosed herein may also be applied to programs and program products.

[0119] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference. [Explanation of symbols]

[0120] 21 CPU (processor) 30 Information Processing Program 31 Generative Model< / url:>

Claims

1. a processor; The processor: Acquire voice data indicating the content of the user's speech; generating a first prompt including the voice data and a predetermined transcription instruction, and inputting the first prompt into a first generative model; and obtaining text data corresponding to the voice data as an output result of the first generative model; generating a second prompt including the text data, user information of a user who executed a generation process for generating a summary result corresponding to the speech data, and a predetermined search instruction, and inputting the second prompt into a second generative model; and obtaining, as an output result of the second generative model, an utterance scene corresponding to the speech data and a corresponding instruction sentence associated with the utterance scene; generating a third prompt including the speech data, the identified speech scene, and an instruction to extract summary information related to a predetermined first perspective, and inputting the third prompt to a third generative model; and obtaining summary information related to the first perspective as an output result of the third generative model; generating a fourth prompt including the speech data, the identified speech scene, summary information related to the first perspective, and an instruction to extract summary information related to a predetermined second perspective, and inputting the fourth prompt to a fourth generative model; and obtaining the summary information related to the second perspective as an output result of the fourth generative model; generating a fifth prompt including the speech data, the identified speech scene, the corresponding instruction sentence, and summary information related to the first perspective and the second perspective, and inputting the fifth prompt into a fifth generative model; and outputting a summary result of the speech content obtained as an output result of the fifth generative model. Information processing system.

2. The processor: acquiring evaluation information indicating a user's evaluation of the output summary of the utterance content; If the evaluation information indicates a negative evaluation, receiving feedback information from the user indicating the reason for the negative evaluation; adjusting the fifth prompt used to generate the negatively rated summary result based on the feedback information; The information processing system according to claim 1 .

3. The processor: inputting the fifth prompt into a plurality of fifth generative models each having a different learning characteristic to generate a plurality of summary results of the utterance content; accepting a user's selection of a summary result to be output from among the generated plurality of summary results; The information processing system according to claim 1 .

4. Acquire voice data indicating the content of the user's speech; generating a first prompt including the voice data and a predetermined transcription instruction, and inputting the first prompt into a first generative model; and obtaining text data corresponding to the voice data as an output result of the first generative model; generating a second prompt including the text data, user information of a user who executed a generation process for generating a summary result corresponding to the speech data, and a predetermined search instruction, and inputting the second prompt into a second generative model; and obtaining, as an output result of the second generative model, an utterance scene corresponding to the speech data and a corresponding instruction sentence associated with the utterance scene; generating a third prompt including the speech data, the identified speech scene, and an instruction to extract summary information related to a predetermined first perspective, and inputting the third prompt to a third generative model; and obtaining summary information related to the first perspective as an output result of the third generative model; generating a fourth prompt including the speech data, the identified speech scene, summary information related to the first perspective, and an instruction to extract summary information related to a predetermined second perspective, and inputting the fourth prompt to a fourth generative model; and obtaining the summary information related to the second perspective as an output result of the fourth generative model; generating a fifth prompt including the speech data, the identified speech scene, the corresponding instruction sentence, and summary information related to the first perspective and the second perspective, and inputting the fifth prompt into a fifth generative model; and outputting a summary result of the speech content obtained as an output result of the fifth generative model. An information processing method in which processing is performed by a computer.

5. Acquire voice data indicating the content of the user's speech; generating a first prompt including the voice data and a predetermined transcription instruction, and inputting the first prompt into a first generative model; and obtaining text data corresponding to the voice data as an output result of the first generative model; generating a second prompt including the text data, user information of a user who executed a generation process for generating a summary result corresponding to the speech data, and a predetermined search instruction, and inputting the second prompt into a second generative model; and obtaining, as an output result of the second generative model, an utterance scene corresponding to the speech data and a corresponding instruction sentence associated with the utterance scene; generating a third prompt including the speech data, the identified speech scene, and an instruction to extract summary information related to a predetermined first perspective, and inputting the third prompt to a third generative model; and obtaining summary information related to the first perspective as an output result of the third generative model; generating a fourth prompt including the speech data, the identified speech scene, summary information related to the first perspective, and an instruction to extract summary information related to a predetermined second perspective, and inputting the fourth prompt to a fourth generative model; and obtaining the summary information related to the second perspective as an output result of the fourth generative model; generating a fifth prompt including the speech data, the identified speech scene, the corresponding instruction sentence, and summary information related to the first perspective and the second perspective, and inputting the fifth prompt into a fifth generative model; and outputting a summary result of the speech content obtained as an output result of the fifth generative model. An information processing program that causes a computer to execute a process.

Citation Information

Patent Citations

  • Processing device, processing method, and processing program

    JP7656741B1

  • Minutes creation support device and program

    JP7681360B1

  • Summary generation method, summary generation system, and computer program

    JP2024167144A

  • JPP7656741B

  • JPP7681360B