Information processing system and information processing method

The system addresses noise and inaccuracy in speech recognition by creating role-specific and user-adapted summaries from meeting audio, ensuring accurate and relevant meeting minutes.

JP2026090834APending Publication Date: 2026-06-03SEMICON ENERGY LAB CO LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SEMICON ENERGY LAB CO LTD
Filing Date
2024-11-22
Publication Date
2026-06-03

Smart Images

  • Figure 2026090834000001_ABST
    Figure 2026090834000001_ABST
Patent Text Reader

Abstract

It helps create optimal meeting minutes that take into account the roles of the viewers. [Solution] An information processing system and information processing method having a function to perform processing using a speech recognition model, a function to perform processing using a language model, a function to receive speech data, perform processing using a speech recognition model and output a conversation record, a function to proofread the conversation record using a language model, and a function to output a user-adjusted summary using a language model and user information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One aspect of the present invention relates to an information processing system and an information processing method using a speech recognition model and a language model, particularly a generative artificial intelligence (AI) model.

[0002] Note that one aspect of the present invention is not limited to the above technical field. Examples of the technical field of one aspect of the present invention disclosed in this specification and the like include semiconductor devices, display devices, light-emitting devices, power storage devices, storage devices, electronic devices, lighting devices, input devices, input / output devices, their driving methods, or their manufacturing methods. The semiconductor device refers to all devices that can function by utilizing semiconductor characteristics.

Background Art

[0003] In recent years, the development of language models using neural networks has been actively carried out, and particularly large language models (LLMs) have attracted attention. A large language model is a natural language processing model learned using a large amount of data. With a large language model, for example, a dialogue model that answers user instructions can be realized. In Non-Patent Document 1, GPT-4 (Generative Pre-trained Transformer 4) (registered trademark) is disclosed as a large language model, and ChatGPT is disclosed as a dialogue model.

[0004] Conventionally, it is known that text data can be extracted from audio data by speech recognition technology to perform speech-to-text conversion.

Prior Art Documents

Non-Patent Documents

[0005]

Non-Patent Document 1

[0006] When attempting to obtain meeting minutes from a meeting using speech recognition, the resulting text data may contain noise and inaccurate information due to the recording environment. Furthermore, meeting discussions often include unnecessary remarks such as interjections, resulting in redundant text data. Additionally, the performance of the speech recognition model may lead to errors, unnecessary repetitions, and other issues in the resulting text data.

[0007] Furthermore, meeting minutes are required to be written in a way that is appropriate to the background and level of understanding of the reader.

[0008] One aspect of the present invention, in view of the above-mentioned problems, aims to support the creation of optimal meeting minutes that take into account the viewer's role (role within a group such as a company). Alternatively, it aims to support the creation of optimal notes or other documents that take into account the viewer's role.

[0009] Furthermore, the description of these problems does not preclude the existence of other problems. Moreover, one aspect of the present invention does not need to solve all of these problems. Other problems can be identified from the description in the specification, drawings, claims, etc. [Means for solving the problem]

[0010] In view of the above problems, one aspect of the present invention comprises a first information processing device, a second information processing device, and a third information processing device, wherein the first information processing device has the function of receiving an instruction sentence, processing it using a language model, and outputting a first response sentence, a second response sentence, a third response sentence, and a fourth response sentence; the second information processing device has the function of receiving audio data, processing it using a speech recognition model, and outputting a conversation record and speaker labeling data; the third information processing device has the function of receiving audio data, transmitting the audio data to the second information processing device and receiving the conversation record and speaker labeling data, creating a first instruction sentence from the conversation record, speaker labeling data, and conversation record proofreading instructions, transmitting the first instruction sentence to the first information processing device and receiving a first response sentence including a proofread conversation record, receiving a role list, and obtaining a role from the role list. This information processing system has the following functions: a function to create a second instruction from a proofread conversation transcript, roles, and summary instructions; a function to send the second instruction to a first information processing device and receive a second response including a summary for each role; a function to create a summary list; a function to add the summaries for each role to the summary list; a function to receive user information; a function to create a third instruction from a proofread conversation transcript, summary list, role list, user information, and role weight creation instructions; a function to send the third instruction to a first information processing device and receive a third response including a first role weight list; a function to create a fourth instruction from a proofread conversation transcript, role list, first role weight list, summary list, and summary modification instructions; a function to send the fourth instruction to a first information processing device and receive a fourth response including a user-adapted summary; and a function to output a user-adapted summary.

[0011] In the above-described information processing system, it is preferable that the third information processing device has the function of outputting a role list, a summary list, and a first role weight list, the function of receiving a second role weight list, and the function of creating a fourth instruction sentence from a proofread conversation record, a role list, a second role weight list, a summary list, and summary correction instructions.

[0012] Furthermore, one aspect of the present invention comprises steps 1 to 16, in the first step, voice data is acquired; in the second step, the voice data is input into a speech recognition model to acquire conversation transcripts and speaker labeling data; in the third step, a first instruction sentence is created from the conversation transcripts, speaker labeling data, and conversation transcript correction instructions; in the fourth step, the first instruction sentence is input into a language model to acquire a first response sentence including a corrected conversation transcript; in the fifth step, a role list is acquired; in the sixth step, roles are acquired from the role list; in the seventh step, a second instruction sentence is created from the corrected conversation transcripts, roles, and summary instructions; in the eighth step, the second instruction sentence is input into a language model to acquire a second response sentence including a summary for each role, the summary for each role is added to a summary list; and the ninth step This information processing method determines whether the entire role list has been processed in each step; if not, it returns to step 6; if processed, it proceeds to step 10, where user information is obtained in step 10; in step 11, a third instruction sentence is created from the corrected conversation transcript, summary list, role list, user information, and role weight creation instructions; in step 12, the third instruction sentence is input into the language model to obtain a third response sentence including the first role weight list; in step 13, a fourth instruction sentence is created from the corrected conversation transcript, role list, first role weight list, summary list, and summary correction instructions; in step 14, the fourth instruction sentence is input into the language model to obtain a fourth response sentence including a user-adjusted summary; and in step 15, a user-adjusted summary is output.

[0013] In the above information processing system, it is preferable that the system has steps 16 to 18 after step 12, in step 16 output a role list, a summary list, and a first role weight list, in step 17 obtain a second role weight list, and in step 18 create a fourth instruction sentence from the corrected conversation transcript, role list, second role weight list, summary list, and summary correction instructions, and then proceed to step 14. [Effects of the Invention]

[0014] According to one aspect of the present invention, it is possible to provide meeting minutes tailored to the viewer's personal information, level of understanding, etc., from audio data of a meeting. Alternatively, it is possible to provide text such as notes tailored to the viewer's personal information, level of understanding, etc., from audio data of a seminar, lecture, etc. [Brief explanation of the drawing]

[0015] [Figure 1] Figure 1 is a schematic diagram illustrating an example of the configuration of an information processing system. [Figure 2] Figure 2 is a block diagram showing an example of the configuration of an information processing system. [Figure 3] Figure 3 is a flowchart showing an example of the processing in an information processing system. [Figure 4] Figure 4 is a flowchart showing an example of the processing in an information processing system. [Figure 5] Figure 5 is a flowchart showing an example of the processing in an information processing system. [Figure 6] Figures 6(A) and 6(B) illustrate the structure of instruction statements used in an information processing system. [Figure 7] Figure 7 illustrates the structure of instruction statements used in an information processing system. [Figure 8] Figure 8 is a diagram illustrating the structure of instruction statements used in an information processing system. [Figure 9] Figure 9 is a flowchart showing an example of the processing in an information processing system. [Modes for carrying out the invention]

[0016] Hereinafter, embodiments will be described with reference to the drawings. However, the embodiments can be implemented in many different ways, and it is easily understood by those skilled in the art that the form and details can be variously changed without departing from the spirit and its scope. Therefore, the present invention is not construed as being limited to the description of the following embodiments.

[0017] In the configuration of the invention described below, the same reference numerals are commonly used for the same parts or parts having the same functions among different drawings, and the repeated description thereof is omitted. Also, when referring to the same function, there may be no particular reference numeral attached.

[0018] In each of the figures described in this specification, the size of each component, the thickness of the layer, or the area may be exaggerated for clarity. Therefore, it is not necessarily limited to that scale.

[0019] The ordinal numbers such as "first", "second", etc. in this specification and the like are attached to avoid confusion of components, and are not numerically limiting. They do not indicate any order or rank such as process order or stacking order. Even for terms without ordinal numbers in this specification and the like, ordinal numbers may be attached in the claims to avoid confusion of components. Even for terms with ordinal numbers in this specification and the like, different ordinal numbers may be attached in the claims. Even for terms with ordinal numbers in this specification and the like, ordinal numbers may be omitted in the claims.

[0020] In this specification and the like, the language model is based on the Transformer architecture and is additionally trained to be a dialogue model. Also, typically, there is a large language model (LLM) as the language model. The LLM is specialized in the text generation function that processes based on the given text data, and the generative AI has not only the text generation function but also an image generation function that processes based on image data. That is, the large language model is one of the generative AIs.

[0021] (Embodiment 1) In this embodiment, an example of the configuration of a text generation system, which is one aspect of the present invention, will be described with reference to Figure 1.

[0022] One embodiment of the present invention is an information processing system that outputs a conversation transcript obtained by converting audio data recorded at meetings, etc., into text data. Furthermore, by utilizing a language model, it becomes possible to remove unnecessary text information, proofread the text, summarize it, and adapt it to the individual.

[0023] Below, we will explain more specific examples with reference to the diagrams.

[0024] <Example of an information processing system configuration 1> The information processing system of this embodiment preferably has a configuration as shown in Figure 1, comprising a first information processing device 10, a second information processing device 20, a third information processing device 30, and an information terminal 40. The third information processing device 30 is connected to the first information processing device 10 and the second information processing device 20 via a network 50. Furthermore, the third information processing device 30 is connected to the information terminal 40 via a network 60.

[0025] Furthermore, in the information processing system of this embodiment, the system configuration shown in Figure 1 is just one example, and the third information processing device 30 may have the functions of at least one of the first information processing device and the second information processing device, or the functions of the third information processing device 30 may be distributed and implemented across multiple information processing devices.

[0026] 《Example configuration of information terminal 40》 In the example configuration of the information processing system 1, the information terminal 40 is operated by the user and can also be called a client computer. In Figure 1, a desktop computer and a smartphone are shown as examples, but a notebook computer or a tablet computer may also be used as the information terminal 40. A tablet computer may be used by connecting a casing with an input unit (typically a keyboard).

[0027] 《Example of the configuration of the first information processing device 10》 Next, we will describe an example configuration of the first information processing device 10.

[0028] The first information processing device 10 can perform processing using a language model. In particular, the first information processing device 10 can perform processing using a model that utilizes a large-scale language model (such as a text generation model or a dialogue model). For example, it can perform processing using large-scale language models such as GPT-4, Llama2, and Llama3. In this specification, the term "language model" includes large-scale language models.

[0029] The first information processing device 10 has the function of receiving instruction sentences. Furthermore, the first information processing device 10 has the function of outputting response sentences to instruction sentences using a language model, and is capable of performing various natural language processing tasks such as translation and summarization.

[0030] A provider of services using an information processing system according to one aspect of the present invention does not necessarily need to own the first information processing device 10 themselves. For example, a service provider can use a portion of the services provided by another business operator as the first information processing device 10.

[0031] 《Example of the configuration of the second information processing device 20》 Next, we will describe an example configuration of the second information processing device 20.

[0032] The second information processing device 20 has the function of receiving audio data, transcribing the audio data into text, and outputting text data. In particular, the second information processing device 20 can perform processing using a speech recognition model. For example, it can perform processing using a speech recognition model such as Whisper®.

[0033] A provider of services using an information processing system according to one aspect of the present invention does not necessarily need to own the second information processing device 20 themselves. For example, a service provider can use a portion of the services provided by another business operator as the second information processing device 20.

[0034] 《Example of the configuration of the third information processing device 30》 Next, an example configuration of the third information processing device 30 will be explained using Figure 2.

[0035] As shown in Figure 2, the third information processing device 30 includes a receiving unit 110, an output unit 120, a storage unit 130, a processing unit 140, and a transmission line 150. In Figure 2, in addition to the first information processing device 10 and the second information processing device 20, an information terminal 40 is shown, and the arrows indicate data transmission and reception. The receiving unit and the output unit together are sometimes called the communication unit. The communication unit enables the third information processing device 30 to send and receive data with the outside world.

[0036] [Reception Desk 110] The reception unit 110 has the function of receiving data from an external source. The reception unit 110 may, for example, use a communication port, or it may use an input device such as a personal computer equipped with communication capabilities.

[0037] For example, data received by the reception unit 110 from the first information processing device 10 may include conversation records output by the language model. Also, for example, data received by the reception unit 110 from the information terminal 40 may include user information and voice data.

[0038] The receiving unit 110 can supply the received data to the storage unit 130, the processing unit 140, or both, via the transmission line 150.

[0039] [Output section 120] The output unit 120 has the function of outputting calculation results and the like to an external device. For example, the output unit 120 can send an instruction to the first information processing device 10. Also, for example, the output unit 120 can send audio data to the second information processing device 20. The output unit 120 may be, for example, a communication port, or a device such as a personal computer equipped with communication capabilities.

[0040] Furthermore, in this embodiment, it is preferable that the audio data is encrypted before transmission and reception.

[0041] [Storage section 130] The memory unit 130 has a memory function. The memory unit 130 is a memory area that can store programs, data, etc. Typical programs include programs executed by the processing unit 140. Data includes data received by the reception unit 110 from the information terminal 40 (e.g., voice data, role list, and user information). It also includes data generated by the speech recognition model and the language model (e.g., text data converted from voice data by the speech recognition model, summaries generated by the language model, etc.). Furthermore, it includes various instructions used by the processing unit 140 when creating instruction statements (conversation record proofreading instructions, summarization instructions, role weight correction instructions, summary correction instructions, etc.).

[0042] The storage unit 130 may have a database. The third information processing device 30 may also have a database separate from the storage unit 130. The third information processing device 30 may have a function to retrieve data from a database located outside the storage unit 130, outside the third information processing device 30, or outside the information processing system. The third information processing device 30 may also have a function to retrieve data from both its own database and an external database.

[0043] The storage unit 130 can use either a storage device or a file server, or both. Furthermore, a database recording the paths of files stored on the file server can be used in the storage unit 130.

[0044] The storage unit 130 has at least one of volatile memory and non-volatile memory. Examples of volatile memory include DRAM (Dynamic Random Access Memory) and SRAM (Static Random Access Memory). Examples of non-volatile memory include ReRAM (Resistive Random Access Memory), PRAM (Phase Change Random Access Memory), FeRAM (Ferroelectric Random Access Memory), MRAM (Magnetoresistive Random Access Memory), and flash memory. The storage unit 130 may also have at least one of NOSRAM (registered trademark) and DOSRAM (registered trademark). The storage unit 130 may also have a recording media drive. Examples of recording media drives include hard disk drives (HDD) and solid state drives (SSD).

[0045] NOSRAM is an abbreviation for "Nonvolatile Oxide Semiconductor Random Access Memory (RAM)". NOSRAM is a type of memory where the memory cell is a 2-transistor (2T) or 3-transistor (3T) gain cell, and the transistors are transistors that use metal oxide in the channel formation region (also called OS transistors). OS transistors have an extremely small current flowing between the source and drain when off, i.e., a leakage current. By utilizing the characteristic of extremely low leakage current, NOSRAM can be used as a non-volatile memory by holding charge corresponding to the data within the memory cell. In particular, NOSRAM can read the stored data without destroying it (non-destructive read), making it suitable for computational processing that involves repeating data read operations a large amount. Because the data capacity of NOSRAM can be increased by stacking it, it can be used as a large-scale cache memory, main memory, or storage memory to improve the performance of semiconductor devices.

[0046] DOSRAM is an abbreviation for "Dynamic Oxide Semiconductor RAM," and refers to RAM with a 1T (transistor) 1C (capacitance) type memory cell. DOSRAM is a type of DRAM formed using OS transistors, and it is a memory that temporarily stores information sent from an external source. DOSRAM is a memory that takes advantage of the low off-current of OS transistors.

[0047] In this specification, "metal oxide" refers to an oxide of a metal in a broad sense. Metal oxides are classified into oxide insulators, oxide conductors (including transparent oxide conductors), oxide semiconductors (also called oxide semiconductors or simply OS), etc. For example, when a metal oxide is used in the semiconductor layer of a transistor, that metal oxide may be referred to as an oxide semiconductor.

[0048] The metal oxide in the channel-forming region preferably contains indium (In). When the metal oxide in the channel-forming region contains indium, the carrier mobility (electron mobility) of the OS transistor increases. For example, indium oxide can be suitably used as the metal oxide in the channel-forming region. Furthermore, the metal oxide in the channel-forming region is preferably an oxide semiconductor containing element M. Element M is preferably at least one of aluminum (Al), gallium (Ga), and tin (Sn). Other elements applicable to element M include boron (B), silicon (Si), titanium (Ti), iron (Fe), nickel (Ni), germanium (Ge), yttrium (Y), zirconium (Zr), molybdenum (Mo), lanthanum (La), cerium (Ce), neodymium (Nd), hafnium (Hf), tantalum (Ta), and tungsten (W). However, multiple elements mentioned above may be combined as element M. Element M is, for example, an element with a high bond energy with oxygen. For example, an element with a higher bond energy with oxygen than indium. Furthermore, the metal oxide containing the channel-forming region is preferably a metal oxide containing zinc (Zn). Metal oxides containing zinc may be more prone to crystallization. For example, indium gallium zinc oxide (also written as IGZO) or indium tin oxide (also written as ITZO®) can be used as the metal oxide containing the channel-forming region.

[0049] The metal oxides present in the channel-forming regions are not limited to indium-containing metal oxides. For example, the metal oxides present in the channel-forming regions may be zinc-tin oxides, gallium-tin oxides, or other metal oxides that do not contain indium but contain zinc, gallium, or tin.

[0050] [Processing Unit 140] The processing unit 140 has the function of performing calculations, analyses, and other processing using data supplied from either or both of the receiving unit 110 and the storage unit 130. The processing unit 140 can supply the processed data to either or both of the storage unit 130 and the output unit 120.

[0051] The processing unit 140 may, for example, have an arithmetic circuit. The processing unit 140 may, for example, have a central processing unit (CPU). In addition, the processing unit 140 may have a GPU (Graphics Processing Unit) in addition to or instead of the CPU. Furthermore, the processing unit 140 may have an NPU (neural processing unit / neural network processing unit).

[0052] The processing unit 140 may have registers and main memory in addition to the CPU. The registers and main memory may also be located within the CPU. The main memory is capable of sending and receiving data with a secondary cache, etc. The main memory includes at least one of volatile memory such as RAM (Random Access Memory) and non-volatile memory such as ROM (Read Only Memory). Furthermore, the main memory may include at least one of NOSRAM and DOSRAM. The main memory may include either or both OS transistors and Si transistors. Note that the configuration of the registers and main memory can be understood by substituting CPU with GPU in this paragraph.

[0053] Examples of RAM include DRAM and SRAM. DRAM or SRAM can also be used by virtually allocating memory space as a workspace for the processing unit 140. The operating system, application programs, program modules, program data, and lookup tables stored in the storage unit 130 are loaded into RAM immediately before execution. The operating system, application programs, program modules, program data, and lookup tables loaded into RAM can be accessed from the processing unit 140.

[0054] ROM can store systems that do not require rewriting. Examples of systems that do not require rewriting include BIOS (Basic Input / Output System) and firmware. Examples of ROM include mask ROM, OTPROM (One Time Programmable Read Only Memory), and EPROM (Erasable Programmable Read Only Memory). Examples of EPROM include UV-EPROM (Ultra-Violet Erasable Programmable Read Only Memory), which allows data to be erased by ultraviolet irradiation, EEPROM (Electrically Erasable Programmable Read Only Memory), and flash memory.

[0055] The processing unit 140 may have a microprocessor such as a DSP (Digital Signal Processor) in addition to the CPU or GPU. Since the DSP is specialized for digital signal processing, it is preferable to include it to control peripheral circuits of the CPU or GPU. The microprocessor may be implemented using a PLD (Programmable Logic Device) operating in hardware such as an FPGA (Field Programmable Gate Array) or FPAA (Field Programmable Analog Array).

[0056] The processing unit 140 enables the third information processing device 30 to create instruction sentences from text data. Specifically, it has the function of creating a first instruction sentence from conversation records, speaker labeling data, and conversation record proofreading instructions 210; the function of creating a second instruction sentence from proofread conversation records, roles, and summary instructions 220; the function of creating a third instruction sentence from proofread conversation records, summary lists, role lists, user information, and role weight creation instructions 230; and the function of creating a fourth instruction sentence from proofread conversation records, role lists, the first role weight list, summary lists, and summary correction instructions 240.

[0057] [Transmission path 150] The transmission line 150 has the function of transmitting data. Data can be transmitted and received between the receiving unit 110, the output unit 120, the storage unit 130, and the processing unit 140 via the transmission line 150. The transmission line 150 may be, for example, an external bus, a LAN (Local Area Network), or the Internet, which is the basis of the World Wide Web (WWW).

[0058] An example of an information processing method used in the information processing system described in this embodiment will be explained with reference to Figures 3 to 8.

[0059] <Example of processing in an information processing system> Figures 3, 4, and 5 are examples of flowcharts illustrating each step of an information processing system according to one aspect of the present invention. Figures 6 to 8 are examples of instruction sentences input to a language model.

[0060] <Step S101> In step S101, the third information processing device 30 acquires audio data from the information terminal 40. The audio data may be, for example, a recording of audio during a meeting. The audio data is formatted digital data, and formats such as wav (RIFF waveform Audio Format), mp3 (MPEG1 audio layer 3), and mp4 (MPEG4 Part 14) can be used.

[0061] The user input corresponding to step S101 is performed at the information terminal 40. The voice data is transmitted to the third information processing device 30 via the reception unit 110 and stored in the storage unit 130.

[0062] <Step S102> In step S102, the third information processing device 30 transmits the audio data to the second information processing device 20 via the output unit 120. The second information processing device 20 inputs the audio data into the speech recognition model and generates a conversation record and speaker labeling data. The second information processing device acquires the conversation record and speaker labeling data via the reception unit 110.

[0063] The conversation transcript is text data created by transcribing audio data and other information from a meeting.

[0064] Speaker labeling data is obtained by identifying the speaker in the transcribed text data of the conversation.

[0065] <Step S103> In step S103, the third information processing device, in its processing unit 140, creates a first instruction sentence from the conversation record, speaker labeling data, and conversation record proofreading instruction 210, as shown in Figure 6(A).

[0066] The conversation record proofreading instruction 210 is an instruction to the language model of the first information processing device 10 to correct interjections, unnecessary repetitions, and errors in the conversation record. For example, the text in the following paragraph can be used as the conversation record proofreading instruction 210.

[0067] "The above transcript was generated from an audio file. Please delete any unnecessary descriptions and correct any errors."

[0068] <Step S104> In step S104, the third information processing device 30 transmits the first instruction to the first information processing device 10 via the output unit 120. The first information processing device 10 inputs the first instruction to the language model and generates a first response sentence including a corrected conversation record. The second information processing device 20 receives the first response sentence via the reception unit 110.

[0069] <Step S105> In step S105, the third information processing device 30 obtains the role list from the storage unit 130.

[0070] A role list is a list of roles within a group such as a company. The role list is prepared in advance by the user and transmitted from the information terminal 40 to the third information processing device 30 via the reception unit 110, and stored in the storage unit 130.

[0071] <Step S106> In step S106, the third information processing device 30 retrieves a role from the role list. The retrieval of a role from the role list involves selecting one of the unselected roles during the repetition of steps S106 to S109, which will be described later.

[0072] <Step S107> In step S107, the third information processing device 30, in its processing unit 140, creates a second instruction sentence from the proofread conversation transcript, role, and summary instruction 220, as shown in Figure 6(B).

[0073] Summarization instruction 220 instructs the language model of the first information processing device 10 to create role-specific summaries from the proofread conversation transcript. The second instruction specifies a role and instructs the language model of the first information processing device 10 to generate a summary adapted to that role by generating a summary from the proofread conversation transcript. For example, the following paragraph can be used in summarization instruction 220.

[0074] "As an expert in the designated role, please create a summary from the conversation transcript."

[0075] <Step S108> In step S108, the third information processing device 30 transmits the second instruction to the first information processing device 10 via the output unit 120. The first information processing device 10 inputs the second instruction into the language model and generates a second response sentence including a summary for each role. The second information processing device 20 receives the second response sentence via the reception unit 110 and adds the summary for each role to the summary list.

[0076] <Step S109> In step S109, if not all roles have been obtained from the role list in step S106, the process returns to step S106; otherwise, the process proceeds to step S110.

[0077] By repeating steps S106 through S109, the summary list becomes a list of summaries for each role included in the role list. For example, when processing audio data from a meeting about system development, if the role selected in step S106 was a system engineer, the summary obtained in step S108 will focus on the requirements, maintenance, etc., of the system that the system engineer is responsible for.

[0078] <Step S110> In step S110, the third information processing device 30 obtains user information from the user. User information includes, for example, if the user is a member of a company, their job title, year of joining the company, age, and job duties.

[0079] The user input corresponding to step S110 is performed at the information terminal 40. User information is transmitted to the third information processing device 30 via the reception unit 110 and stored in the storage unit 130. Alternatively, the user input corresponding to step S110 may be omitted by reading the user information stored in the storage unit 130.

[0080] <Step S111> In step S111, the third information processing device 30, in its processing unit 140, creates a third instruction statement from the corrected conversation record, summary list, role list, user information, and role weight creation instruction 230, as shown in Figure 7.

[0081] The role weight creation instruction 230 instructs the language model of the first information processing device 10 to evaluate the relevance of each summary for each role as a weight, taking user information into consideration, and output it. Furthermore, it is preferable that the third instruction statement be written in a way that clearly shows the correspondence between the role list and the summary list. It is also preferable that the role weight creation instruction 230 includes an example of the weight format. Examples of weight formats include XML format and JSON format. For example, the text in the following paragraph can be used in the role weight creation instruction 230.

[0082] "We have created summaries for the above transcription results using multiple roles. For each summary, please weight its relevance to the user information and output the weights in the following format." <role> Roll 1< / role> <weight> Weight 1< / weight> <role> Roll 2< / role> <weight> Weight 2< / weight> "

[0083] <Step S112> In step S112, the third information processing device 30 transmits the third instruction to the first information processing device 10 via the output unit 120. The first information processing device 10 inputs the third instruction to the language model and generates a third response sentence including the first role weight list. The second information processing device 20 receives the third response sentence via the reception unit 110.

[0084] The first role weight list is output by the language model of the first information processing device 10 by associating the summary list with user information. The first role weight list is represented by each role and its weight. The weight can be represented as a string or a number. In the case of a string, for example, it can be represented as strong, weak, large, medium, small, heavy, light, etc. In the case of a number, it can be represented as a floating-point number, integer, or fraction. Table 1 is a table representation of the first role weight list, and is an example of the first role weight list obtained from the results of using user information of programmers with long experience in the third instruction statement. As shown in Table 1, the role weight list represents the weight of each role within the company, such as programmer and system engineer, as a floating-point number. Table 1 shows an example in which the role weights of programmers and experts are output as high as a result of using user information of programmers with long experience in the third instruction statement.

[0085] [Table 1]

[0086] <Step S113> In step S113, the third information processing device 30, in its processing unit 140, creates a fourth instruction statement from the corrected conversation record, the role list, the first role weight list, the summary list, and the summary correction instruction 240, as shown in Figure 8.

[0087] The summary modification instruction 240 instructs the language model of the first information processing device 10 to output a summary adapted for the user from the summary list, taking into account the first role weight list. For example, the text in the following paragraph can be used in the summary modification instruction 240.

[0088] "Based on the weights of each role as indicated in the weight information, please create a summary of the meeting minutes using the summaries in the summary list."

[0089] <Step S114> In step S114, the third information processing device 30 transmits the fourth instruction to the first information processing device 10 via the output unit 120. The first information processing device 10 inputs the fourth instruction into the language model and generates a fourth response sentence including a user-adapted summary. The second information processing device 20 receives the fourth response sentence via the reception unit 110.

[0090] A user-adapted summary is a summary adapted to the user from the summaries included in the summary list, according to the first role weight list. For example, if the first role weight list is as shown in Table 1, the user-adapted summary is generated by primarily reconstructing the summaries for the programmer role and the expert role in the summary list.

[0091] <Step S115> In step S115, the third information processing device 30 transmits the user-adapted summary to the information terminal 40 via the output unit 120, and the information terminal 40 outputs the user-adapted summary to the user.

[0092] By using such an information processing system, it is possible to provide users with a summary tailored to their needs from audio data recorded during meetings and other events.

[0093] <Modification of information processing systems> Figure 9 shows a modified flow chart of the information processing system. In Figure 9, steps S116 to S118 are added after step S112.

[0094] <Step S116> In step S116, the third information processing device 30 transmits the role list, summary list, and first role weight list to the information terminal 40 via the output unit 120, and the information terminal 40 outputs the role list, summary list, and first role weight list to the user.

[0095] <Step S117> In step S117, the third information processing device 30 obtains the second role weight list from the information terminal 40 via the reception unit 110.

[0096] In step S117, the second role weight list is the first role weight list output to the user in step S116, with the user modifying the string or numerical values ​​representing the weights for each role. By directly modifying the role weight list, users can create a summary that more accurately reflects their roles.

[0097] <Step S118> In step S118, the third information processing device 30, in its processing unit 140, creates a fourth instruction sentence from the corrected conversation record, the role list, the second role weight list, the summary list, and the summary correction instructions.

[0098] After completing step S118, proceed to step S114. This will allow the user to obtain a summary that is more tailored to them. [Explanation of Symbols]

[0099] 10 Information Processing Devices 20 Information Processing Devices 30 Information Processing Devices 40 Information terminals 50 Networks 60 Networks 110 Reception Department 120 Output section 130 Storage section 140 Processing Unit 150 transmission lines 210 Instructions for proofreading the conversation transcript 220 Summary Instructions Instructions for creating 230 roll weights 240 Summary correction instructions

Claims

1. It comprises a first information processing device, a second information processing device, and a third information processing device. The first information processing device has the function of receiving an instruction, processing it using a language model, and outputting a first response, a second response, a third response, and a fourth response. The second information processing device has the function of receiving audio data, processing it using a speech recognition model, and outputting conversation records and speaker labeling data. The third information processing device has a function for receiving the audio data, The function includes transmitting the aforementioned audio data to the second information processing device and receiving the conversation record and the speaker labeling data, A function to create a first instruction sentence from the aforementioned conversation transcript, the aforementioned speaker labeling data, and the conversation transcript proofreading instructions, A function to transmit the first instruction sentence to the first information processing device and to receive the first response sentence including a corrected conversation record, The ability to receive role lists, A function to retrieve a role from the aforementioned role list, A function to create a second instruction sentence from the aforementioned proofread conversation transcript, the aforementioned role, and summary instructions, A function to transmit the second instruction to the first information processing device and receive the second response including a summary for each role, A function to create summary lists, A function to add summaries for each role to the summary list, A function to receive user information, A function to create a third instruction sentence from the aforementioned proofread conversation transcript, summary list, role list, user information, and role weight creation instructions, A function to transmit the third instruction statement to the first information processing device and receive the third response statement including the first roll weight list, A function to create a fourth instruction sentence from the aforementioned corrected conversation transcript, the role list, the first role weight list, the summary list, and the summary revision instructions, A function to transmit the fourth instruction to the first information processing device and to receive the fourth response including a user-adapted summary, An information processing system having a function to output the user-adapted summary.

2. In claim 1, The third information processing device has a function to output the role list, the summary list, and the first role weight list, The function to receive a second list of role weights, An information processing system having a function to create a fourth instruction sentence from the aforementioned corrected conversation transcript, the role list, the second role weight list, the summary list, and the summary correction instructions.

3. The process comprises steps 1 through 15, In the first step described above, audio data is acquired, In the second step described above, the audio data is input to the speech recognition model to obtain conversation transcripts and speaker labeling data. In the third step described above, a first instruction sentence is created from the conversation transcript, the speaker labeling data, and the conversation transcript proofreading instructions. In the fourth step, the first instruction sentence is input into the language model to obtain a first response sentence including a corrected conversation transcript. In the fifth step described above, obtain the role list, In the sixth step described above, obtain a roll from the roll list, In the seventh step, a second instruction is created from the corrected conversation transcript, the role, and the summary instructions. In the eighth step, the second instruction is input to the language model, a second response is obtained that includes a summary for each role, and the summaries for each role are added to the summary list. In step 9, determine whether all of the role list has been processed. If not, return to step 6; if processed, proceed to step 10. In the tenth step described above, user information is obtained, In the 11th step, a third instruction is created from the corrected conversation transcript, the summary list, the role list, the user information, and the role weight creation instructions. In the 12th step, the third instruction is input to the language model, and a third response including the first role weight list is obtained. In the 13th step, a fourth instruction sentence is created from the corrected conversation transcript, the role list, the first role weight list, the summary list, and the summary revision instructions. In the 14th step, the fourth instruction is input to the language model, and a fourth response sentence including a user-adapted summary is obtained. An information processing method that outputs the user-adapted summary in the 15th step.

4. In claim 3, The process is followed by steps 16 through 18, In the 16th step, the role list, the summary list, and the first role weight list are output. In step 17 above, obtain a second roll weight list, An information processing method that, in step 18, creates a fourth instruction sentence from the corrected conversation transcript, the role list, the second role weight list, the summary list, and the summary correction instructions, and then proceeds to step 14.