Information processing device, data output method, and data output program
The information processing device uses LLMs to generate and adjust the granularity of meeting minutes, addressing the issue of information density in conventional tools, allowing users to customize the level of detail in visualized meeting summaries.
Patent Information
- Application Number
- PCT/JP2024/027630
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-01
- Publication Date
- 2026-02-05
AI Technical Summary
Conventional tools for creating meeting minutes, especially those using AI, often provide too much or too little information, failing to meet user requirements in terms of information density.
An information processing device that utilizes large-scale language models (LLMs) to generate meeting minutes in a graph structure, allowing for adjustable granularity through user-defined parameters to visualize the information at an optimal level.
Enables visualization of meeting minutes that aligns with user preferences by adjusting the amount of information displayed, providing an overview without overwhelming the user.
Smart Images

Figure JP2024027630_05022026_PF_FP_ABST
Abstract
Description
Information processing device, data output method, and data output program
[0001] The present invention relates to an information processing device, a data output method, and a data output program.
[0002] In remote meetings such as web conferences and video conferences, if you miss a discussion, there is a psychological hurdle of asking questions and it is difficult to look back at the minutes, making it difficult to keep up with the discussion.
[0003] For these reasons, tools that use AI (Artificial Intelligence) to create meeting minutes are becoming more common. Furthermore, because chart-based minutes make it easier to grasp the main points than text-based minutes, visualization technology has been proposed that uses generative AI to generate images from transcribed text.
[0004] Hylable, [Retrieved July 22, 2024], Internet <URL: https: / / www.hylable.com / > AI Minutes Taker, [Retrieved July 22, 2024], Internet <URL: https: / / gijirokukun.com / >
[0005] However, the amount of information presented to users using conventional technologies such as the minutes-taking tools and visualization technologies described above can sometimes be too much or too little compared to the amount of information desired by the user, so there is room for improvement in providing the amount of information that meets the user's requirements.
[0006] Therefore, an object of the present invention is to provide an information processing device, a data output method, and a data output program that can realize visualization of an amount of information that meets a user's request.
[0007] In order to solve the above-mentioned problems and achieve the object, the information processing device of the present invention has an acquisition unit that acquires spoken text corresponding to spoken voice input in a conference system, a first model execution unit that inputs the spoken text into a large-scale language model, thereby causing the large-scale language model to output a first script in which a graph structure corresponding to minutes of a meeting in the conference system is described in a markup language, a second model execution unit that inputs the first script and parameters that specify the granularity of the graph structure into the large-scale language model, thereby causing the large-scale language model to output a second script in which a graph structure corresponding to the granularity is described, and an output unit that outputs graph data in which the graph structure corresponding to the granularity is visualized based on the second script.
[0008] According to the present invention, it is possible to realize visualization of an amount of information that meets a user's request.
[0009] FIG. 1 is a block diagram showing an example of the functional configuration of an information processing device. FIG. 2 is a schematic diagram (1) showing an example of the input / output form of the LLM 5. FIG. 3 is a diagram (1) showing an example of spoken text. FIG. 4 is a diagram (1) showing an example of a system prompt. FIG. 5 is a diagram (1) showing an example of an LLM output. FIG. 6 is a schematic diagram (2) showing an example of the input / output form of the LLM. FIG. 7 is a diagram (2) showing an example of a system prompt. FIG. 8 is a diagram (2) showing an example of an LLM output. FIG. 9 is a diagram showing an example of visualized data. FIG. 10 is a schematic diagram showing an example of a method for changing granularity specification parameters. FIG. 11 is a flowchart showing the procedure for data output processing. FIG. 12 is a diagram showing an example of a hardware configuration.
[0010] Hereinafter, an information processing device, a data output method, and a data output program according to the present disclosure will be described with reference to the accompanying drawings. Note that this embodiment merely illustrates one example or one aspect, and the following description does not limit the structure, action, function, properties, characteristics, methods, uses, etc. according to the present disclosure.
[0011] <Overall Configuration> Fig. 1 is a block diagram showing an example of the functional configuration of an information processing device 10. For example, Fig. 1 shows the information processing device 10 providing a data output function that outputs visualization data based on a first script in which spoken text corresponding to speech spoken at a meeting is input into a large-scale language model to describe a graph structure corresponding to meeting minutes, and a second script in which a graph structure of a specified granularity is re-input into the large-scale language model to re-describe the graph structure.
[0012] In one aspect, the information processing device 10 may be realized by a server device. For example, the information processing device 10 can provide the above-described data output function as a cloud service by executing PaaS (Platform as a Service) type middleware or SaaS (Software as a Service) type application.
[0013] 1, the information processing device 10 can be communicatively connected to the conference system 3 via a network NW. For example, the network NW may be realized by any type of communication network, such as the Internet or a local area network (LAN), whether wired or wireless.
[0014] The conference system 3 is a system that provides a communication function for sharing video, audio, materials, etc. between multiple locations. For example, the conference system 3 may be realized by any type of web conference system, such as an on-premise type, a cloud type, or a browser type. However, the conference system 3 is not limited to this, and may also be realized by a video conference system.
[0015] Note that, although an example in which the data output function is provided as a cloud service is given here, the present invention is not limited to this. For example, the data output function may be provided on-premise. Furthermore, the data output function may be packaged as one function of a service or application provided by a service provider that provides services related to remote meetings.
[0016] In addition, the above-described data output function is implemented as a client-server system, but is not limited to this. For example, the above-described data output function may be provided standalone by causing an application running on a user terminal participating in a conference held by the conference system 3 to execute processing corresponding to the above-described data output function on the user terminal.
[0017] <Configuration of Information Processing Device 10> Next, the functional configuration of the information processing device 10 that provides the above-mentioned data output function will be described. Fig. 1 schematically shows blocks related to the data output function of the information processing device 10. As shown in Fig. 1, the information processing device 10 has a communication control unit 11, a storage unit 13, and a control unit 15. Note that Fig. 1 only shows a selection of functional units related to the above-mentioned data output function, and the information processing device 10 may be provided with functional units other than those shown.
[0018] The communication control unit 11 is a functional unit that controls communication with other devices such as the conference system 3. In one embodiment, the communication control unit 11 can be realized by a network interface card such as a LAN card. In one aspect, the communication control unit 11 receives voice data or text data of utterances inputted by voice in a conference held in the conference system 3, or outputs visualized data outputted by the data output function to user terminals participating in the conference.
[0019] The storage unit 13 is a functional unit that stores various types of data. In one embodiment, the storage unit 13 may be realized by internal, external, or auxiliary storage of the information processing device 10. For example, the storage unit 13 stores an utterance text DB 13A and parameter setting data 13B. Note that the respective data in the utterance text DB 13A and the parameter setting data 13B will be described later together with scenes in which each data is referenced or registered.
[0020] The control unit 15 is a functional unit that performs overall control of the information processing device 10. For example, the control unit 15 may be realized by a hardware processor. As shown in FIG. 1 , the control unit 15 has an acquisition unit 15A, a first model execution unit 15B, a second model execution unit 15C, an output unit 15D, and a change unit 15E. Note that the control unit 15 may also be realized by hardwired logic or the like.
[0021] The acquisition unit 15A is a processing unit that acquires speech text corresponding to speech input to the conference system 3. In one aspect, the acquisition unit 15A can acquire speech text in which speech audio data is converted into text data using a transcription function, or so-called transcription function, provided by the conference system 3. For example, when using the transcription function, the acquisition unit 15A can acquire, along with the speech text, a speaker, a timestamp for the start or end of the speech section, and the like, for each speech section detected from an audio signal input in real time to the conference system 3. In this way, each time new speech text is acquired by the acquisition unit 15A, the new speech text is additionally registered in the speech text DB 13A. This allows the speech text acquired in real time from the start of the conference to be accumulated in chronological order. Note that while an example using the transcription function has been given here, it goes without saying that a pair of speaker and speech text can be acquired for each speech section by acquiring speech audio data and performing speech section detection, speaker estimation, and dictation.
[0022] The first model execution unit 15B and the second model execution unit 15C are both processing units that execute input / output control for large-scale language models, so-called LLMs (Large Language Models).
[0023] Of these, the first model execution unit 15B inputs spoken text into the LLM, while the second model execution unit 15C re-inputs the output of the LLM to which the spoken text has been input into the LLM, which is where the difference lies in their operation.
[0024] More specifically, the first model executing unit 15B can start processing each time the utterance text acquired by the acquiring unit 15A is additionally registered in the utterance text DB 13A. For example, the first model executing unit 15B inputs the utterance text stored in the utterance text DB 13A to the LLM, thereby causing the LLM to output a first script in which a structure corresponding to the minutes of the conference system 3 is written in a markup language.
[0025] Fig. 2 is a schematic diagram (1) showing an example of an input / output form of the LLM 5. Fig. 2 shows a schematic diagram of input / output control to the LLM 5 executed by the first model executing unit 15B. As shown in Fig. 2, the first model executing unit 15B inputs a user prompt 21, in which utterance text T stored in the utterance text DB 13A is embedded, and a system prompt 41 to the LLM 5.
[0026] Fig. 3 is a diagram (1) showing an example of utterance text. Fig. 3 shows an example of utterance text T to be embedded in the user prompt 21 shown in Fig. 2. As shown in Fig. 3, the utterance text T contains the utterances of each speaker in chronological order from the beginning of the conference, Suzuki, Harada, Takahashi, ..., Yamamoto, Takahashi, Suzuki, converted into text. In this way, the utterance text T, which is an accumulation of a series of utterance texts starting from the start of the conference, is embedded in the user prompt 21.
[0027] FIG. 4 is a diagram (1) showing an example of a system prompt. FIG. 4 also shows an example of the system prompt 41 shown in FIG. 2. As shown in FIG. 4, the system prompt 41 includes an instruction element 41A. This instruction element 41A includes instructions for expressing the structure of a flowchart corresponding to the minutes of a meeting based on inclusion relationships, relevance, and discussion transitions. Furthermore, the instruction element 41A includes instructions for extracting pairs of assertion phrases and the speakers of the assertion phrases for each assertion phrase included in the spoken text. The system prompt 41 also includes an example sentence 41B, embedded as a context element, in Mermaid notation, which has a graph structure with the spoken text as nodes and speakers as edges. Furthermore, the system prompt 41 includes an output format element, as underlined in FIG. 4, that specifies the format of the LLM5 output as Mermaid notation.
[0028] In this way, by implementing prompt engineering that imparts regularity to the chart description in the output of LLM5, it is possible to stabilize the output of LLM5 even when there is variation in the content of the text input to LLM5. Furthermore, by extracting pairs of speech phrases and speakers in units of assertions, even when one speech section contains multiple topics, it is possible to divide it into multiple nodes and describe the flowchart.
[0029] The LLM 5 receives the user prompt 21 and the system prompt 41 and outputs a first script SC1 written in Markdown notation, for example, Mermaid notation.
[0030] FIG. 5 is a diagram (1) showing an example of an LLM output. FIG. 5 shows the first script SC1 output by the LLM 5 shown in FIG. 2. As shown in FIG. 5, the first script SC1 describes a flowchart code block that proceeds left or right using Mermaid notation. For example, in the example shown in FIG. 5, a total of 55 nodes are defined in the flowchart, from ID1 "Enhancing security for public web services" to ID55 "Instructions for moving to plan execution." Furthermore, a total of 53 edges are defined in the flowchart, described from line 2 to line 54 of "flowchartLR" shown in FIG. 5 (lines 14 to 39 are omitted). In this way, the first script SC1 is obtained, describing the structure of a flowchart corresponding to the details of the entire spoken text T.
[0031] Then, the second model execution unit 15C inputs the first script SC1 and parameters specifying the granularity of the graph structure into the LLM5, causing the LLM5 to output the second script SC2 describing the graph structure corresponding to the granularity.
[0032] FIG. 6 is a schematic diagram (2) showing an example of an input / output form of the LLM5. FIG. 6 illustrates the input / output control of the LLM5 executed by the second model execution unit 15C. As shown in FIG. 6, the second model execution unit 15C inputs to the LLM5 a user prompt 22 embedded with the first script SC1 input by the first model execution unit 15B, i.e., the first script SC1 output by the LLM5 to which the utterance text T is input. Furthermore, the second model execution unit 15C inputs to the LLM5 a system prompt 42 embedded with a specification P1 of the range included in the parameter setting data 13B, a specification P2 of the maximum number of nodes, and a specification P3 of the maximum number of stages.
[0033] These "range," "maximum number of nodes," and "maximum number of levels" correspond to examples of granularity specification parameters that specify the granularity of the graph structure. Among these, "range" may refer to the range of graph data defined by the description in the first script SC1 that is focused on as the target of attention for the second script SC2. For example, by specifying an arbitrary node from the graph data defined by the description in the first script SC1, all nodes located in a lower hierarchy than the specified node can be specified as the "range." Furthermore, "maximum number of nodes" may refer to the maximum number of nodes constituting the flowchart defined by the description in the second script SC2. Furthermore, "maximum number of levels" may refer to the number of levels constituting the flowchart defined by the description in the second script SC2, in other words, the maximum depth or number of levels.
[0034] FIG. 7 is a diagram (2) showing an example of the system prompt 42. FIG. 7 shows an example of the system prompt 42 shown in FIG. 6. As shown in FIG. 7, the system prompt 42 includes a directive for narrowing the range of the flowchart defined by the description of the first script SC1 to be focused on as the target of attention of the second script SC2 to a range equal to or less than the specified range P1 set in the parameter setting data 13B. Furthermore, the system prompt 42 includes a directive for re-editing and summarizing the nodes included in the flowchart defined by the description of the first script SC1 until the number of nodes constituting the flowchart defined by the description of the second script SC2 is equal to or less than the specified maximum number of nodes P2 and the number of stages constituting the flowchart is equal to or less than the specified maximum number of stages P3.
[0035] By such system prompt 42, the first script SC1, which describes the structure of a flowchart corresponding to the details of the entire spoken text T, is summarized according to the range specification P1, the maximum number of nodes specification P2, and the maximum number of stages specification P3.
[0036] When these user prompts 22 and system prompts 42 are input, the LLM 5 outputs a second script SC2 in which the first script SC1, which describes the structure of a flowchart corresponding to the details of the entire spoken text T, is summarized according to a range specification P1, a maximum number of nodes specification P2, and a maximum number of stages specification P3.
[0037] Here, at the time of initial visualization, it is possible to use, as an example only, the granularity specification parameters registered as system settings in the parameter setting data 13B, namely, the range specification P1, the maximum number of nodes specification P2, and the maximum number of stages specification P3. For example, the parameter setting data 13B may be system-set with the range specification P1, the maximum number of nodes specification P2, and the maximum number of stages specification P3 that allow the flowchart defined by the description of the second script SC2 to be visualized at the maximum size that fits within the Web conference screen displayed on the user terminal or within a display area allocated to a part of the screen.
[0038] In this way, it is not necessary to visualize the first script SC1, which describes the structure of a flowchart corresponding to the details of the entire utterance text T. This is because visualizing the first script SC1 is likely to overwhelm the user with information. For example, in the example of the first script SC1 shown in Figure 6, graph data corresponding to a flowchart including a total of 55 nodes and a total of 53 edges would be generated, which would result in the graph data being displayed on the web conference screen or in a display area of a portion thereof becoming overcrowded.
[0039] Therefore, it is possible to visualize a second script SC2, i.e., a cut-out summary of the first script SC1, which describes the structure of a flowchart summarized according to the range specification P1, the maximum number of nodes specification P2, and the maximum number of stages specification P3.
[0040] Fig. 8 is a diagram (2) showing an example of the output of LLM5. Fig. 8 shows the second script SC2 output by LLM5 shown in Fig. 6. Furthermore, Fig. 8 shows, as an example of range specification P1, the result of specifying the entire flowchart defined by the description of first script SC1, i.e., the range from the node with ID1, which is the top root node, to the bottom node. Furthermore, Fig. 8 shows the result of specifying "20" as an example of the maximum number of nodes specification P2 and "9" as an example of the maximum number of stages specification P3.
[0041] 8, the description in the second script SC2 summarizes the entire flowchart defined by the description in the first script SC1 down to 13 nodes, which is less than the specified maximum number of "20." Furthermore, the description in the second script SC2 summarizes the entire flowchart defined by the description in the first script SC1 down to 9 stages, which is less than the specified maximum number of stages of "9." In this way, the description in the second script SC2 makes it possible to realize a cut-and-paste summary that allows an overview of the entire flowchart defined by the description in the first script SC1.
[0042] Here, we have given an example of using granularity specification parameters set by the system at the time of initial visualization, but it goes without saying that it is also possible to accept user specification of granularity specification parameters from the time of initial visualization.
[0043] The output unit 15D is a processing unit that executes output control for user terminals and the like included in the conference system 3. In one aspect, the output unit 15D outputs graph data in which a graph structure corresponding to the granularity specification parameters is visualized based on the second script SC2 obtained as a result of input / output control to the LLM 5 by the second model executing unit 15B. Such visualized data may be displayed on a GUI screen, such as a web conference screen, shared by user terminals participating in a conference held in the conference system 3. For example, the output unit 15D can input the second script SC2, for example, a code block in which a flowchart is written in Mermaid notation, into a graph creation tool, thereby causing the graph creation tool to draw the graph data.
[0044] Fig. 9 is a diagram showing an example of visualization data. Fig. 9 shows an example in which graph data G1 visualized from the second script SC2 shown in Fig. 8 is displayed in a partial display area of the Web conference screen provided by the conference system 3. As shown in Fig. 9, the graph data G1 summarizes the details of the entire utterance text T down to the number of nodes "13" and the number of lines "9." Therefore, the graph data G1 can realize visualization of an amount of information at a level that allows an overview of the entire utterance text T, for example, visualization of an information density of "coarse."
[0045] The change unit 15E is a processing unit that changes the granularity specification parameters set in the parameter setting data 13B. In one aspect, the change unit 15E can accept user settings of the granularity specification parameters from user terminals participating in a conference held in the conference system 3 via left and right click operations of a mouse. For example, an operation to increase the maximum number of nodes and the maximum number of stages can be assigned to a left click, while an operation to decrease the maximum number of nodes and the maximum number of stages can be assigned to a right click.
[0046] 10 is a schematic diagram showing an example of a method for changing a granularity parameter. For example, FIG. 10 shows an example in which the range specification is fixed for the entire flowchart defined by the description of the first script SC1, and a user setting of the granularity parameter is accepted via left-click or right-click. Furthermore, FIG. 10 shows an example in which, per left-click operation, the currently set maximum number of nodes is multiplied by "3" and the currently set maximum number of stages is added by "5," while per right-click operation, the currently set maximum number of nodes is multiplied by "1 / 3" and the currently set maximum number of stages is subtracted by "5."
[0047] 10 , when a single left click is received while the graph data G1 shown in FIG. 9 is displayed in the display area 60 of the Web conference screen, the system settings set in the parameter setting data 13B, i.e., the maximum number of nodes of "20" and the maximum number of stages of "9," are changed as follows: The maximum number of nodes of "20" is multiplied by "3," the value of one left click, to change it to "60," and the maximum number of stages of "9" is increased by "5," the value of one left click, to change it to "14." In this case, a system prompt 42 in which the new maximum number of nodes and the new maximum number of stages are embedded, and a user prompt 22 in which the first script SC1 is embedded, are input to the LLM 5. As a result, the LLM 5 outputs the second script SC2 after the granularity specification parameters have been changed.
[0048] Then, the graph data G2 visualized from the second script SC2 after the granularity specification parameter change is displayed. Here, let us assume that the width and height of the display area 60 of the Web conference screen are "W" and "H," respectively, and that the nodes and edges of the graph data G1 and G2 are visualized at the same size. Let us assume that the graph range 61 that fits the entire graph data G2 has a width of "2W" and a height of "2H." In this case, the display scope of the entire graph data G2 to be displayed in the display area 60 of the Web conference screen can be determined based on the cursor coordinate position where a left-click operation was accepted on the display area 60 of the Web conference screen.
[0049] In one aspect, the display scope of the graph data G2 can be set by superimposing the cursor coordinate position on the display area 60 of the Web conference screen on a corresponding point of the cursor coordinate position on a range 61 that contains the entire graph data G2. For example, when the cursor coordinate position on the display area 60 of the Web conference screen is (2, 2), the corresponding point in the graph range 61 of the graph data G2 is (4, 4), where the ratio of the x coordinate corresponding to the horizontal direction from the origin (0, 0) in the upper left and the y coordinate corresponding to the vertical direction is equal to the ratio of the cursor coordinate position. The display scope 60 shown by the dashed line in FIG. 10 is set by superimposing the cursor coordinate position (2, 2) on the display area 60 of the Web conference screen on the corresponding point (4, 4) on the graph range 61 of the graph data G2.
[0050] In this way, since it is possible to accept user settings for any granularity specification parameters via left and right click operations with the mouse, it is possible to realize visualization of the amount of information that meets the user's requirements.
[0051] 10 shows an example in which the user setting of the granularity designation parameter is accepted via a left click or a right click, but the user setting of the granularity designation parameter can also be accepted via other mouse operations, such as rotating the mouse wheel, for example, forward or backward. Furthermore, it goes without saying that the user setting of the granularity designation parameter can be accepted via GUI operations, not limited to mouse operations.
[0052] <Processing Flow> Next, the processing flow of the information processing device 10 according to this embodiment will be described. Fig. 11 is a flowchart showing the procedure of the data output processing. This processing is merely an example and can be started when a request for displaying graph data is received from a user terminal.
[0053] 11, the acquisition unit 15A acquires the utterance text T stored in the utterance text DB 13A (step S101). Then, the first model execution unit 15B generates the user prompt 21 in which the utterance text T acquired in step S101 is embedded (step S102).
[0054] Next, the first model executing unit 15B inputs the user prompt 21 generated in step S102 and the system prompt 41 used to structure the details of the entire spoken text T to the LLM 5 (step S103).
[0055] Then, the second model executing unit 15C generates a user prompt 22 in which the first script SC1 output by the LLM 5 as a result of the input in step S103 is embedded (step S104).
[0056] Furthermore, the second model executing unit 15C generates a system prompt 42 in which a specification P1 of the range included in the parameter setting data 13B, a specification P2 of the maximum number of nodes, and a specification P3 of the maximum number of stages are embedded (step S105).
[0057] Then, the second model executing unit 15C inputs the user prompt 22 generated in step S104 and the system prompt 42 generated in step S105 to the LLM 5 (step S106).
[0058] Then, the output unit 15D displays the graph data in which the graph structure corresponding to the granularity specification parameters is visualized based on the second script SC2 output by the LLM5 as a result of the input in step S106 in the display area of the web conference screen displayed on the user terminal (step S107).
[0059] Thereafter, the change unit 15E waits for a mouse click operation on the display area of the Web conference screen (Step S108). If a mouse click operation is accepted (Yes in Step S109), the change unit 15E determines whether the click operation is a left click (Step S110).
[0060] At this time, if the click operation is a left click (Yes at step S110), the change unit 15E reflects the user settings for increasing the maximum number of nodes and the maximum number of stages in the parameter setting data 13B (step S111).
[0061] On the other hand, if the click operation is a right click (No at step S110), the change unit 15E reflects the user settings for decreasing the maximum number of nodes and the maximum number of stages in the parameter setting data 13B (step S112).
[0062] Thereafter, the process proceeds to step S104, and the process from step S104 is repeated. As a result, the system prompt 42 in which the changed granularity specification parameter is embedded and the user prompt 22 in which the first script SC1 is embedded are input to the LLM 5, and the output of the second script SC2 in which the granularity specification parameter has been changed can be obtained from the LLM 5.
[0063] The process shown in FIG. 11 is merely an example and may continue until the Web conference screen is closed.
[0064] <Summary> As described above, the information processing device 10 according to this embodiment outputs visualization data based on a first script SC1 that inputs speech text corresponding to speech spoken at a conference into the LLM 5 and describes a graph structure corresponding to the minutes of the conference, and a second script that re-inputs the first script SC1 into the LLM 5 and re-describes a graph structure of a specified granularity. Therefore, the information processing device 10 according to this embodiment can visualize an amount of information that meets a user's request.
[0065] <Exercising Creative Ability> The matters described in this embodiment, such as specific examples of speech text and prompts, and the types of LLMs, are merely examples and can be changed. Furthermore, the order of processes in the flowchart described in this embodiment can be changed or some processes can be skipped within a consistent range.
[0066] <System> The information including the processing procedures, control procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, any one or more of the functional units of the acquisition unit 15A, the first model execution unit 15B, the second model execution unit 15C, the output unit 15D, and the change unit 15E of the information processing device 10 may be configured as separate devices.
[0067] Furthermore, the components of each device shown in the figure are functional concepts and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown. In other words, all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc. Note that each configuration may also be a physical configuration.
[0068] Furthermore, all or any part of the processing performed by the illustrated device may be realized by a program executed by a hardware processor such as an MPU (Micro-Processing Unit) or a CPU (Central Processing Unit), or may be realized as hardware using wired logic.
[0069] <Hardware> Next, an example of the hardware configuration of the information processing device 10 described in this embodiment will be described. For example, the information processing device 10 can be implemented by installing a program that realizes the functions of the information processing device 10 on a computer. For example, the information processing device 10 can be implemented by having the computer execute the program, which is provided as package software or online software. The computer referred to here includes desktop or notebook personal computers, rack-mounted server computers, and the like. Furthermore, the computer also includes smartphones, mobile phones, PHS (Personal Handyphone System) and other mobile communication terminals, as well as PDAs (Personal Digital Assistants). The functions of the information processing device 10 may also be implemented on a cloud server.
[0070] An example of a computer that executes the above program (data output program) will be described using Fig. 12. As shown in Fig. 12, the computer 1000 includes, for example, a memory 1010, a CPU 1020, a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.
[0071] The memory 1010 includes a read-only memory (ROM) 1011 and a random access memory (RAM) 1012. The ROM 1011 stores a boot program such as a basic input / output system (BIOS). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to a mouse 1110 and a keyboard 1120, for example. The video adapter 1060 is connected to a display 1130, for example.
[0072] 12, the hard disk drive 1090 stores, for example, an OS 1091, an application program 1092, a program module 1093, and program data 1094. The storage unit 13 described in the above embodiment is provided in, for example, the hard disk drive 1090 or the memory 1010.
[0073] Then, the CPU 1020 reads out the program module 1093 and the program data 1094 stored in the hard disk drive 1090 into the RAM 1012 as necessary, and executes the above-mentioned procedures.
[0074] The program module 1093 and program data 1094 relating to the data output program are not limited to being stored in the hard disk drive 1090, but may be stored in a removable storage medium, for example, and read by the CPU 1020 via the disk drive 1100. Alternatively, the program module 1093 and program data 1094 relating to the program may be stored in another computer connected via a network such as a LAN or a WAN (Wide Area Network), and read by the CPU 1020 via the network interface 1070.
[0075] 3 Conference system 5 LLM 10 Information processing device 11 Communication control unit 13 Storage unit 13A Speech text DB 13B Parameter setting data 15 Control unit 15A Acquisition unit 15B First model execution unit 15C Second model execution unit 15D Output unit 15E Change unit
Claims
1. An information processing device comprising: an acquisition unit that acquires spoken text corresponding to spoken voice input in a conference system; a first model execution unit that inputs the spoken text into a large-scale language model, thereby causing the large-scale language model to output a first script in which a graph structure corresponding to minutes of the conference system is described in a markup language; a second model execution unit that inputs the first script and parameters that specify the granularity of the graph structure into the large-scale language model, thereby causing the large-scale language model to output a second script in which a graph structure corresponding to the granularity is described; and an output unit that outputs graph data in which the graph structure corresponding to the granularity is visualized based on the second script.
2. The information processing device described in claim 1, characterized in that the first model execution unit inputs into the large-scale language model a prompt that uses the spoken text as input data, has the spoken text as nodes, and has a graph structure with the speaker of the spoken text as edges, and includes example sentences written in the markup language as context elements.
3. The information processing device described in claim 1, characterized in that the second model execution unit inputs a prompt to the large-scale language model that includes an instruction element to narrow down the range of graph data defined by the description of the first script to be focused on as the target of attention of the second script to a specified range or less.
4. The information processing device described in claim 1, characterized in that the second model execution unit inputs a prompt to the large-scale language model including an instruction element to re-edit and summarize the nodes included in the graph data defined by the description of the first script until the number of nodes constituting the graph data defined by the description of the second script becomes less than or equal to a specified maximum number of nodes.
5. The information processing device described in claim 1, characterized in that the second model execution unit inputs a prompt to the large-scale language model including an instruction element to re-edit and summarize the nodes included in the graph data defined by the description of the first script until the number of levels of the graph data defined by the description of the second script becomes less than or equal to a specified maximum number of levels.
6. The information processing device of claim 1, further comprising a modification unit that accepts changes to parameters that specify the granularity of the graph structure, and wherein the second model execution unit inputs the first script and parameters that specify the granularity after the modification by the modification unit into the large-scale language model, thereby causing the large-scale language model to output a second script that describes a graph structure that corresponds to the granularity after the modification.
7. A data output method executed by an information processing device, comprising: an acquisition step of acquiring spoken text corresponding to spoken voice input in a conference system; a first model execution step of inputting the spoken text into a large-scale language model, thereby causing the large-scale language model to output a first script in which a graph structure corresponding to minutes of the conference system is described in a markup language; a second model execution step of inputting the first script and parameters specifying the granularity of the graph structure into the large-scale language model, thereby causing the large-scale language model to output a second script in which a graph structure corresponding to the granularity is described; and an output step of outputting graph data in which the graph structure corresponding to the granularity is visualized based on the second script.
8. A data output program for causing a computer to execute the following steps: an acquisition step of acquiring spoken text corresponding to spoken voice input in a conference system; a first model execution step of inputting the spoken text into a large-scale language model, thereby causing the large-scale language model to output a first script in which a graph structure corresponding to minutes of a conference in the conference system is described in a markup language; a second model execution step of inputting the first script and parameters specifying the granularity of the graph structure into the large-scale language model, thereby causing the large-scale language model to output a second script in which a graph structure corresponding to the granularity is described; and an output step of outputting graph data in which the graph structure corresponding to the granularity is visualized based on the second script.