Display information generation method, information processing device, program, and information processing system
The method addresses the challenge of classifying and arranging meeting statements by logical stance and chronological order, allowing clear visualization of discussion dynamics and progression.
Patent Information
- Application Number
- JP2026030628
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2026-02-27
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2046-02-27
AI Technical Summary
Conventional meeting minute creation systems fail to adequately classify and arrange statements based on their pro/con or neutral stance and chronological order, making it difficult to intuitively grasp the logical stance and progression of a meeting discussion.
A method to determine logical stance attributes for each statement, classify and arrange them into multiple regions, and display information based on chronological order, incorporating integration of similar statements and relationship estimation to visualize the discussion structure.
Enables intuitive understanding of the logical stance and chronological progression of meeting discussions by categorizing statements into conforming, neutral, and opposing regions, highlighting important statements, and displaying relationships between them.
Smart Images

Figure 0007910830000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to information processing for generating display information for visualizing minutes data based on audio data or text data related to a meeting.
Background Art
[0002] Conventionally, in order to record the content of a meeting, a minutes creation system has been proposed that records meeting audio, converts it into text by voice recognition or character recognition, and then automatically edits and distributes a minutes file (see, for example, Patent Document 1).
[0003] In addition, a technique has also been proposed for logically structuring posting content and generating and providing information visually illustrated for each opposing claim (see, for example, Patent Document 2).
[0004] Furthermore, techniques for improving the efficiency of meeting information sharing have also been proposed, such as inputting the speech content spoken in a meeting or text such as chat, and listing and outputting specific sentences based on type determination (see, for example, Patent Document 3).
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Patent Document 2
Patent Document 3
Non-Patent Documents
[0006]
Non-Patent Document 1
[0007] However, conventional meeting minute creation systems, such as those described in Patent Document 1, primarily focus on generating meeting minutes by transcribing and automatically editing the meeting content. They do not adequately disclose how to classify and arrange each statement in a display area according to its pro- or con- or neutral stance (logical stance), and how to generate and output display information to arrange it in chronological order.
[0008] Furthermore, while visualization techniques like those described in Patent Document 2 are useful in that they generate visualization information for each opposing argument, the framework for generating display information suitable for visualizing meeting minutes data is not clear. This framework involves determining logical stance attributes for multiple statements included in meeting logs, etc., classifying and arranging them into multiple domains (agree, neutral, disagree) according to those logical stance attributes, and outputting a combination of display information arranged based on the order of statements.
[0009] Therefore, with conventional technology, there is a problem in that it is difficult to intuitively grasp from meeting minutes data which statements are categorized as supporting, opposing, or neutral (questions, reservations, etc.) when reviewing the structure of a meeting's discussion, and how that progression unfolded chronologically.
[0010] Therefore, the present invention aims to provide a method for generating display information that allows the approval, disapproval, or neutral stance (logical stance) of each statement, as well as its chronological progression, to be grasped in the same display space, based on audio data or text data relating to a meeting, and to provide an apparatus, program, and system for executing said display information generation method. [Means for solving the problem]
[0011] The inventors of the present application have found that, in generating and outputting display information for visualizing meeting minutes data based on audio or text data related to a meeting, in order to grasp the pro / con or neutral stance of each statement and its chronological progression on the same display space, it is effective to determine attributes including a logical stance attribute indicating the pro / con or neutral stance included in the statement, generate display information for classifying and arranging into multiple areas according to the logical stance attribute, and display information for arranging according to the order of statements (chronological order), and output display information including these.
[0012] In other words, a display information generation method according to one aspect of the present invention includes: an attribute determination step of determining one or more attributes, including at least the logical stance attribute, for each of a plurality of statements contained in audio data or text data relating to a meeting; a first display information generation step of generating first display information for classifying and arranging target data corresponding to the plurality of statements into one of a plurality of regions in a display space based on the logical stance attribute; a second display information generation step of generating second display information for arranging the target data in the display space based on the order in which the plurality of statements are spoken; and an output step of outputting display information including the first display information and the second display information.
[0013] The aforementioned multiple regions may be configured to consist of a conforming region, a neutral region, and an opposing region corresponding to the logical stance attribute. Furthermore, the neutral region may be configured to include target data corresponding to the question and / or hold.
[0014] Furthermore, in another aspect of the present invention, the invention further includes an integration step to generate target data by integrating statements of the same meaning from among the plurality of statements, and the first display information and the second display information may be configured to be generated for the integrated target data. The determination of the same meaning may be configured to be based on at least one of temporal proximity, semantic similarity, or consecutive statements by the same speaker.
[0015] Furthermore, in another aspect of the present invention, the configuration may further include a relationship display information generation step, which involves estimating the relationship between the plurality of statements regarding counterarguments or references, generating relationship display information corresponding to the relationship, and including it in the display information. The estimation of the relationship may be based on explicit references, conjunctions, or the degree of agreement of the conclusions of the statements. The relationship display information may also include directional information indicating the direction of the counterargument or reference.
[0016] Furthermore, in another aspect of the present invention, the attribute may include an importance attribute in addition to the logical stance attribute, and the display information may be configured to include a display variable that controls the emphasis of the target data based on the importance attribute. The importance attribute may include the importance of the statement, and the emphasis may be configured to be determined based on the importance as a parameter relating to the attributes of the characters and / or the attributes of the border. The importance may be configured to be calculated based on at least one of the following: that it is a novel idea, that it includes a logical structure from evidence to conclusion, that it includes content that negates existing statements, or the number of subsequent responses.
[0017] In another aspect of the present invention, the attribute may include a quantitative attribute in addition to the logical stance attribute, and the display information may include a display variable that controls the size of the display area for the target data based on the quantitative attribute. The quantitative attribute may include the number of characters in the statement, and the size of the display area may be determined based on the number of characters.
[0018] In another aspect of the present invention, the logical stance attribute may be determined as either synchronized, neutral, or opposed, and the first display information may be configured to include information for arranging the target data in synchronized, neutral, and opposed regions in accordance with the logical stance attribute.
[0019] In another aspect of the present invention, the attribute determination step may be configured to include (a) a process of estimating or generating the attribute using a generative AI including a large-scale language model, and / or (b) a process of determining the attribute by logic based on predetermined rules or models.
[0020] In another aspect of the present invention, the display information may be configured to include, for each of the target data, at least region identification information indicating to which of the plurality of regions the data belongs, and coordinate information indicating a position within the display space.
[0021] In another aspect of the present invention, the order of speaking is defined based on a speaking order index assigned to the plurality of speeches, and the second display information may be generated based on the speaking order index.
[0022] According to the present invention configured as described above, in the minutes data, it is possible to grasp, on the same display space, to which of the synchronization side / opposing side / neutral (question, reservation, etc.) each speech belongs and the transition thereof in time series.
[0023] Furthermore, the present invention may be realized as a display information generation device (information processing device) that generates and outputs display information necessary for display based on voice data or text data related to a meeting. The device includes a processor and a memory, and the processor executes a program stored in the memory to execute each of the above steps (attribute determination, first display information generation, second display information generation, output, etc.).
[0024] The present invention may also be realized as a program for causing a computer to execute each of the above steps.
[0025] Furthermore, the present invention may be realized as a display information generation system (information processing system) including a terminal device and a server device. The terminal device acquires voice data or text data related to a meeting and transmits it to the server device. The server device executes each of the above steps (attribute determination, first display information generation, second display information generation, etc.) to generate display information, and the terminal device performs display based on the display information.
Advantages of the Invention
[0026] According to the invention of claim 1, it is possible to generate and output display information necessary for display so that the logical stance (for, against, or neutral) of each statement and its chronological progression can be grasped on the same display space, based on audio data or text data relating to a meeting.
[0027] According to the invention of claim 2, by classifying and arranging statements into multiple domains consisting of a conforming domain, a neutral domain, and an opposing domain corresponding to logical stance attributes, it is possible to clearly distinguish and understand which position each statement belongs to.
[0028] According to the invention of claim 3, by including target data corresponding to questions or reservations in the neutral area, it is possible to grasp the status of the discussion, including statements that cannot be immediately classified as for or against (questions, reservations, etc.).
[0029] According to the invention of claim 4, by integrating statements of the same nature to generate target data and generating display information for the integrated target data, it is possible to visualize the main points of the discussion while suppressing the repetition of redundant statements.
[0030] According to the invention of claim 5, by making a determination of the same meaning based on temporal proximity, semantic similarity, or consecutive statements by the same speaker, it is possible to integrate the groups of statements made in a meeting in a way that reflects the actual situation and to visualize them with high visual clarity.
[0031] According to the invention of claim 6, it is possible to grasp the interactions (counter-reference relationships) in a discussion by estimating the relationships between counterarguments or references between statements, generating relationship display information, and including it in the display information.
[0032] According to the invention of claim 7, it is possible to automatically estimate the relationship between counterarguments or references and reflect it in visualization by estimating the relationship based on the degree of agreement of explicit references, conjunctions, or the conclusion of statements.
[0033] According to the invention of claim 8, by including directional information in the relationship indication information, it is possible to clearly understand which statement a counterargument or reference leads to which statement.
[0034] According to the invention of claim 9, by including an importance attribute in addition to the logical stance attribute, and including a display variable in the display information that controls the degree of emphasis based on the importance attribute, it is possible to highlight statements according to their importance.
[0035] According to the invention of claim 10, by determining the emphasis level based on importance as a parameter relating to the attributes of the characters and / or the attributes of the border, it is possible to visually highlight and grasp important statements.
[0036] According to the invention of claim 11, by calculating importance based on any of the following: a novel idea, logical structure, content that negates existing statements, or the number of subsequent responses, it is possible to extract and emphasize noteworthy statements in a discussion using objective indicators.
[0037] According to the invention of claim 12, by including quantitative attributes in addition to logical stance attributes, and including a display variable in the display information that controls the size of the display area based on the quantitative attributes, it is possible to reflect the amount of speech (e.g., the number of characters) in the visualization.
[0038] According to the invention of claim 13, by using the number of characters as a quantitative attribute and determining the size of the display area based on the number of characters, it is possible to intuitively grasp the magnitude of the amount of speech.
[0039] According to the invention of claim 14, by determining the logical stance attribute as either conforming, neutral, or opposing, and including information in the first display information that places the target data in the conforming area, neutral area, and opposing area corresponding to the logical stance attribute, it is possible to clarify the correspondence between the classification of positions and their placement, and to grasp the conflict structure of the argument.
[0040] According to the invention of claim 15, attributes can be determined by a generative AI including a large-scale language model or by logic based on predetermined rules or models, and logical stance attributes, etc., can be flexibly estimated from meeting data and reflected in the generation of display information.
[0041] According to the invention of claim 16, by generating display information including region identification information and coordinate information for each target data, the arrangement within the display space can be treated as machine-readable data and used for display processing, redrawing, etc.
[0042] According to the invention of claim 17, by defining the order of speech based on a speech order index and generating second display information based on the speech order index, it is possible to realize a time-series arrangement based on clear sequential information.
[0043] According to the invention of claim 18, the above-mentioned display information generation method can be realized as an information processing device, and it is possible to provide a display information generation process based on meeting data as a device configuration.
[0044] According to the invention of claim 19, the above-mentioned display information generation method can be provided as a program for causing a computer to execute it, and the display information generation process can be implemented in various computer environments.
[0045] According to the invention of claim 20, it can be realized as an information processing system including a terminal device and a server device, and it is possible to efficiently operate by dividing the tasks of data acquisition and display on the terminal side and attribute determination and display information generation on the server side.
[0046] Furthermore, no specific configuration that allows the approval, disapproval, or neutral stance (logical stance) of each statement, as well as its chronological progression, to be grasped on the same display space based on audio data or text data related to a meeting, as described in the aforementioned Patent Documents 1-3 and Non-Patent Documents 1 and 2, is described at all. [Brief explanation of the drawing]
[0047] [Figure 1] This is a schematic diagram showing an example configuration of a display information generation system according to one embodiment of the present invention. [Figure 2] This is a functional block diagram of the terminal device. [Figure 3] This is a functional block diagram of the server device. [Figure 4] This is a sequence diagram showing an example of the processing sequence (main embodiment) between the terminal device and the server device. [Figure 5] This flowchart shows an example of the processing flow for generating display information. [Figure 6] This is a conceptual diagram that represents the classification and arrangement of logical stance attributes into multiple areas (first display information) and the arrangement based on the order of statements (chronological order) (second display information) on the same display space. [Figure 7] This is a conceptual diagram illustrating an example of determining attributes (such as logical stance attributes) of statements included in meeting logs, etc., using a generation AI (such as LLM) or a predetermined logic. [Figure 8] This is a conceptual diagram illustrating an example of integrating multiple statements of the same meaning to generate a group of statements (the target data after integration). [Figure 9] This is a conceptual diagram illustrating an example of estimating relationships (counterarguments / references, etc.) between multiple statements or integrated target data, and generating relationship display information. [Figure 10] This is a conceptual diagram illustrating an example of determining emphasis (e.g., border thickness) based on importance attributes. [Figure 11] This is a conceptual diagram illustrating an example of determining the size of the display area based on quantitative attributes (e.g., number of characters). [Figure 12A] This diagram shows an excerpt of an example of meeting log input (meeting minutes text). [Figure 12B] This figure shows examples of logical stance attributes (correspondence table) assigned to statements included in the input example (excerpt) shown in Figure 12A. [Figure 13A] This figure shows an example of a display generated based on meeting logs, etc. (the entire meeting log). [Figure 13B] This figure shows an example of the display following Figure 13A. [Modes for carrying out the invention]
[0048] A display information generation method, information processing device, program, and information processing system according to embodiments of the present invention will be described with reference to Figures 1 to 13B. Below, a display information generation system 1 shown in Figure 1 will be exemplified as an example of an information processing system. A server 20 will be exemplified as an example of an information processing device, and a program executed by the server 20 will be exemplified as an example of a program. Furthermore, a processing procedure following the processing flow shown in Figure 5 will be exemplified as an example of a display information generation method.
[0049] Figure 1 is a schematic diagram showing an example configuration of a display information generation system 1 according to one embodiment of the present invention. The display information generation system 1 comprises a terminal 10, a server 20, and a communication network 30 connecting the terminal 10 and the server 20.
[0050] Terminal 10 can acquire input data equivalent to meeting logs (for example, audio data D1 and text data D2) and transmit it to server 20 via communication network 30. Terminal 10 can also display information D8 received from server 20.
[0051] Based on the input data received from terminal 10, server 20 determines attributes (including at least logical stance attributes) for multiple statements included in the meeting log, etc., generates display information based on those attributes, and sends it back to terminal 10. Note that Figure 1 is an example to explain the division of responsibilities among processing entities, and some or all of the processing may be performed on terminal 10, or it may be performed by server 20 alone.
[0052] Figure 2 is a functional block diagram of terminal 10. Terminal 10 comprises an input acquisition unit 11, a transmission unit 12, a reception unit 13, and a display unit 14.
[0053] The input acquisition unit 11 acquires audio from the meeting or text such as meeting logs and generates input data. The transmission unit 12 sends the input data (e.g., D1, D2) acquired by the input acquisition unit 11 to the server 20. The receiving unit 13 receives display information D8 from the server 20. The display unit 14 displays multiple statements corresponding to the statements on the display space based on the display information D8.
[0054] The terminal 10 may include a storage unit 15 and a control unit 16 as needed. The storage unit 15 can store input data or display information, etc. The control unit 16 controls the operation of each part of the terminal 10.
[0055] Figure 3 is a functional block diagram of the server 20. The server 20 comprises an input receiving unit 21, a pre-processing unit 22, an attribute determination unit 23, a display information generation unit 24, and an output unit 25.
[0056] The input receiving unit 21 receives input data transmitted from the terminal 10. The pre-processing unit 22 may perform processing on the input data that contributes to the extraction of speech (for example, speech recognition, text formatting, speaker separation, timestamping, etc.).
[0057] The attribute determination unit 23 determines attributes, including logical stance attributes, for each of the multiple statements included in the meeting log, etc. (or the group of statements after integrating similar statements). The attribute determination unit 23 may be implemented by a generation AI unit 23a using a generation AI (LLM, etc.), or by a predetermined logic unit 23b using predetermined logic (rules, models, etc.). Note that attribute determination is not limited to a single inference, but may be performed as multiple dialogues or stepwise processing.
[0058] The display information generation unit 24 includes at least a first display information generation unit 24a and a second display information generation unit 24b. The first display information generation unit 24a generates first display information for classification and placement in one of several areas within the display space based on logical stance attributes. The second display information generation unit 24b generates second display information for placement within the display space based on the order of statements (chronological order). The display information generation unit 24 may also include a relationship display information generation unit 24c, if necessary, for generating relationship display information that represents the relationships between statements (counterarguments, references, etc.).
[0059] The output unit 25 outputs display information D8, including first display information and second display information (and relationship display information as needed), to the terminal 10. The server 20 may also include a storage unit 26 and a control unit 27 as needed. The storage unit 26 can store input data, statements, attributes, display information, etc. The control unit 27 controls the operation of each part of the server 20.
[0060] Next, the flow of the display information generation process performed between terminal 10 and server 20 will be explained based on Figures 4 and 5. In the following, the basic configuration of the present invention will be explained focusing on (i) attribute determination including logical stance attributes, (ii) generation of first display information, (iii) generation of second display information, and (iv) output of display information. On the other hand, preprocessing, utterance extraction, and similar meaning integration are positioned as preliminary or incidental processes for realizing the above basic configuration, and may be performed on the server 20 side or terminal 10 side depending on the implementation, or may be omitted as appropriate.
[0061] Herein, in this specification, "meeting log, etc." refers to information representing the content of statements made in a meeting, and includes audio data D1, text data D2, or a combination thereof. Furthermore, "statement" refers to a unit of statement extracted from the meeting log, etc., and can be defined, for example, as a speaking turn separated by a speaker, a speech interval within a predetermined time window, or a sentence unit. In addition, "group of statements" refers to a combined unit obtained by integrating multiple statements when they are judged to have the same meaning. In the following explanation, "statement" may be treated as including a group of statements unless otherwise inconsistent.
[0062] Furthermore, "target data" refers to data that is processed in response to multiple statements contained in a meeting log, etc., and includes, for example, statement data D3, or a group of statements after integration of similar content (integrated target data D3a, D3b, etc.). In other words, target data may be composed of individual statements, or it may be composed of a group of statements formed by integrating multiple statements. In the following explanation, unless otherwise inconsistent, target data will be treated as a concept that includes both statements and groups of statements.
[0063] Furthermore, "logical stance attribute" refers to an attribute that represents the position of a statement on a given topic, proposition, or agenda item, and may include attributes such as agreement (agreement), disagreement, or neutrality (which may also include questioning, reservation, etc.). "Display space" refers to a two- or three-dimensional space (screen, canvas, etc.) for arranging displays corresponding to multiple statements. Furthermore, "multiple areas" refers to divisions set within the display space for classifying and arranging statements according to logical stance attributes. For example, it may be divided into multiple areas such as an agreement area (agreement area), a neutral area, and an disagreement area.
[0064] Figure 4 is a sequence diagram showing an example of a processing sequence (main embodiment) for terminal 10 and server 20. Terminal 10 acquires input data corresponding to conference logs, etc. (e.g., voice data D1 and text data D2) using an input acquisition unit 11, and transmits the input data to server 20 using a transmission unit 12. Server 20 receives the input data using an input receiving unit 21, and may perform necessary preprocessing using a preprocessing unit 22. Specific examples of preprocessing include speech recognition for voice data D1, normalization of text data D2, formatting of line breaks and symbols, speaker estimation, timestamping, and removal of unnecessary parts.
[0065] Server 20 extracts multiple statements from meeting logs, etc. Statement extraction may be performed based on, for example, speaker boundaries, line breaks, periods, time information, or segments of speech recognition results. This generates statement data D3. Statement extraction may be performed as part of preprocessing, or as a separate step prior to attribute determination. Furthermore, after statement extraction, multiple statements may be merged to generate a group of statements. The determination of merger to the same purpose may be performed based on, for example, temporal proximity between statements, semantic similarity, or consecutive statements by the same speaker. This generates the merged target data (e.g., D3a, D3b, etc.).
[0066] Server 20 determines attributes, including logical stance attributes, for each of multiple statements using the attribute determination unit 23. Here, the attribute determination unit 23 may be implemented by a generation AI unit 23a using a generation AI (LLM, etc.), or by a predetermined logic unit 23b using predetermined logic (rules, statistical models, machine learning models, etc.). When using a generation AI, estimation results can be obtained by giving questions or extraction instructions regarding statements via prompts, etc., but such estimation is not limited to a single inference and may be performed as multiple dialogues or step-by-step processing. Furthermore, the estimation results may vary depending on the type, version, settings, etc. of the model, but since the present invention does not fix a specific estimation result and relates to information processing that determines attributes including logical stance attributes from meeting logs, etc., and generates display information based on those attributes, the above variations are permissible. As a result of attribute determination, attribute data D4 is generated.
[0067] When determining attributes using generative AI, the input and output formats are examples and not limited to those described, but a configuration that can be implemented by those skilled in the art may be as follows: The input includes at least utterance identification information (e.g., utterance ID) corresponding to the target data and the text of the target data (utterance turn or group of utterances), and may optionally include information representing the topic or agenda, information representing potential conclusions, and information representing the surrounding context. The output includes at least utterance ID and logical stance attributes (e.g., agree, neutral, disagree) corresponding to the utterance ID, and may optionally include confidence level or score, and information representing the basis for the estimation.
[0068] For example, a possible input instruction (prompt) for the generating AI could be: "For the following utterance text, classify the logical stance on the topic as either agree, neutral, or disagree, and output the result along with the utterance ID." An example of the output format could be structured data in JSON format, such as: {"utterance_id": "A123", "stance": "agree", "confidence": 0.85}.
[0069] The server 20 generates at least first and second display information using the display information generation unit 24. The first display information generation unit 24a generates first display information D5 for classifying and placing into one of several regions within the display space based on the logical stance attribute. Specifically, for example, it can generate region identification information and coordinate information for placing statements with a logical stance attribute of agreement (agreement) in the agreement region, statements with an disagreement attribute in the disagreement region, and statements with a neutral stance in the neutral region. Note that the number of regions is not limited to three, but may be divided into more regions. Furthermore, if the logical stance is treated as a numerical value (score), the intervals in the display space may be associated with multiple regions according to the range of the score, and the statements may be classified and placed within those intervals.
[0070] The second display information generation unit 24b generates second display information D6 for placement within the display space based on the order of statements (chronological order). Specifically, for example, it can assign a statement order index to each statement and generate information representing placement coordinates or display order according to the statement order index. Here, "placement based on statement order (chronological order)" means that the arrangement is based on the chronological order of the statements. For example, the placement coordinates or display order may be determined according to the statement order index. In addition, while the arrangement is based on the order of statements (chronological order), as an auxiliary measure to make the display easier to read, some of the statements (such as statements corresponding to conclusions) may be placed in a separate frame and highlighted.
[0071] The output unit 25 generates display information D8, which includes the first display information D5 and the second display information D6, and transmits it to the terminal 10. The terminal 10 receives the display information D8 via the receiving unit 13 and displays the information based on the display information D8 via the display unit 14. This realizes the output of display information as the basic configuration of the present invention.
[0072] Figure 5 is a flowchart showing an example of the processing flow for the display information generation method. The processing shown in Figure 5 may be executed by, for example, the server 20, but some or all of the processing may be executed by the terminal 10, or the server 20 may execute it alone.
[0073] In step S501, input data (D1, D2) is acquired. The input data may be acquired by terminal 10 and sent to server 20, or it may be acquired directly by server 20.
[0074] In step S502, multiple statements are extracted from the meeting log, etc., and statement data D3 is generated. Step S502 may be performed, for example, as part of preprocessing.
[0075] Next, as an example, in step S503, multiple statements of the same nature may be integrated to generate a group of statements (target data after integration). Step S503 can be positioned as a process for expanding the target unit for subsequent attribute determination and display information generation from statements to groups of statements.
[0076] In step S504, for each of the multiple statements, attributes including the logical stance attribute are determined, and attribute data D4 is generated. Step S504 may be executed by a generation AI or predetermined logic.
[0077] In step S505, first display information is generated. That is, first display information D5 is generated for classifying and placing into one of several areas within the display space based on the logical stance attribute.
[0078] In step S506, second display information is generated. That is, second display information D6 is generated to be placed in the display space based on the order of statements (chronological order).
[0079] In step S507, display information D8, which includes the first display information D5 and the second display information D6, is output. Display information D8 is sent to, for example, terminal 10 and displayed on terminal 10.
[0080] Thus, steps S504 to S507 shown in Figure 5 correspond to attribute determination, generation of first display information, generation of second display information, and output of display information, which are the basic configurations of the present invention. On the other hand, steps S502 and S503 are positioned as preliminary or incidental processes for forming the target units for attribute determination from meeting logs, etc., and may be combined or omitted as appropriate depending on the implementation.
[0081] Next, the content and generation examples of the display information of the present invention will be described based on Figures 6 to 13B. In the following, the core of the explanation will be the representation of classification arrangement based on logical stance attributes and arrangement based on the order of statements (chronological order) on the same display space, while integration of the same purpose, relationships between statements, emphasis based on importance, and the size of the display area based on quantitative attributes will be described as examples of additional elements that reinforce the core.
[0082] Herein, in this specification, "attribute" refers to information assigned to a statement included in a meeting log, etc., and includes at least a logical stance attribute, and, as necessary, various other information such as topic, agenda, speaker, sentiment, importance, responsiveness, orientation, or novelty. Attributes may be estimated or calculated based on the content of the statement, the surrounding context, speaker information, time information, etc.
[0083] Furthermore, in this specification, "first display information" refers to information for classifying and arranging a statement in one of several areas within the display space based on logical stance attributes, and includes, for example, area identification information, coordinate information within the area, or a combination thereof. "Second display information" refers to information for arranging statements within the display space based on the order of statements (chronological order), and includes, for example, a statement order index, coordinate information corresponding to the statement order index, display order information, or a combination thereof.
[0084] Furthermore, the above coordinate information only needs to be information that can identify the position in the display space, and its representation format is not limited. For example, the display space may be treated as a two-dimensional coordinate system, and the position may be represented as a pair of X and Y coordinates. Alternatively, the display space may be treated as a grid divided into a grid, and the position may be represented as a pair of row and column numbers. In addition, multiple lanes or rows (for example, lanes corresponding to a synchronization area, a neutral area, and an opposite area) may be set up within the display space, and the placement position may be represented by lane identification information and the position within the lane.
[0085] Furthermore, the representation of the arrangement based on the order of speaking in the second display information is not limited to coordinate information. For example, it may be configured as information that controls the display order according to the speaking order index, or as information that assigns the speaking order index to the display elements.
[0086] Figure 6 is a conceptual diagram that shows how to represent the classification and arrangement of statements into multiple areas corresponding to logical stance attributes (first display information) and the arrangement based on the order of statements (chronological order) (second display information) on the same display space. The display space is, for example, a two-dimensional space set on the screen of the display unit 14, and multiple areas for classifying statements according to logical stance attributes may be set within this display space. The multiple areas may be divided into, for example, a conformity area, a neutral area, and an opposition area. As a result, each statement is classified into one of the areas according to its logical stance attribute and placed in a predetermined position within that area.
[0087] Furthermore, within the same display space, each statement can be arranged based on the order of utterance (chronological order). For example, they may be arranged vertically or horizontally according to the order of utterance, or their coordinate values may be determined according to the utterance order index. This enables a display that makes it easy to simultaneously grasp the distribution of logical stances and the progression of utterances. Note that the number of regions is not limited to three, but may be divided into more regions. Also, when treating logical stances as numerical values (scores), the intervals in the display space may be associated with multiple regions according to the range of the score, and the statements may be classified and arranged within those intervals.
[0088] Figure 7 is a conceptual diagram illustrating an example of determining attributes (such as logical stance attributes) of statements included in meeting logs, etc., using a generative AI (such as an LLM) or predetermined logic. In this specification, "generative AI" refers to an example of an information processing model that performs inference or generation based on input text, etc., and includes, for example, a large-scale language model (LLM). Generative AI may include, for example, a trained model. Furthermore, generative AI and predetermined logic may be used in combination; for example, preprocessing (speaker separation, timestamping, normalization, etc.) may be performed by predetermined logic, and attribute determination, including logical stance attributes, may be performed by generative AI. Furthermore, "predetermined logic" refers to an example of a process that estimates or calculates attributes from statements using rules, judgment conditions, statistical models, machine learning models, etc. Attribute determination is not limited to logical stance attributes, but may also include various attributes such as topic, agenda, speaker, sentiment, importance, and responsiveness.
[0089] When using generative AI, logical stance attributes can be estimated by taking the content of a statement, the surrounding context, topic information, etc., as input. Estimation is not limited to a single inference, but may be performed as multiple dialogues or step-by-step processing. Furthermore, the estimation results may vary depending on the prompt, model type, version, settings, etc., but since the present invention does not fix a specific estimation result and relates to information processing that determines attributes including logical stance attributes from meeting logs, etc., and generates display information based on those attributes, the above variations are permissible.
[0090] Figure 8 is a conceptual diagram illustrating an example of integrating multiple statements of the same meaning to generate a group of statements (the target data after integration). Integrating statements of the same meaning is a process that groups together statements with the same or similar meaning from, for example, a meeting log, and treats them as a single integrated unit (group of statements). By performing this integration, even when there are many statements of the same meaning, the complexity of the display can be suppressed, making it easier to grasp the overall picture of the statements.
[0091] The determination of whether statements have the same meaning can be made based, for example, on temporal proximity, semantic similarity, or consecutive statements by the same speaker. These criteria are just examples, and other criteria may be used. Furthermore, merging statements of the same meaning is not mandatory, and each statement may be treated as a separate item to be displayed without merging. In the following explanation, "statement" may be treated as including the group of statements after merging, unless otherwise inconsistent.
[0092] Figure 9 is a conceptual diagram illustrating an example of estimating relationships (counterarguments, references, etc.) between multiple statements or integrated target data, and generating relationship display information. In this specification, "relationship display information" refers to information for representing relationships between statements in display space, and can be represented, for example, by arrows, lines, links, or notes. Relationships between statements can be estimated using clues such as quoted expressions, negative expressions, references to the same topic, or response relationships between speakers. If the relationship has direction (for example, from reference source to referenced source, from counterargument source to counterargument source), the direction may be indicated by arrows, etc.
[0093] Generating relationship information is an additional element that makes the displayed information easier to understand, but it is not mandatory. For example, relationship information may not be generated, and only classification arrangements based on logical stance attributes and arrangements based on the order of statements may be displayed.
[0094] Figure 10 is a conceptual diagram illustrating an example of determining emphasis (e.g., border thickness) based on importance attributes. In this specification, "importance attribute" refers to an attribute that represents the importance of a statement, and can be determined based on, for example, its contribution to the conclusion, its relevance to the agenda, the number of responses, or a predetermined evaluation index. "Emphasis" refers to a display parameter used to control visibility or prominence in the display, and can be expressed by, for example, border thickness, font weight, font size, presence or absence of hatching, etc. By determining emphasis based on importance attributes, important statements can be made easier to grasp visually.
[0095] Figure 11 is a conceptual diagram illustrating an example of determining the size of the display area based on quantitative attributes (e.g., number of characters). In this specification, "quantitative attributes" refer to attributes that represent the amount of speech, including, for example, the number of characters, the number of words, the duration of speech, the volume (speech intensity), or an indicator representing the amount of information. By determining the size of the display area based on quantitative attributes, it becomes easier to intuitively grasp the differences in the amount of speech. Note that the emphasis and size of the display area shown in Figures 10 and 11 are examples of display methods, and other display variables (e.g., color, shade, shape, etc.) may be used as auxiliary elements. However, since the core of this invention is the classification arrangement based on logical stance attributes and the arrangement based on the order of speech, it is desirable to use these as auxiliary elements based on these assumptions.
[0096] Figure 12A shows an excerpt of an example of input data (meeting minutes text) for meeting logs, etc. The input example shown in Figure 12A is just one example for space limitations, and actual meeting logs, etc., may be handled as text data D2 containing a larger number of statements.
[0097] Figure 12B shows an example (correspondence table) of logical stance attributes assigned to statements included in the input example (excerpt) shown in Figure 12A. The logical stance attributes shown in Figure 12B are just examples, and the attributes obtained as a result of attribute determination may include other attributes depending on the implementation.
[0098] Figures 13A and 13B show examples of displays generated based on meeting logs (the entire meeting log). In the display examples shown in Figures 13A and 13B, multiple statements included in the meeting log are classified and placed in one of several areas based on logical stance attributes, and are also arranged on the display space based on the order of statements (chronological order). This allows for an intuitive understanding of the distribution of approvals and disapprovals in a meeting and their progression within the same display space.
[0099] Figures 13A and 13B are display examples that reflect the entire meeting log, while Figures 12A and 12B are input examples and correspondence tables that show only a portion of the meeting log due to space limitations. Therefore, Figures 12A and 12B are examples of how to present input examples, and the ability to determine attributes and generate display information based on the entire meeting log remains unchanged.
[0100] According to the embodiment described above, multiple statements included in a meeting log, etc., can be represented in the same display space based on both a classification arrangement based on logical stance attributes and an arrangement based on the order of statements (chronological order), making it possible to intuitively grasp the distribution of approvals and disapprovals and their trends.
[0101] Furthermore, according to the embodiments described above, by adding elements such as integration of similar concepts, relationship display information, emphasis based on importance, and size of display areas based on quantitative attributes, it is possible to make it easier to grasp the relationships between statements or important statements while suppressing the complexity of the display.
[0102] The display information generation method, information processing apparatus, program, and information processing system according to the present invention are not limited to the embodiments described above, and various modifications and improvements are possible within the scope of the claims. Modifications are described below, but these may be combined as appropriate, as long as they do not contradict each other.
[0103] First, the division of processing responsibilities is just one example. In the embodiment described above, the terminal 10 is responsible for input acquisition and display, and the server 20 is responsible for attribute determination and display information generation, but the system is not limited to this. For example, if the input data is stored in the server 20, the server 20 may be configured to perform everything from acquiring the input data to outputting the display information. Also, if the terminal 10 has sufficient computing resources, the terminal 10 may be configured to perform a series of processes including attribute determination and display information generation. Furthermore, the system may be configured to distribute the execution across multiple information processing devices, such as by having different devices handle attribute determination and display information generation.
[0104] Next, the execution method for attribute determination is just one example. In the embodiments described above, examples were given of determining attributes using a generative AI (LLM, etc.) or predetermined logic, but the invention is not limited to these. For example, when using a generative AI, attributes may be determined by a single inference, or by multiple dialogues or stepwise processing. Examples of stepwise processing include a method in which statements are extracted, topics are estimated, logical stance is estimated, and importance is estimated, and these are executed sequentially in multiple processing stages. Furthermore, the prompts, model types, versions, settings, etc. of the generative AI can be changed as appropriate, and the attribute determination results may vary as a result. However, in the implementation of the present invention, it is sufficient to determine attributes including logical stance attributes from meeting logs, etc., and to generate display information based on those attributes.
[0105] Furthermore, the predetermined logic is not limited to rule-based determination, but may also include statistical models or machine learning models. For example, the system may be configured to estimate logical stance attributes using a classifier trained on past meeting logs. Moreover, the system may be configured to use a generative AI in combination with the predetermined logic. For example, the system may be configured to perform speaker separation or topic estimation using the predetermined logic, and then input the results into the generative AI to estimate logical stance attributes.
[0106] Next, the representation of logical stance attributes and the division into multiple regions are examples. In the embodiments described above, the example shown was the division into multiple regions based on three values: for, neutral, and against, but it is not limited to this. For example, "question" or "reserve" may be included in "neutral." Also, the multiple regions are not limited to three, but may be divided into more regions such as leaning towards for, leaning towards neutral, leaning towards against, etc. Furthermore, logical stance attributes may be represented as numerical values (scores), and intervals on the display space may be associated with multiple regions according to the range of the score, and classified and arranged within those intervals. In this case, the display may appear to be arranged continuously, but the classification and arrangement can be achieved by dividing according to the range of the score.
[0107] Next, the arrangement based on the order of remarks (chronological order) is just one example. In the embodiments described above, the arrangement on the display space according to the order of remarks was illustrated, but it is not limited to this. For example, while using the order of remarks as the basis, remarks corresponding to conclusions, remarks corresponding to summaries of topics, or remarks of high importance may be fixed in separate frames and highlighted. Furthermore, the representation of arrangement based on the order of remarks is not limited to arrangement by coordinates, but may also be represented by display order, lanes, hierarchy, or animation.
[0108] Next, the form of similar-purpose integration is just one example. In the embodiment described above, we exemplified a case where multiple statements are integrated based on similar purpose to generate a group of statements, but this is not limited to that. For example, similar-purpose integration may be based on any of the following: temporal proximity, semantic similarity, or consecutive statements by the same speaker, or a combination of multiple criteria. Furthermore, similar-purpose integration is just one example of how to simplify the display of statements, and each statement may be treated as a display target as is without integration. Moreover, the group of statements generated as a result of similar-purpose integration may be represented by extracting representative statements, summarizing multiple statements, or linking statements, etc.
[0109] Next, the relationships between statements and the forms of relationship indicator information are examples. In the embodiments described above, we have illustrated cases where relationships such as counterarguments or references are estimated and relationship indicator information is represented by arrows, etc., but we are not limited to this. For example, relationship estimation may be based on clues such as quoted expressions, negative expressions, agreement on topics, or response relationships between speakers. Direction may also be expressed as from the source of the reference, from the source of the counterargument to the target of the counterargument, etc. The representation of relationship indicator information is not limited to arrows, but may also be expressed by line segments, links, notes, or the proximity of indicators, etc. Furthermore, relationship indicator information is an additional element, and the system may be configured not to generate it.
[0110] Next, the manner in which importance attributes and emphasis levels are expressed is merely an example. In the embodiments described above, an example was given in which emphasis levels such as border thickness are determined based on importance attributes, but this is not limited to this. For example, importance attributes may be calculated based on the degree of contribution to the conclusion, the degree of relevance to the agenda, the number of responses, or predetermined evaluation indicators. Emphasis is not limited to border thickness, but may also be expressed by font thickness, font size, hatching, or bringing the display to the forefront.
[0111] Next, the quantitative attributes and the size of the display area are examples. In the embodiments described above, the case in which the size of the display area is determined based on quantitative attributes such as the number of characters was illustrated, but the invention is not limited to this. For example, the quantitative attributes may include the number of words, speech time, or an indicator representing the amount of information. The size of the display area may be controlled by area, width, height, etc.
[0112] Furthermore, the auxiliary elements of the display are merely examples. In the embodiments described above, examples were given of expressing differences through emphasis or the size of the display area, but the system is not limited to these. For example, display variables such as color, shade, shape, icon, or animation may be used as auxiliary elements. However, it is desirable to construct the display with a classification arrangement based on logical stance attributes and an arrangement based on the order of statements (chronological order) as the core, and to use these auxiliary elements for the purpose of reinforcing that core.
[0113] The variations described above are merely examples and are not limited to them. These variations may be combined as appropriate, as long as they do not contradict each other. [Explanation of Symbols]
[0114] 1. Display Information Generation System (Information Processing System) 10. Terminal (Terminal device) 11 Input acquisition unit 12 Transmitter 13 Receiving Unit 14 Display section 15 Storage section 16 Control Unit 20 Servers (Information Processing Devices) 21 Input Reception Section 22 Pre-processing section 23 Attribute determination section 23a Generation AI Department 23b Designated logic section 24 Display information generation section 24a First display information generation section 24b Second display information generation section 24c Relationship display information generation unit 25 Output section 26 Memory section 27 Control Unit 30 Communication Network D1 Audio Data D2 Text Data D3 Statement Data Target data after D3a integration Target data after D3b integration D4 Attribute Data D5 1st display information D6 2nd display information D7 Relationship Display Information D8 Display information S501 Input Acquisition S502 Speech Extraction S503 Integration of the Same Purpose S504 Attribute determination S505 First display information generation S506 Second display information generation S507 Display Information Output
Claims
1. A method for generating display information for visualizing meeting minutes data, which is performed by a computer, and which generates and outputs display information necessary for display based on audio data or text data related to the meeting. The computer performs an attribute determination step in which it determines, for each of the multiple statements contained in the audio data or text data, one or more attributes that include at least one logical stance attribute indicating a pro, con, or neutral position contained in the statement, A first display information generation step in which the computer generates first display information for classifying and arranging target data corresponding to the plurality of statements into one of a plurality of regions in the display space based on the logical stance attribute, The computer generates a second display information generation step in which it generates second display information for arranging the target data in the display space based on the order in which the plurality of statements are made, The computer includes an output step of outputting display information including the first display information and the second display information, The display information includes, for each target data, at least region identification information indicating which of the plurality of regions it belongs to, and coordinate information indicating its position within the display space. The order of statements is defined based on the statement order index assigned to the multiple statements, and the second display information is generated based on the statement order index. A method for generating display information, wherein the second display information is based on the arrangement according to the order of statements, and includes information for highlighting statements that correspond to conclusions or summaries of agenda items by fixing them in a separate frame.
2. An information processing device that generates and outputs display information necessary for display based on audio data or text data related to a meeting, A processor and memory are provided, and the processor executes a program stored in the memory, For each of the multiple statements contained in the aforementioned audio data or text data, determine one or more attributes that include at least one logical stance attribute indicating a pro, con, or neutral position contained in the statement. Based on the aforementioned logical stance attribute, first display information is generated for classifying and arranging the target data corresponding to the multiple statements into one of multiple regions within the display space. Based on the order in which the multiple statements are made, second display information is generated for arranging the target data within the display space. Output display information including the first display information and the second display information, The display information includes, for each target data, at least region identification information indicating which of the plurality of regions it belongs to, and coordinate information indicating its position within the display space. The order of statements is defined based on the statement order index assigned to the multiple statements, and the second display information is generated based on the statement order index. The second display information includes, based on the arrangement according to the order of statements, information for highlighting statements that correspond to conclusions or summaries of agenda items by fixing them in a separate frame.
3. On the computer, A step of determining one or more attributes for each of the multiple statements contained in audio or text data related to a meeting, including at least one logical stance attribute that indicates the position of approval, disapproval, or neutrality contained in the statement. A step of generating first display information for classifying and arranging target data corresponding to the plurality of statements into one of a plurality of regions in the display space, based on the aforementioned logical stance attribute, A step of generating second display information for arranging the target data within the display space based on the order in which the multiple statements are made. A step of outputting display information including the first display information and the second display information, A program to execute, The display information includes, for each target data, at least region identification information indicating which of the plurality of regions it belongs to, and coordinate information indicating its position within the display space. The order of statements is defined based on the statement order index assigned to the multiple statements, and the second display information is generated based on the statement order index. The program includes information for highlighting statements that correspond to conclusions or summaries of agenda items, while being based on the order in which the statements were made, by fixing them in a separate box.
Citation Information
Patent Citations
Data display instruction method
JP2003076639A
System for taking proceedings
JP2005175627A
Information processing device and information processing method
JP2012194963A
Information processing device, information processing method and program
JP2017016566A
Information sharing support system, information sharing support device, information sharing support method, and program
JP2017215943A