Information processing device and information processing method
Patent Information
- Application Number
- PCT/JP2025/012149
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2026-10-01
Smart Images

Figure JP2025012149_01102026_PF_FP_ABST
Abstract
Description
Information Processing Apparatus and Information Processing Method
[0001] The present disclosure relates to an information processing apparatus and an information processing method.
[0002] Meeting agent technologies such as facilitation are known as technologies for analyzing participants' opinions, ideas, and the like in meetings, smoothly advancing the progress of proceedings, and supporting consensus building (see Patent Document 1). Implementation of such meeting agent technology requires real-time interaction based on voice input and output for conversations between users and the system.
[0003] On the other hand, in recent years, research on meeting agent technology using large language models (hereinafter referred to as "LLMs") has progressed. In implementing such meeting agent technology, it has been proposed to use an interactive system (e.g., Realtime API, etc.) that enables bidirectional real-time communication and voice interaction between a client and a server.
[0004] Japanese Unexamined Patent Application Publication No. 2022-024657
[0005] However, since existing LLM technology is designed for one-to-one interaction, when a single interactive system is applied to a meeting with multiple participants, there arises a problem that the interactive system reacts to voices that are not response requests to the LLM (for example, calls to other participants). Therefore, with meeting agent technology using LLMs, it has been difficult to simultaneously execute immediate response processing that responds to response requests in real time while performing computational processing such as estimating discussion status for facilitation.
[0006] In view of the above, an object of the present disclosure is to enable simultaneous execution of computational processing for meeting facilitation and immediate response processing.
[0007] The information processing device relating to this disclosure comprises a first interactive system and a second interactive system, the first interactive system including a data acquisition unit that acquires discussion utterance data or data relating to a response request, and a generation unit that generates a first prompt for requesting the generation of facilitation advice based on the discussion situation estimated from the discussion utterance data, the second interactive system including a reception unit that receives the response request data or the first prompt from the first interactive system, a content acquisition unit that inputs a second prompt or the first prompt for the response request based on the response request data to a generation AI model, thereby acquiring a response to the response request or the facilitation advice from the generation AI model, and an output unit that outputs the response or the facilitation advice.
[0008] According to this disclosure, computational processing and immediate response processing for meeting facilitation can be performed simultaneously.
[0009] This is a diagram of the entire system including the information processing device. This is a diagram showing an example of a prompt that generates advice regarding facilitation. This is a diagram showing an example of a prompt that is generated when a question is received. This is a flowchart showing the process performed by the interactive system 10A. This is a flowchart showing the process performed by the interactive system 10B. This is a diagram for explaining the process performed by the phase estimation unit. This is a diagram for explaining the calculation of the teamwork level. This is a diagram for explaining an example of the process performed by the decision unit in the divergence phase or clarification phase. This is a diagram showing a modified configuration of the configuration in Figure 1. This is a diagram showing an example of the hardware configuration of the information processing device.
[0010] [Configuration of the Information Processing Device] Figure 1 shows a configuration diagram of System 1, including the Information Processing Device 10 related to this disclosure. As shown in Figure 1, System 1 comprises an external server 20 on which a Large-Scale Language Model (LLM) 20A runs, and the Information Processing Device 10. The Information Processing Device 10 comprises an interactive system 10A that performs computational processing for facilitation, and an interactive system 10B that performs immediate response processing. These interactive systems 10A and 10B may be configured, for example, by a Realtime API. Interactive system 10A corresponds to the "first interactive system," and interactive system 10B corresponds to the "second interactive system." The Information Processing Device 10 has the following configuration in order to realize the functions related to this disclosure.
[0011] The interactive system 10A includes a data acquisition unit 11 that acquires discussion utterance data or data related to response requests (voice data or text data), and a generation unit 12 that generates a first prompt for requesting the generation of facilitation advice based on the discussion status estimated from the discussion utterance data. An example of the above first prompt is shown in Figure 2. Figure 2 shows an example of a prompt when the current discussion phase is in the convergence phase, but the narrowing of opinions is not satisfactory, and therefore the generation of facilitation advice is requested.
[0012] The interactive system 10B includes a reception unit 13 that receives data related to the response request or a first prompt from the interactive system 10A, a content acquisition unit 14 that inputs a second prompt or a first prompt for the response request based on the data related to the response request to the LLM 20A, and obtains a response to the response request or advice regarding facilitation from the LLM 20A, and an output unit 15 that outputs a response or advice regarding facilitation. An example of the second prompt is shown in Figure 3. Figure 3 shows an example of a prompt when requesting an answer to a question in the role of a facilitator.
[0013] While various configurations can be adopted for the generation unit 12 in the interactive system 10A, this embodiment describes an example configuration and process in which the generation unit 12 estimates the discussion phase and generates a prompt requesting appropriate facilitation advice according to the discussion phase.
[0014] As shown in Figure 1, the generation unit 12 includes a speech recognition unit 12A that performs speech recognition processing on the acquired speech, a phase estimation unit 12B that estimates the phase of the discussion based on the spoken content data, a progress calculation unit 12C that calculates an index value for an index that should be obtained as the progress of the discussion according to the estimated phase, a situation estimation unit 12D that estimates the state of the discussion based on the calculated index value, and a determination unit 12E that determines the content of a first prompt requesting the generation of facilitation advice based on the evaluation of the estimated state of the discussion and outputs it to the interactive system 10B. The processing by the generation unit 12 with this configuration will be described in detail later.
[0015] [Regarding the processing performed in the information processing device] The processing performed in the information processing device 10 (processing related to the information processing method of this disclosure) will be described below in accordance with the flowcharts in Figures 4 and 5. For example, when a user (operator) of the information processing device 10 inputs a predetermined execution start command, the processing in Figure 4 is started in the information processing device 10.
[0016] First, the data acquisition unit 11A acquires the speech data of the discussion or the data related to the request for a response (voice data or text data) (step S1 in Figure 4). Here, the speech data of the discussion and the voice data related to the request for a response are acquired by a microphone set up in the conference room, etc. The text data related to the request for a response is acquired by an input device (keyboard, etc.) of the information processing device 10 (not shown).
[0017] Next, the speech recognition unit 12A performs speech recognition processing on the acquired speech data (step S2). For example, it acquires speech data for each participant by identifying the speaker based on existing technology using the acquired speech data, and then acquires speech content data for each participant by converting the acquired speech data for each participant into text data based on existing speech recognition technology.
[0018] Here, the speech recognition unit 12A determines, based on the speech recognition result, whether or not it is a question posted by the owner (data related to a request for an answer) (i.e., whether or not it is speech data for a discussion) (step S3). If the determination is that it is a question posted by the owner, the speech recognition unit 12A transfers the data related to the question posted to the interactive system 10B (step S4).
[0019] On the other hand, if the determination in step S3 is that the question is not submitted by the owner, that is, if it is text data converted from the discussion utterance data (speech content data for each participant), the speech recognition unit 12A transfers the text data (speech content data for each participant) to the phase estimation unit 12B. Thereafter, the phase estimation unit 12B, the progress calculation unit 12C, the situation estimation unit 12D, and the decision unit 12E perform the processing described later to generate a first prompt requesting the generation of facilitation advice (step S5). The processing content in step S5 will be described in detail later. After that, the decision unit 12E transfers the generated first prompt to the interactive system 10B (step S6).
[0020] The interactive system 10B starts executing the process shown in Figure 5 as a trigger when it receives the data related to the question submission transferred in step S4 or the first prompt transferred in step S6.
[0021] First, the reception unit 13 receives data related to the response request or a first prompt from the interactive system 10A (step S11) and transfers it to the content acquisition unit 14. The content acquisition unit 14 determines whether the received data is a question posted by the owner (step S12), and if it is a question posted by the owner, it generates a second prompt for the response request based on the data related to the response request (step S13). On the other hand, if the determination in step S12 is not a question posted by the owner, i.e., if it is a first prompt, step S13 is skipped.
[0022] In the next step S14, the content acquisition unit 14 inputs the received first prompt or the second prompt generated in step S13 to the LLM 20A, thereby querying the LLM 20A for a response to the response request or advice regarding facilitation (step S14).
[0023] Subsequently, the content acquisition unit 14 receives a response from the LLM 20A (a response to a response request or advice regarding facilitation) (step S15), and the output unit 15 outputs the received response (a response to a response request or advice regarding facilitation) (step S16). This allows the user of the information processing device 10 to recognize the response to the response request or advice regarding facilitation. In this case, "output" can take various forms, such as display output, print output, or data transmission to an external location of the information processing device 10.
[0024] Through the above-described embodiment, computational processing for meeting facilitation and real-time response processing can be performed simultaneously. For example, when applied to a group discussion, it is possible to perform processing to estimate the discussion status while simultaneously performing processing to respond to user questions in real time.
[0025] [Details of the prompt generation process] The details of the prompt generation process in step S5 of Figure 4 are described below.
[0026] The phase estimation unit 12B estimates the phase of the discussion based on the speech content data showing the content of each participant's speech (step S5A in Figure 4). Specifically, the phase estimation unit 12B estimates which of the following phases of the discussion is represented, as shown in Figure 6: - Divergent phase, which focuses on generating more ideas; - Convergent phase, which focuses on converging opinions; - Clarification phase, which focuses on analyzing the factors of the problem. As one estimation method, the phase estimation unit 12B may obtain an appropriate phase estimation result from the LLM 20A by inputting a prompt to the LLM 20A instructing it to estimate the phase of the discussion from the content of the speech data.
[0027] Next, the progress calculation unit 12C calculates indicator values for the indicators that should be determined as the progress of the discussion according to the estimated phase (step S5B in Figure 4). Specifically, as shown in Figure 6, if the estimated phase is the clarification phase, the progress calculation unit 12C calculates the number of factors, the number of categories, and the degree of factor abstraction as the progress of the discussion; if the estimated phase is the divergence phase, it calculates the number of ideas, the number of categories, and the degree of idea abstraction as the progress of the discussion; and if the estimated phase is the convergence phase, it calculates the degree of convergence of opinions in the discussion as the progress of the discussion.
[0028] The following is an example of how to calculate the individual indicator values mentioned above that relate to the progress of the discussion.
[0029] Regarding the "number of ideas," the progress calculation unit 12C extracts the ideas contained in the statement content data (i.e., the ideas that have come up in the discussion so far) and calculates the number of extracted ideas as the number of ideas. For example, the progress calculation unit 12C may send a prompt to the LLM 20A to inquire about the number of ideas contained in the statement content data, along with the statement content data itself, and calculate the number of ideas responded to by the LLM as the number of ideas. The "number of factors" may also be calculated using the same method as the "number of ideas" described above.
[0030] Regarding the "number of categories," the progress calculation unit 12C extracts ideas included in the utterance data (i.e., ideas that have come up in the discussion so far), categorizes the extracted ideas according to predetermined criteria, and calculates the number of categories obtained as the number of categories. For example, the phase estimation unit 12B may, similar to the number of ideas described above, send a prompt to the LLM 20A to inquire about the number of categories included in the utterance data, along with the utterance data itself, and calculate the number of categories answered by the LLM as the number of categories. Alternatively, as another method, the number of categories may be calculated by inputting the target utterance data into a machine learning model trained by machine learning, where the utterance data is the explanatory variable and the categories or the number of categories is the dependent variable, and obtaining the output categories or the number of categories.
[0031] Regarding the "idea abstraction level," the progress calculation unit 12C estimates it, for example, by following these steps: (1) The progress calculation unit 12C divides a sentence related to a certain idea, included in the spoken content data of each participant, into multiple words by morphological analysis; (2) It calculates the co-occurrence probability of each divided word with a word list included in a dictionary containing words with multiple meanings for word sense disambiguation (WSD); (3) Then, for each sentence related to the idea included in the spoken content data, it repeats (1) and (2) above to calculate the co-occurrence probability of each word in the entire sentence related to the idea, calculates the average of the obtained co-occurrence probabilities, and calculates the obtained average value as the idea abstraction level for that idea. Note that the "abstraction level" may also be information indicating whether the idea is sufficiently specific (or abstract) in relation to the topic, and specifically may be a numerical value in the range of 0 to 100, or information indicating which level it is among several predetermined levels. Note that the "factor abstraction level" may also be calculated using the same method as the idea abstraction level described above.
[0032] Regarding the "degree of opinion convergence," the progress calculation unit 12C vectorizes the content of the statements using methods such as TF-IDF and BERT embedding, clusters the resulting vectors into multiple clusters using methods such as K-means and Hierarchical Clustering, and calculates the degree of opinion convergence in the discussion based on the reduction in the number of clusters or the bias in the number of clusters. Specifically, the degree of opinion convergence is calculated by following a series of steps including (a) vectorization of the content of the statements, (b) clustering of the vectors, and (c) evaluation of the degree of opinion convergence. The following will explain using a specific example. Here, the following is assumed as an example of discussion data including the agenda and the log of the statements made during the discussion. Topic: "Pricing for the new product" Log of discussion comments: Comment 1: "I think we should set the price low, especially to attract early customers." Comment 2: "I agree with lowering the price, but that could lower the profit margin." Comment 3: "Considering the profit margin, shouldn't we also consider setting the price a little higher?" Comment 4: "As a strategy to attract early customers, how about distributing discount coupons?" Comment 5: "That's a good idea. I think it's smart to attract early customers with discounts."
[0033] In the hypothetical example described above, in "(a) Vectorization of the utterance content," the progress calculation unit 12C first tokenizes the utterance content. For example, tokenizing utterance 1 above yields the following: ["I", "set", "low", "price", "should", "I think", "especially", "to acquire", "initial customers"] Next, the progress calculation unit 12C removes stop words. Here, frequently occurring but low-information words (e.g., "I", "is", "but") are removed, and then the content is vectorized using TF-IDF and an arbitrary machine learning model (e.g., BERT).
[0034] In the next step, "(b) Vector Clustering," the progress calculation unit 12C classifies (clusters) the content of the statements into several clusters using methods such as K-means and Hierarchical Clustering. For example, the content of statements 1 to 5 above is classified (clustered) into the following three clusters. Here, the number of elements is always 2. Cluster 1: "Set price low" "Acquire initial customers" Cluster 2: "Profit margin" "Set price high" Cluster 3: "Discount coupon" "Initial customer strategy"
[0035] In the next step, "(c) Evaluation of the degree of convergence of opinions," the progress calculation unit 12C checks whether the discussion is converging based on whether the number of clusters has decreased or whether there is a large bias in the number of clusters. For example, whether the number of clusters has decreased can be determined by the value of (current number of clusters) / (number of clusters at the start of the discussion), and whether there is a large bias in the number of clusters can be determined by the value of (variance of the number of elements in each cluster at the current time) / (variance of the number of elements in each cluster at the start of the discussion).
[0036] Next, the situation estimation unit 12D estimates the state of the discussion based on the calculated index values (step S5C in Figure 4). Specifically, based on the calculated index values, the situation estimation unit 12D estimates the following scores regarding the state of the discussion: "teamwork level" based on the distribution of participant statements during the discussion, "statement level" indicating the degree to which important statements are made among all participants' statements, and "agenda relevance level" indicating the degree to which statements relevant to the agenda are made within a certain time. In this embodiment, an example is shown in which all three—teamwork level, statement level, and agenda relevance level—are estimated, but at least one of the above three may be estimated. The calculation of the above three scores will be explained below.
[0037] Regarding the "degree of teamwork," the situation estimation unit 12D calculates it using a statistic called the coefficient of variation, based on the distribution of participant statements during the discussion. The higher the value, the greater the degree of teamwork, within the range of 0.0 to 1.0. The coefficient of variation is the value obtained by dividing the standard deviation by the mean. It is a unitless numerical value used to relatively evaluate the variability of data with different units, and the relationship between data and variability relative to the mean.
[0038] As shown in Figure 7, (i) The situation estimation unit 12D calculates the average number of times each participant has spoken. If participant A has spoken 3 times, participant B has spoken 8 times, and participant C has spoken 7 times, the average number of times everyone has spoken is calculated to be "6 times". (ii) The situation estimation unit 12D calculates the standard deviation σ from the variance of each participant's number of speaking times relative to the calculated average. In the example in Figure 7, the variance of each participant's number of speaking times relative to the average is "4.666...", so the standard deviation σ is calculated to be "2.160...". (iii) Then, the situation estimation unit 12D calculates the coefficient of variation using the following formula (1), and as shown in Figure 7, "36%" is obtained. Coefficient of variation = (standard deviation / mean) (1) (iv) Furthermore, the situation estimation unit 12D calculates the degree of teamwork using the following formula (2), and as shown in Figure 7, "64%" is obtained. Teamwork score = 100 - coefficient of variation (2) If the calculated value exceeds 100%, the teamwork score is set to 100%, and if the calculated value is less than 0%, the teamwork score is set to 0%.
[0039] Next, regarding "speaker level," the situation estimation unit 12D calculates the speaker level, which indicates the degree to which each participant's statement is important, using the following procedure: (i) As a preprocessing step, the situation estimation unit 12D performs tokenization (word segmentation) and normalization (unification of character types, case conversion) for each participant's statement. (ii) Then, the situation estimation unit 12D calculates the speaker level for each participant's statement from the information obtained in (i) above. As for the calculation method here, existing methods for calculating the importance of words appearing in a document, such as TF-IDF (Term Frequency-Inverse Document Frequency) and Okapi BM25, may be used.
[0040] Furthermore, the "relevance of the agenda" is calculated by the situation estimation unit 12D using, for example, the following first method, second method, etc.
[0041] In the first method, the situation estimation unit 12D determines whether or not there have been any statements related to the purpose of the agenda within a certain period of time. For example, the situation estimation unit 12D may provide the LLM 20A with statement content data and information on the purpose of the agenda, and send a prompt instructing it to "determine whether or not the various statements that can be grasped from the statement content data are related to the purpose of the agenda, and to determine the degree of relevance of the statements to the purpose of the agenda in the overall discussion within a predetermined numerical range (for example, a range of 0.0 to 1.0: the higher the number, the higher the degree of relevance)," and estimate the degree of relevance answered by the LLM 20A as the agenda relevance.
[0042] In the second method, the situation estimation unit 12D determines whether there are any statements related to the agenda within a certain time. Specifically, the situation estimation unit 12D (i) vectorizes the agenda and each statement, (ii) calculates the similarity (e.g., cosine similarity) between the vectorized agenda and each vectorized statement, and (iii) calculates the minimum value among the multiple similarities calculated as the agenda relevance. It should be noted that using the minimum value as described above has the advantage of being able to confirm whether a certain theme relevance is maintained as a whole statement, even if some statements deviate from the topic.
[0043] Hereinafter, a specific example of the second method will be described. Here, the following is assumed as an example of discussion data including an agenda and discussion utterance logs. Agenda: "On the progress of AI technology and its social impact" Discussion utterance log: Utterance a: "AI technology is progressing rapidly." Utterance b: "The issue of data privacy is becoming increasingly important." Utterance c: "Technological development of autonomous vehicles is progressing." Utterance d: "New regulations for privacy protection are required."
[0044] The situation estimation unit 12D vectorizes the above agenda and each utterance, and calculates the cosine similarity between the vectorized agenda and each vectorized utterance using existing technology. Assume that this yields the following cosine similarity values. Similarity for utterance a: 0.777764051839938 Similarity for utterance b: 0.7738009682636184 Similarity for utterance c: 0.7976624298419579 Similarity for utterance d: 0.7404025807062681 Here, the situation estimation unit 12D calculates the minimum value among the plurality of calculated similarity values as the agenda relevance. In the above example, since the similarity for utterance d is the minimum value, 74% is obtained as the agenda relevance.
[0045] Next, the determination unit 12E evaluates the discussion situation based on the estimated score relating to the discussion situation (step S5D in FIG. 4), and generates a prompt requesting facilitation advice based on the evaluation result (step S5E). Since the processing of these steps S5D and S5E differs depending on the discussion phase, they will be described individually below.
[0046] If the discussion phase is in the convergence phase, the decision unit 12E uses the degree of opinion convergence A1 calculated by the progress calculation unit 12C, the degree of teamwork A2 estimated by the situation estimation unit 12D, and the degree of topic relevance A3 to calculate the progress of convergence B1 (B1=A1), the degree of teamwork B2 (B2=A2), and the degree of topic relevance B3 (B3=A3) as evaluation values for evaluating the discussion situation. The thresholds t1 to t3 for evaluating the above evaluation values B1 to B3 are set in advance using training data, etc. The decision unit 12E evaluates whether the discussion situation is appropriate in terms of each evaluation value B1 to B3, depending on whether each evaluation value B1 to B3 exceeds the corresponding threshold (step S5D). At this time, evaluation values that do not exceed the corresponding threshold are selected as "missing features".
[0047] Furthermore, the decision unit 12E generates a prompt for the selected evaluation value to request facilitation advice, for example, as follows (step S5E): If the missing feature is "Convergence Progress B1", it presents a new perspective for selection and asks a question to solicit opinions on the multiple options that have gathered votes from that perspective. If the missing feature is "Teamwork B2", it asks a question to check if there are any points to notice based on the opinions of others. If the missing feature is "Agenda Relevance B3", it presents comments to summarize the opinions and ideas that have been expressed so far. If there are multiple evaluation values selected as "missing features", one evaluation value should be selected according to a predetermined priority order (for example, evaluation value B3 → B2 → B1 priority order).
[0048] On the other hand, when the discussion phase is a divergence phase or an elucidation phase, the determining unit 12E, based on scores indicating the degree related to the estimated discussion situation (idea abstraction degree X1, number of ideas X2, number of categories X3, teamwork degree X4, and agenda relevance degree X5), obtains, as evaluation values: (a) idea concreteness Y1 obtained by evaluating the concreteness of ideas, (b) discussion comprehensiveness Y2 obtained by evaluating the comprehensiveness of discussion, (c) satisfaction degree of all participants Y3 obtained by evaluating the satisfaction degree of participants, and (d) agenda relevance degree Y4 obtained by evaluating the agenda relevance degree; and evaluates whether the discussion is in an appropriate situation from the viewpoint related to each evaluation value depending on whether the obtained four evaluation values Y1 to Y4 exceed a predetermined threshold for each evaluation value (step S5D). At this time, an evaluation value that does not exceed the corresponding threshold is selected as a "missing feature".
[0049] Further, the determining unit 12E generates a prompt requesting facilitation advice according to one evaluation value selected as a missing feature (step S5E). If there are a plurality of selected evaluation values, the determining unit 12E selects one evaluation value based on a predetermined priority order, and generates a prompt requesting facilitation advice according to the selected one evaluation value. Hereinafter, a processing example of steps S5D and S5E when the discussion phase is a divergence phase or an elucidation phase will be described.
[0050] As an example, the determining unit 12E calculates (a) the evaluation value "idea concreteness Y1" obtained by evaluating the concreteness of ideas according to the following formula (3). Y1=1-X1 (3) As shown in Fig. 8, when X1=0.7, X2=3, X3=12, X4=0.8, and X5=0.7, idea concreteness Y1=1-0.7=0.3 is calculated.
[0051] Further, the determining unit 12E calculates (b) the evaluation value "discussion comprehensiveness Y2" obtained by evaluating the comprehensiveness of discussion according to the following formula (4). Y2=X2+X3 (4) From the numerical example (X1 to X5) in Fig. 8 described above, discussion comprehensiveness Y2=3+12=15 is calculated. Note that the above calculation method (addition of X2 and X3) is an example, and another calculation method such as multiplication of X2 and X3 may be adopted.
[0052] Furthermore, the decision unit 12E adopts the teamwork score X4 obtained in step S5D as the evaluation value "total satisfaction Y3" which evaluates the level of satisfaction of the participants. That is, Y3 = X4 = 0.8.
[0053] Similarly, the decision unit 12E adopts the agenda relevance X5 obtained in step S5D as the evaluation value "Agenda Relevance Y4" which evaluates the degree of agenda relevance. That is, Y4 = X5 = 0.7.
[0054] The decision unit 12E then determines whether each of the four evaluation values Y1 to Y4 exceeds the corresponding threshold t1 to t4, and selects the evaluation values that did not exceed the threshold as "missing features". Here, as shown in Figure 8, the threshold t1 = 0.5 for evaluation value Y1, the threshold t2 = 20 for evaluation value Y2, the threshold t3 = 0.7 for evaluation value Y3, and the threshold t4 = 0.5 for evaluation value Y4 are predetermined by prior machine learning or the like.
[0055] Furthermore, the decision unit 12E generates prompts for facilitation advice according to the selected evaluation value, i.e., the "missing feature," for example, as follows: • If the missing feature is Idea Specificity Y1, it generates a prompt for facilitation advice to encourage deeper exploration of the ideas that have emerged from the discussion so far. • If the missing feature is Discussion Coverage Y2, it generates a prompt for facilitation advice to present new perspectives and encourage the generation of new ideas. • If the missing feature is Everyone's Agreement Y3, it generates a prompt for facilitation advice to check if there are any points to notice based on the opinions of others. • If the missing feature is Agenda Relevance Y4, it generates a prompt for facilitation advice to summarize the opinions and ideas that have emerged so far.
[0056] In the numerical example in Figure 8, evaluation values Y1 and Y2 are selected as "missing features" because they did not exceed their respective thresholds t1 and t2. If multiple evaluation values are selected in this way, evaluation value Y1 is selected because it has a higher priority than evaluation value Y2, according to a predetermined priority order (for example, evaluation value Y4 → Y1 → Y3 → Y2).
[0057] The "priority order of evaluation values Y4 → Y1 → Y3 → Y2" in the above example follows the following principles: ・For topic relevance Y4, the first priority was to produce an output related to the topic. The idea was that it is important to produce as much output as possible, regardless of its nature. ・For idea specificity Y1, the second priority was that the output be specific. The idea was that even if the output focuses on only one perspective or is biased towards someone's opinion, the more specific the output, the easier it is to move the discussion forward. ・For everyone's satisfaction Y3, the third priority was that all participants be satisfied. Even if the output focuses on only one perspective or the idea is excellent, since discussion is a collaborative effort involving multiple participants, it is desirable that the output be as satisfactory as possible for all participants, without being biased towards anyone's opinion. ・For comprehensiveness of discussion Y2, the fourth priority was that the discussion be comprehensive. The idea was that the more diverse the perspectives considered, the better the idea.
[0058] [Variations of System 1] System 1 is not limited to the configuration shown in Figure 1 above, and may also be configured in which the LLM20A is implemented in the information processing device 10, as shown in Figure 9. As mentioned above, the information processing device 10 may be a mobile terminal such as a smartphone, mobile phone, smartwatch, or wearable device, and the LLM20A may be implemented in such a mobile terminal. This configuration can be realized by installing an application that performs the functions of the LLM20A in the information processing device 10.
[0059] The gist of this disclosure is found in the following [1] to [8]. [1] An information processing device comprising a first interactive system and a second interactive system, the first interactive system including: a data acquisition unit that acquires discussion utterance data or data relating to a response request; a generation unit that generates a first prompt for requesting the generation of facilitation advice based on the discussion situation estimated from the discussion utterance data; the second interactive system including: a reception unit that receives data relating to the response request or the first prompt from the first interactive system; a content acquisition unit that inputs a second prompt or the first prompt for the response request based on the data relating to the response request to a generation AI model, thereby acquiring a response to the response request or the facilitation advice from the generation AI model; and an output unit that outputs the response or the facilitation advice. [2] The information processing device according to [1], wherein, when both the data relating to the response request and the first prompt are received, the content acquisition unit first inputs the second prompt to the generating AI model to obtain a response to the response request, and then inputs the first prompt to the generating AI model to obtain advice regarding the facilitation. [3] The information processing device according to [1] or [2], wherein, when the data relating to the response request is audio data, the generation unit converts the audio data into text data by speech recognition, and the receiving unit accepts the converted text data as data relating to the response request. [4] The information processing device according to any one of [1] to [3], wherein the data acquisition unit acquires the data relating to the response request when a predetermined operation is performed. [5] The information processing device according to any one of [1] to [4], wherein the information processing device is comprised of a terminal carried by one of the participants in the discussion.[6] An information processing method comprising: a first interactive system provided in an information processing device acquires utterance data of a discussion or data relating to a request for a response; the first interactive system generates a first prompt for requesting the generation of facilitation advice based on the discussion situation estimated from the utterance data of the discussion; a second interactive system provided in the information processing device receives data relating to the request for a response or the first prompt from the first interactive system; the second interactive system inputs a second prompt for the request for a response or the first prompt based on the data relating to the request for a response to a generation AI model, thereby acquiring a response to the request for a response or advice relating to facilitation from the generation AI model; and the second interactive system outputs the response or advice relating to facilitation.
[0060] [Explanation of terms, explanation of hardware configuration (Figure 10), etc.] The block diagram used in the description of the above embodiment shows functional units. These functional blocks (components) are realized by any combination of at least one of hardware and software. Furthermore, the method of realizing each functional block is not particularly limited. That is, each functional block may be realized using one device that is physically or logically coupled, or it may be realized using two or more physically or logically separated devices that are directly or indirectly connected (for example, using wired, wireless, etc.). A functional block may be realized by combining the above one device or the above multiple devices with software.
[0061] Functions include, but are not limited to, judgment, decision, determination, calculation, calculation, processing, derivation, investigation, exploration, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, assumption, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating (mapping), and assigning. For example, a functional block (configuration part) that enables transmission is called a transmitting unit or transmitter. In all cases, as mentioned above, the method of implementation is not particularly limited.
[0062] For example, the information processing device in one embodiment of the present disclosure may function as a computer that performs processing of the information processing method of the present disclosure. Figure 10 is a diagram showing an example of the hardware configuration of the information processing device 10 according to one embodiment of the present disclosure. The above-described information processing device 10 may be physically configured as a computer device including a processor 1001, memory 1002, storage 1003, communication device 1004, input device 1005, output device 1006, bus 1007, etc.
[0063] In the following explanation, the term "device" can be replaced with "circuit," "device," "unit," etc. The hardware configuration of the information processing device 10 may include one or more of the devices shown in the figure, or it may be configured to omit some of the devices.
[0064] Each function in the information processing device 10 is realized by loading predetermined software (programs) onto hardware such as the processor 1001 and memory 1002, which allows the processor 1001 to perform calculations, control communication by the communication device 1004, and control at least one of data reading and writing in the memory 1002 and storage 1003.
[0065] The processor 1001 controls the entire computer, for example, by running an operating system. The processor 1001 may be composed of a central processing unit (CPU) that includes interfaces with peripheral devices, control units, arithmetic units, registers, etc.
[0066] Furthermore, the processor 1001 reads programs (program code), software modules, data, etc., from at least one of the storage 1003 and the communication device 1004 into the memory 1002 and executes various processes accordingly. The program used is one that causes the computer to execute at least a part of the operations described in the above embodiment. For example, the data acquisition unit 11 of the information processing device 10 may be implemented by a control program stored in the memory 1002 and operated on the processor 1001, and other functional blocks may be implemented similarly. The above-described various processes have been explained as being executed by one processor 1001, but they may be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. The program may also be transmitted from a network via a telecommunications line.
[0067] The memory 1002 is a computer-readable recording medium and may consist of at least one of the following: ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), etc. The memory 1002 may also be called a register, cache, main memory, etc. The memory 1002 can store executable programs (program code), software modules, etc., for carrying out a wireless communication method according to one embodiment of the present disclosure.
[0068] The storage 1003 is a computer-readable recording medium and may consist of at least one of the following: an optical disc such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disc, a digital multipurpose disc, a Blu-ray® disc), a smart card, flash memory (e.g., a card, a stick, a key drive), a floppy® disk, a magnetic strip, etc. The storage 1003 may also be called an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, server, or other suitable medium including at least one of the memory 1002 and the storage 1003.
[0069] The communication device 1004 is hardware (transceiver / receiver device) for communicating between computers via at least one of a wired network and a wireless network, and is also referred to as a network device, network controller, network card, communication module, etc. The communication device 1004 may be configured to include, for example, a high-frequency switch, duplexer, filter, frequency synthesizer, etc., in order to implement at least one of frequency division duplex (FDD) and time division duplex (TDD).
[0070] The input device 1005 is an input device that accepts input from an external source (e.g., a keyboard, mouse, microphone, switch, button, sensor, etc.). The output device 1006 is an output device that outputs to an external source (e.g., a display, speaker, LED lamp, etc.). The input device 1005 and the output device 1006 may be configured as an integrated unit (e.g., a touch panel).
[0071] Furthermore, each device, such as the processor 1001 and memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or different buses may be configured for each device.
[0072] Furthermore, the information processing device 10 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), and an FPGA (Field Programmable Gate Array), and some or all of each functional block may be realized by such hardware. For example, the processor 1001 may be implemented using at least one of these hardware components.
[0073] The notification of information is not limited to the embodiments described herein and may be carried out by other means. For example, the notification of information may be carried out by physical layer signaling (e.g., DCI (Downlink Control Information), UCI (Uplink Control Information)), upper layer signaling (e.g., RRC (Radio Resource Control) signaling, MAC (Medium Access Control) signaling, broadcast information (MIB (Master Information Block), SIB (System Information Block))), other signals, or combinations thereof. RRC signaling may also be called RRC messages, and may be, for example, RRC Connection Setup messages, RRC Connection Reconfiguration messages, etc.
[0074] Each aspect / embodiment described in this disclosure refers to LTE (Long Term Evolution), LTE-A (LTE-Advanced), SUPER 3G, IMT-Advanced, 4G (4th generation mobile communication system), 5G (5th generation mobile communication system), 6th generation mobile communication system (6G), xth generation mobile communication system (xG) (xG (where x is, for example, an integer or decimal)), FRA (Future Radio Access), NR (new Radio), New radio access (NX), Future generation radio access (FX), W-CDMA (registered trademark), GSM (registered trademark), CDMA2000, UMB (Ultra Mobile Broadband), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20 may apply to at least one system utilizing UWB (Ultra-WideBand), Bluetooth®, or other appropriate systems, and to next-generation systems extended, modified, created, or defined based thereon. Alternatively, multiple systems may be applied in combination (e.g., a combination of at least one of LTE and LTE-A with 5G).
[0075] The processing procedures, sequences, flowcharts, etc., of each aspect / embodiment described in this disclosure may be reordered, provided they do not contradict each other. For example, the methods described in this disclosure present various step elements using exemplary order and are not limited to the specific order presented.
[0076] The specific operations described in this disclosure as being performed by a base station may, in some cases, be performed by its upper node. In a network consisting of one or more network nodes having a base station, it is clear that various operations performed for communication with a terminal can be performed by the base station and at least one other network node (for example, an MME or S-GW, but not limited to these). Although the above example illustrates the case where there is one other network node besides the base station, it may also be a combination of multiple other network nodes (for example, an MME and an S-GW).
[0077] Information can be output from a higher layer (or lower layer) to a lower layer (or higher layer). Input and output may also occur via multiple network nodes.
[0078] Input and output information may be stored in a specific location (e.g., memory) or managed using a management table. Input and output information may be overwritten, updated, or appended to. Output information may be deleted. Input information may be sent to other devices.
[0079] The determination may be made by a value represented by one bit (0 or 1), by a boolean value (true or false), or by a numerical comparison (for example, a comparison with a predetermined value).
[0080] Each aspect / embodiment described in this disclosure may be used individually, in combination, or switched between as needed during implementation. Furthermore, notification of specific information (e.g., notification that "X is") is not limited to explicit notification, but may also be implicit (e.g., by not providing such notification).
[0081] Although the present disclosure has been described in detail above, it will be clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the intent and scope of the present disclosure as defined by the claims. Therefore, the descriptions in the present disclosure are illustrative and not intended to be restrictive in any way.
[0082] Software should be broadly interpreted to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, and so on, whether they are called software, firmware, middleware, microcode, hardware description languages, or by any other name.
[0083] Furthermore, software, instructions, information, etc., may be transmitted and received via a transmission medium. For example, if software is transmitted from a website, server, or other remote source using at least one of wired technology (such as coaxial cable, fiber optic cable, twisted pair, or digital subscriber line (DSL)) and wireless technology (such as infrared or microwave), then at least one of these wired and wireless technologies is included in the definition of a transmission medium.
[0084] The information, signals, etc. described in this disclosure may be represented using any of the various different techniques. For example, the data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
[0085] In addition, terms used in this disclosure and terms necessary for understanding this disclosure may be replaced with terms having the same or similar meanings. For example, at least one of the channel and symbol may be a signal (signaling). Also, a signal may be a message. Furthermore, a component carrier (CC) may be called a carrier frequency, cell, frequency carrier, etc.
[0086] The terms “system” and “network” as used in this disclosure are interchangeable.
[0087] Furthermore, the information, parameters, etc., described in this disclosure may be expressed using absolute values, relative values from a given value, or other corresponding information. For example, wireless resources may be indicated by an index.
[0088] The names used for the parameters described above are not restrictive in any way. Furthermore, mathematical formulas and other expressions using these parameters may differ from those expressly disclosed in this disclosure. Various channels (e.g., PUCCH, PDCCH, etc.) and information elements can be identified by any suitable name, and therefore, the various names assigned to these various channels and information elements are not restrictive in any way.
[0089] In this disclosure, terms such as “Base Station (BS),” “wireless base station,” “fixed station,” “NodeB,” “eNodeB (eNB),” “gNodeB (gNB),” “access point,” “transmission point,” “reception point,” “transmission / reception point,” “cell,” “sector,” “cell group,” “carrier,” and “component carrier” may be used interchangeably. Base stations may also be referred to by terms such as macrocell, small cell, femtocell, and picocell.
[0090] A base station can accommodate one or more (e.g., three) cells. If a base station accommodates multiple cells, the entire coverage area of the base station can be divided into multiple smaller areas, each of which may also be provided with communication services by a base station subsystem (e.g., a Remote Radio Head (RRH)). The terms “cell” or “sector” refer to part or all of the coverage area of at least one of the base station and / or base station subsystems that provide communication services in that coverage.
[0091] In this disclosure, the transmission of information by a base station to a terminal may be interpreted as the base station instructing the terminal to perform control or operation based on the information.
[0092] In this disclosure, terms such as "Mobile Station (MS)," "user terminal," "User Equipment (UE)," and "terminal" may be used interchangeably.
[0093] A mobile station may also be referred to by those skilled in the art as a subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, or several other appropriate terms.
[0094] At least one of the base station and the mobile station may be called a transmitting device, a receiving device, a communication device, etc. At least one of the base station and the mobile station may be a device mounted on a mobile body, the mobile body itself, etc.
[0095] The term "moving object" refers to any object that can move, regardless of its speed. This also includes cases where the moving object is stationary. The term "moving object" includes, but is not limited to, vehicles, transport vehicles, automobiles, motorcycles, bicycles, connected cars, excavators, bulldozers, wheel loaders, dump trucks, forklifts, trains, buses, handcarts, rickshaws, ships and other watercraft, airplanes, rockets, satellites, drones (registered trademarks), multicopters, quadcopters, balloons, and anything carried on them.
[0096] Furthermore, the mobile entity may be one that autonomously drives based on operational commands. It may be a vehicle (e.g., a car, an airplane), an unmanned mobile entity (e.g., a drone, an autonomous vehicle), or a robot (manned or unmanned). Note that at least one of the base station and the mobile station may be a device that does not necessarily move during communication operations. For example, at least one of the base station and the mobile station may be an IoT (Internet of Things) device such as a sensor.
[0097] Furthermore, the term "base station" in this disclosure may be interpreted as "user terminal." For example, the various aspects / embodiments of this disclosure may be applied to a configuration in which communication between a base station and a user terminal is replaced with communication between multiple user terminals (which may be called, for example, D2D (Device-to-Device), V2X (Vehicle-to-Everything), etc.). Also, terms such as "uplink" and "downlink" may be interpreted as terms corresponding to terminal-to-terminal communication (for example, "side"). For example, uplink channel, downlink channel, etc., may be interpreted as side channel.
[0098] As used in this disclosure, the terms “determining” and “determining” may encompass a wide variety of actions. “Determining” may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, or inquiring (e.g., searching in a table, database, or other data structure), or ascertaining. “Determining” may also include receiving (e.g., receiving information), transmitting (e.g., sending information), inputting, outputting, or accessing (e.g., accessing data in memory). Furthermore, "judgment" and "decision" can include considering something as having been "judged" or "decided" after resolving, selecting, choosing, establishing, comparing, etc. In other words, "judgment" and "decision" can include considering something as having been "judged" or "decided" after some action. Also, "judgment (decision)" can be reinterpreted as "assuming," "expecting," or "considering."
[0099] The terms “connected,” “coupled,” or any variation thereof, mean any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are “connected” or “coupled” with each other. The coupling or connection between elements may be physical, logical, or a combination thereof. For example, “connection” may be reinterpreted as “access.” As used in this disclosure, two elements may be considered to be “connected” or “coupled” with each other using at least one of one or more wires, cables, and printed electrical connections, and, in some non-limiting and non-exclusive examples, electromagnetic energy having wavelengths in the radio frequency domain, microwave domain, and optical (both visible and invisible) domain.
[0100] In this disclosure, the phrase "based on" does not mean "based solely on" unless otherwise specified. In other words, the phrase "based on" means both "based solely on" and "based at least on."
[0101] Any reference to elements using the designations “first,” “second,” etc., as used in this disclosure does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient way to distinguish between two or more elements. Accordingly, references to the first and second elements do not imply that only two elements may be employed, or that the first element must precede the second element in any way.
[0102] In the configuration of each of the above devices, "means" may be replaced with "part," "circuit," "device," etc.
[0103] Where the terms “include,” “including,” and variations thereof are used in this disclosure, these terms are intended to be inclusive, as is the term “comprising.” Furthermore, the term “or” as used in this disclosure is not intended to mean exclusive OR.
[0104] In this disclosure, if articles are added through translation, such as a, an, and the in English, this disclosure may include the fact that the noun following these articles is plural.
[0105] In this disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "combine" may be interpreted similarly to "different."
[0106] 1...System, 10...Information processing device, 10A, 10B...Interactive system, 11...Data acquisition unit, 12...Generation unit, 12A...Speech recognition unit, 12B...Phase estimation unit, 12C...Progress calculation unit, 12D...Status estimation unit, 12E...Decision unit, 13...Reception unit, 14...Content acquisition unit, 15...Output unit, 20...External server, 20A...LLM, 1001...Processor, 1002...Memory, 1003...Storage, 1004...Communication device, 1005...Input device, 1006...Output device, 1007...Bus.
Claims
1. An information processing device comprising a first interactive system and a second interactive system, wherein the first interactive system includes a data acquisition unit that acquires discussion utterance data or data relating to a response request, and a generation unit that generates a first prompt for requesting the generation of facilitation advice based on the discussion situation estimated from the discussion utterance data, and the second interactive system includes a reception unit that receives data relating to the response request or the first prompt from the first interactive system, a content acquisition unit that inputs a second prompt or the first prompt for the response request based on the data relating to the response request to a generation AI model, thereby acquiring a response to the response request or the facilitation advice from the generation AI model, and an output unit that outputs the response or the facilitation advice.
2. The information processing apparatus according to claim 1, wherein, when both the data relating to the response request and the first prompt are received, the content acquisition unit first inputs the second prompt to the generating AI model to obtain a response to the response request, and then inputs the first prompt to the generating AI model to obtain advice regarding the facilitation.
3. If the data relating to the response request is audio data, the generation unit converts the audio data into text data by speech recognition, and the receiving unit receives the converted text data as the data relating to the response request, as described in claim 1.
4. The information processing apparatus according to claim 1, wherein the data acquisition unit acquires data relating to the response request when a predetermined operation is performed.
5. The information processing device according to claim 1, wherein the information processing device is comprised of a terminal carried by one of the participants in the discussion.
6. An information processing method comprising: a first interactive system provided by an information processing device acquires utterance data of a discussion or data relating to a request for a response; the first interactive system generates a first prompt for requesting the generation of facilitation advice based on the discussion situation estimated from the utterance data of the discussion; a second interactive system provided by the information processing device receives data relating to the request for a response or the first prompt from the first interactive system; the second interactive system inputs a second prompt for the request for a response or the first prompt based on the data relating to the request for a response to a generation AI model, thereby acquiring a response to the request for a response or advice relating to facilitation from the generation AI model; and the second interactive system outputs the response or advice relating to facilitation.