Support apparatus
The dialogue progress support device addresses the lack of specific guidance and high power consumption in dialogue support by preprocessing text data to extract relevant units and using a large language model, enabling efficient and targeted question generation.
Patent Information
- Application Number
- JP2024001195
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-07-22
AI Technical Summary
Existing dialogue support technologies, such as those described in Patent Document 1, only provide general notifications of dialogue issues without specific guidance, and the processing required for advanced dialogue techniques like the SPIN method consumes significant power due to high computational complexity.
A dialogue progress support device that preprocesses text data to extract processing units of a given length, removes irrelevant parts, and uses a large language model to generate questions related to the gist of the dialogue, thereby reducing computational complexity and power consumption.
The device effectively supports dialogue progress by generating targeted questions and providing specific guidance while significantly reducing power consumption through efficient processing.
Smart Images

Figure 2025107774000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a device for assisting the progress of a dialogue.
Background Art
[0002] Various dialogues exemplified by sales activities, various meetings, inquiry responses, etc. are being conducted at business sites. Dialogues at business sites have various purposes. By communicating the matters to be conveyed, the purpose of the dialogue is further achieved. Therefore, there is a demand for a technology that supports the achievement of the purpose of the dialogue.
[0003] Regarding the technology for supporting the achievement of the purpose of a dialogue, Patent Document 1 discloses a meeting management device having a voice recognition unit, an output unit, a keyword storage unit, a keyword extraction unit, a tabulation unit, and the like. In the meeting management device of Patent Document 1, positive keywords and the like defined as representing the smooth progress of a meeting and negative keywords and the like defined as representing a troubled progress are stored in advance in the keyword storage unit. Then, the tabulation unit of Patent Document 1 classifies and tabulates the phrases extracted by the keyword extraction unit from the speech analyzed by the voice recognition unit, and determines whether the progress of the meeting is smooth based on the tabulation. And the output unit of Patent Document 1 outputs the determination to an external terminal. With the above configuration and the like, Patent Document 1 can notify the participants to that effect when it is determined from the content of the speech that the meeting is troubled or off-topic, and promote cooperation for the smooth progress of the meeting.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, the technology of Patent Document 1 only notifies that a meeting is in dispute or derailed, and there is room for further improvement in terms of providing specific guidance to support the achievement of the purpose of the dialogue. Specific guidance is exemplified by, for example, solutions to problems (solutions), marketing strategies, agency policies, sales policies, and marketing measures, sales measures, and sales-related education such as landing pages.
[0006] When achieving the purpose of a sales-related dialogue, a talk script using various dialogue techniques such as the SPIN method is beneficial. However, since the dialogue using dialogue techniques is sophisticated, the generation of such a talk script is not easy. By the way, processing related to dialogue using a large language model trained with a large amount of text data is being carried out. Such processing can realize advanced language processing such as generating a talk script along the SPIN method. However, since the processing using a large language model is complex and sophisticated, it can consume a large amount of power.
[0007] The present invention has been made in view of such circumstances. The object of the present invention is to achieve both supporting the achievement of the purpose of the dialogue and suppressing the power consumption in such support.
Means for Solving the Problem
[0008] As a result of intensive studies to solve the above problems, the present inventors found that the above object can be achieved by performing preprocessing to reduce the amount of input data and then performing processing using a large language model when generating questions related to the above support. And the present inventors have completed the present invention. Specifically, the present invention provides the following.
[0009] The present invention provides a dialogue progress support device, comprising: a processing unit extraction unit that extracts a processing unit of a given length from text data of a dialogue; and a question generation unit that generates, by processing using a large language model including the processing unit as an input, a question regarding one speaker related to the processing unit, wherein the processing unit extraction unit removes a portion deviating from the gist of the utterance from the processing unit, and the question generation unit generates a question that prompts presentation of information to be provided from the one speaker to another speaker regarding the gist of the utterance included in the processing unit.
[0010] The present invention preprocesses text data by a processing unit extraction unit. Thereafter, the present invention performs processing using a large language model with the preprocessed text data (processing unit) as an input by a question generation unit. Through this processing, the present invention generates a question that prompts presentation of information to be provided from one speaker to another speaker regarding the gist of the utterance. As a result, the speaker supported by the present invention can convey matters to be conveyed in the dialogue in a natural form of an answer to the generated question. By repeating such questions, the present invention can support obtaining solutions to initial problems, marketing measures, sales measures, and sales-related education.
[0011] Incidentally, the computational complexity of a large language model usually increases in proportion to the square of the length of the input. Therefore, if the entire text data generated from a long dialogue is input to the large language model, the computational complexity becomes enormous. When the computational complexity becomes enormous, the power consumption in the computer becomes enormous. Due to sustainable development goals (SDGs) and the like, energy efficiency is also required in the computer. Therefore, consuming enormous power in the computer is not preferable from the perspective of energy efficiency. To suppress power consumption, the present invention extracts a processing unit of a given length from the text data by the processing unit extraction unit. Thereby, the present invention reduces the computational complexity related to the large language model compared to the case where the entire text data is used as an input. The following is an explanation of how the present invention reduces the computational complexity related to the large language model.
[0012] In the following description, let the length of the text data to be processed be L, and the given length be C. When the entire text data is used as the input, the computational complexity in a large language model is proportional to the square of the length of the input, i.e., L squared. Therefore, when the entire text data is used as the input, the order of the computational complexity is O(L 2 ^2). Thus, in such a process, when the length L of the entire text data increases, the computational complexity becomes extremely large, proportional to the square of the length L. On the other hand, when the processing unit of the present invention is used as the input, the computational complexity per operation in a large language model is proportional to the square of the length of the input, i.e., C squared. That is, the order of the computational complexity per operation is a constant computational complexity determined by the given length C, O(C 2 ^2). Since the entire text data is divided into the given length, such a process is executed L / C (rounded up) times. Here, C 2 ×L / C = CL. Therefore, when processing the entire text data with the processing unit as the input, the order of the computational complexity is O(CL). Thus, the present invention can suppress the computational complexity when the length L of the entire text data increases to the minimum computational complexity proportional to the length L.
[0013] By the way, the dialogue data may include parts that deviate from the gist of the utterance. Such parts are, for example, unnecessary characters such as input mistakes or line breaks and blank characters generated by speech recognition processing, interjections, speech hesitations, repeated utterances, etc. The processing unit extraction unit of the present invention removes such parts from the processing unit. Thereby, the present invention makes the length of the input in a large language model even shorter. Therefore, the present invention further reduces the computational complexity related to the large language model.
[0014] With the above configuration, the present invention significantly reduces the power consumption in the process of assisting the progress of the dialogue. Therefore, the present invention can achieve both assisting in achieving the purpose of the dialogue and suppressing the power consumption in such assistance.
Advantages of the Invention
[0015] The present invention can achieve both supporting the achievement of the purpose of dialogue and suppressing power consumption in such support.
Brief Description of the Drawings
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Embodiments for Carrying Out the Invention
[0017] The following is a detailed description of an example of an embodiment of the present invention with reference to the drawings.
[0018] <System S> FIG. 1 is a block diagram showing the hardware configuration and software configuration of the system S of the present embodiment. The following is an explanation using FIG. 1 regarding an example of a preferable aspect of the hardware configuration and software configuration in the system S of the present embodiment. The system S includes a support device 1 and a terminal T configured to be communicable with each other via a network N and configured to support the progress of dialogue.
[0019] 〔Regarding Dialogue〕 The "dialogue" in the present embodiment is not particularly limited. The "dialogue" in the present embodiment includes various dialogues exemplified by, for example, sales activities, various meetings, inquiry responses, and the like.
[0020] 〔Support Device 1〕 The support device 1 includes a control unit 11, a storage unit 12, a communication unit 13, and the like. The type of the support device 1 is not particularly limited. Examples of such types include a server, a cloud server, and the like. The cloud server is, for example, Amazon Web Services (AWS, registered trademark), etc. When the type of the support device 1 is AWS, the support process described later is realized using a service that automatically allocates computing resources such as AWS Lambda. Thereby, an administrator or the like of the support device 1 can realize the support process described later without managing the computing resources by themselves.
[0021] [Control Unit 11] The control unit 11 includes a CPU (Central Processing Unit), a RAM (Random Access Memory), a ROM (Read Only Memory), and the like.
[0022] The control unit 11 cooperates with the storage unit 12 and / or the communication unit 13 as necessary. And the control unit 11 realizes the software components of the program of the present embodiment executed by the support device 1. The software components include a data acquisition unit 111, a speech recognition unit 112, a processing unit extraction unit 113, a question generation unit 114, a purpose acquisition unit 115, an achievement degree evaluation unit 116, an advice set generation unit 117, an output unit 118, and the like.
[0023] [Storage Unit 12] The storage unit 12 is a device in which data and / or files are stored, and has a storage unit that non-temporarily stores data by means of a hard disk, a semiconductor memory, a recording medium, a memory card, and the like.
[0024] The storage unit 12 may have a mechanism that enables connection to a storage device or a storage system such as a NAS (Network Attached Storage), a SAN (Storage Area Network), a cloud storage, a file server, and / or a distributed file system via the network N.
[0025] The memory unit 12 stores programs executed by a microcomputer, an interactive database 121, a large language model, the association between types of conversations and the target(s) corresponding to such types, etc.
[0026] (Interactive database 121) In the interactive database 121, the data of the conversation and various data related to the conversation extracted, generated, etc. by the support device 1 of the present embodiment are stored in an associated manner. The data of the conversation is, for example, the voice of the conversation, the text of the conversation, etc. The various data related to the conversation include, for example, text data obtained by voice recognition of the voice of the conversation, processing units extracted from the conversation, the target of the conversation, the degree of achievement of the conversation, a collection of advice for the conversation, and other data.
[0027] The data of the conversation is stored in the interactive database 121 as text data related to the voice of the conversation. In addition, the support device 1 acquires the target of the conversation and stores it in the interactive database 121. Then, the support device 1 evaluates the degree of achievement of the stored target of the conversation and stores it in the interactive database 121. In addition, the support device 1 generates a collection of advice for the conversation and stores it in the interactive database 121.
[0028] It is preferable that the voice of the conversation and other data are stored in the interactive database 121 in association with a conversation ID for identifying the conversation. Thereby, the support device 1 can store, acquire, etc. the data of the conversation and other data using the conversation ID.
[0029] FIG. 2 is an example of the dialogue database 121. In the example shown in FIG. 2, as data of the dialogue related to the first dialogue which is a management meeting, a dialogue ID "T0001" and a voice of the dialogue "T0001.M4A" etc. are stored in association with each other. Also, in the said example, a video of the dialogue "T0001.MP4" is stored. Also, in the said example, text data "(First speaker) Since everyone is here, let's start the regular management meeting. ··· First, let's proceed from the performance report for July. The sales amount was ··· -3 million. The F rate was ··· +0.2 points. The L rate was ··· -0.4 points. The cost profit was ··· -1 million. The operating profit was ··· -3 million." is stored in association with the voice of the dialogue related to the first dialogue. Also, in the said example, the text data of the first dialogue is divided and stored into a first-1 processing unit "(First speaker) Since everyone is here, let's start the regular management meeting. ··· First, let's proceed from the performance report for July." and a first-2 processing unit "The sales amount was ··· -3 million. The F rate was ··· +0.2 points. The L rate was ··· -0.4 points. The cost profit was ··· -1 million. The operating profit was ··· -3 million." etc. based on a given length. In the said example, a target "Report the specifications such as sales amount, F rate, L rate, cost profit, operating profit, etc." and the achievement degree of the said target "All reported" are stored in association with the voice of the dialogue related to the first dialogue.
[0030] Also, in the example shown in FIG. 2, as the data of the conversation related to the second conversation which is a business activity, the conversation ID "T0002" and the voice of the conversation "T0002.M4A", etc. are stored in association with each other. Further, the text data of the second conversation "(Second speaker) First, please allow me to introduce our company. Our company is... Here is what our company has done, but what is important is your company's... Can you tell me about your company's situation? (Third speaker) Our company is... Since the corona pandemic,... Therefore, newly... (Second speaker) Thank you. What is your company's strength? (Third speaker) Yes. Our company's..." is stored. The text data of the second conversation is divided and stored into the second-1 processing unit "(Second speaker) First, please allow me to introduce our company. Our company is...", the second-2 processing unit "What is important is your company's... Can you tell me about your company's situation?", the second-3 processing unit "(Third speaker) Our company is... Since the corona pandemic,... Therefore, newly...", etc. based on a given length. In this example, the first goal "Introduce one's own company", the second goal "Ask about the other party's goal", and the third goal "Ask about the other party's issues" are stored in association with the voice of the conversation related to the second conversation. Also, in this example, the achievement degree of the first goal "Introduced, but it would be even better if there is an explanation of △△", the achievement degree of the second goal "Obtained the other party's goal", and the achievement degree of the third goal "Obtained the other party's issues" are stored.
[0031] (Large language model) The large language model of this embodiment is not particularly limited as long as it is a language model that executes natural language processing for generating a response to an input prompt. The "language model" mentioned here is a type of probability model used in natural language processing, and it is a model for probabilistically predicting how likely a given word or sentence is to occur as natural language. Specifically, the language model predicts the next sentence by calculating the occurrence probability of a given sentence, comparing the occurrence probabilities of multiple sentences, etc. Thereby, the language model can automatically generate the most likely sentence based on the context related to the given sentence.
[0032] The large language model of this embodiment is preferably a large language model that is trained with at least 500 GB or more of a large amount of text, such as OpenAI's ChatGPT-3.5, ChatGPT-4, Meta AI's LLaMa, etc., and has at least 10 billion or more parameters. With such a large language model, the support device 1 can appropriately perform various processes related to support processes (described later) such as generation of questions for one speaker related to a processing unit.
[0033] [Communication unit 13] The communication unit 13 is not particularly limited as long as it can connect the support device 1 to the network N and enable communication with the terminal T or the like. Examples of the communication unit 13 include a wireless device compatible with a mobile phone network, a device connectable to a wireless LAN, and a network card compatible with the Ethernet standard.
[0034] [Network N] The type of the network N is not particularly limited as long as it enables communication between the support device 1 and the terminal T or the like. Examples of the type of the network N include the Internet, a mobile phone network, a wireless LAN, etc.
[0035] [Terminal T] The terminal T is, for example, a personal computer, a laptop computer, a smartphone, a tablet terminal, or the like. The terminal T can execute processes such as displaying information provided from the support device 1 and providing video and audio of the dialogue to the support device 1.
[0036] [Main flowchart of support process] FIG. 3 is a main flowchart showing an example of a preferred flow of the support process executed by the support device 1 of this embodiment. FIG. 4 is a figure following FIG. 3. The following is an example of a preferred flow of the support process executed by the support device 1 of this embodiment using FIGS. 3 and 4.
[0037] The following is an example of a procedure for generating questions and the like based on the text data of a dialogue obtained by speech recognition processing for the speech of the dialogue. The support processing of the present embodiment may use, as an input, the text data of a dialogue that is a repetition of questions and answers obtained by repeating text input and text display. The support processing related to the text data of the dialogue is realized in the same manner as the support processing related to the speech of the dialogue, except that the speech recognition step is not executed. And the support processing related to the text data of the dialogue has the same effect as the support processing related to the speech of the dialogue.
[0038] [Step S1: Determine whether to acquire dialogue data] The control unit 11 executes the data acquisition unit 111 in cooperation with the storage unit 12 and the communication unit 13. Then, the control unit 11 executes, by the data acquisition unit 111, a process of determining whether to acquire dialogue data (step S1, data acquisition step). If it is determined to acquire, the control unit 11 acquires the speech or the like. Then, the control unit 11 moves the process to step S2. If it is not determined to acquire, the control unit 11 returns the process to step S1.
[0039] The procedure for determining whether to acquire dialogue data in the data acquisition step is not particularly limited. The procedure may be, for example, a procedure for determining to acquire dialogue data when the dialogue data is provided from the acquisition source of the dialogue data. The "dialogue data" referred to here includes, for example, data related to the speech of the dialogue, text data of the dialogue, and the like. The "data related to the speech of the dialogue" may be data that further includes a video corresponding to the speech of the dialogue. The acquisition source of the data in the data acquisition step is, for example, the terminal T. The support device 1 may acquire the dialogue data stored in the storage unit 12. The format of the data is not particularly limited.
[0040] The dialogue data exemplified by the speech or the like acquired in the data acquisition step is preferably related to the ongoing dialogue. Thereby, the support device 1 can provide the user with questions and the like for prompting the presentation of information to be provided from one speaker to another speaker in real time.
[0041] [Step S2: Perform speech recognition processing] The control unit 11 executes the speech recognition unit 112 in cooperation with the storage unit 12. Then, the control unit 11 executes a process of performing speech recognition processing on the speech acquired in the data acquisition step by the speech recognition unit 112 (step S2, speech recognition step). The control unit 11 moves the process to step S3. The speech recognition processing includes a procedure for generating text data of the dialogue included in the speech.
[0042] The method of speech recognition in the speech recognition step is not particularly limited. In order to generate questions, evaluate the degree of achievement, etc. based on the standing positions of each speaker with respect to the target, it is preferable that the speech recognition in the speech recognition step includes a procedure for identifying the speaker for each utterance. The support device 1 generates text data to which data capable of identifying the timing of the utterance is attached in the speech recognition step. Such data is, for example, a time stamp.
[0043] [Step S3: Store text data] The control unit 11 executes a process of storing the text data of the dialogue such as the text data generated in the speech recognition step in the dialogue database 121 (step S3, text data storage step). The control unit 11 moves the process to step S4.
[0044] Subsequently, the support device 1 executes a processing unit extraction step including a first processing unit extraction step and a second processing unit extraction step. The first processing unit extraction step is a step of extracting a processing unit of a given length from the text data. The second processing unit extraction step is a step of removing a part that deviates from the gist of the utterance from the processing unit. Steps S4 and S5 are an example of the processing unit extraction step.
[0045] [Step S4: Extract a processing unit of a given length] The control unit 11 executes the processing unit extraction unit 113 in cooperation with the storage unit 12. Then, the control unit 11 executes a process of extracting a processing unit of a given length from the text data stored in the text data storage step by the processing unit extraction unit 113 (step S4, first processing unit extraction step). The control unit 11 moves the process to step S5.
[0046] The "given length" is preferably the number of characters of predetermined text data exemplified by, for example, 300 characters or the like. Since the "given length" is the number of characters, the support device 1 can generate a question (described later) that prompts the presentation of information for each fixed amount of speech for text data not relying on speech recognition processing. Also, since the "given length" is the number of characters, the support device 1 can generate a question (described later) that prompts the presentation of information for each fixed amount of speech in text data related to speech recognition processing, regardless of whether the speaker is speaking quickly or slowly. The "predetermined number of characters of text data" is not particularly limited, but is preferably included in the following numerical range. The upper limit of the "predetermined number of characters of text data" is preferably 600 characters, more preferably 400 characters, and even more preferably 300 characters. Thereby, the support device 1 can shorten the time from when the information to be presented is not presented until a question is generated. The lower limit of the "predetermined number of characters of text data" is preferably 150 characters, more preferably 250 characters, and even more preferably 300 characters. Thereby, the support device 1 can prevent itself from determining whether the information to be presented is presented in a state where the statements serving as judgment criteria are not sufficiently complete.
[0047] When the text data is text data related to speech recognition processing, the "given length" is preferably a predetermined time length exemplified by, for example, one minute. Since the "given length" is a time length, the support device 1 can generate a question (described later) that prompts the presentation of information at regular intervals. The "predetermined time length" is not particularly limited, but is preferably included in the following numerical range. The upper limit of the "predetermined time length" is preferably 3 minutes, more preferably 2 minutes, and even more preferably 1 minute. Thereby, the support device 1 can shorten the time from when the information to be presented is not presented until a question is generated. The lower limit of the "predetermined time length" is preferably 30 seconds, more preferably 50 seconds, and even more preferably 1 minute. Thereby, the support device 1 can prevent determining whether the information to be presented is presented in a state where the statements serving as the basis for the determination are not sufficiently complete.
[0048] [Step S5: Remove parts deviated from the gist] The control unit 11 executes a process of removing parts deviated from the gist from the processing unit of the given length extracted in the first processing unit extraction step from the processing unit extraction unit 113 (step S5, second processing unit extraction step). The control unit 11 moves the process to step S6.
[0049] In the second processing unit extraction step, parts deviated from the gist are removed from the processing unit by various natural language processes. The "various natural language processes" include, for example, a procedure for removing parts deviated from the gist from the processing unit using a pattern that defines parts deviated from the gist. Such a procedure has a relatively small processing amount. Therefore, when using such a procedure, the support device 1 can suppress the power consumption related to the procedure.
[0050] (Parts deviated from the gist) The parts that deviate from the gist are, for example, the following parts. Mistaken input: parts that are mistakenly input, such as "......", "desu ne", etc.; Silent parts: parts where there is no speech and blank characters, line breaks, and / or symbols indicating a silent state are generated by speech recognition processing; Interjections: mere interjections such as "ah", "eh", "hmm", "yes yes"; Speech pauses: parts where the speech stops in the middle of a phrase, such as "two thousand and, in the case of 2022..."; Retractions: parts where the spoken phrase is cancelled and another phrase is spoken, such as "In the case of Company B, excuse me, in the case of Company A"; Duplication of the same phrase due to speech recognition processing: parts where, unlike the actual conversation due to speech recognition processing, the same phrase is duplicated in the text data, such as "The results for 2022 are as follows. The results for 2022 are as follows."
[0051] If such parts are included in the processing unit, the length of the input in the large language model will become long. In addition, since such parts are unclear, they can increase the processing load in the large language model and make the processing result inaccurate. Therefore, the second processing unit extraction step can further reduce the computational amount related to the large language model and obtain a more accurate processing result.
[0052] [Step S6: Generate a question using a large language model] The control unit 11 executes the question generation unit 114 in cooperation with the storage unit 12. Then, the control unit 11 executes a process of generating a question by using a large language model that includes the processing unit extracted in the processing unit extraction step as an input from the question generation unit 114 (step S6, question generation step). The control unit 11 moves the process to step S5. The question generated here is a question that prompts the presentation of information to be provided from one speaker to another regarding the gist of the speech included in the processing unit.
[0053] The input given to the large language model is, for example, a prompt including an instruction such as "generate questions to prompt the speaker to supplement information" and the above-described processing unit. In order to generate appropriate questions according to the purpose of the dialogue, it is preferable that the instruction includes additional instructions corresponding to the given purpose of the dialogue. For example, when the purpose of the dialogue is to promote products and / or services, an additional instruction such as "for this sales pitch" may be corresponding. When this additional instruction is included, the above instruction becomes "generate questions to prompt the speaker to supplement information for this sales pitch".
[0054] The input given to the large language model preferably includes additional instructions that convey the purpose of the dialogue. The purpose of the dialogue is, for example, the product planning stage, the marketing strategy stage, the sales stage, etc. Thereby, the support device 1 can generate more appropriate questions that reflect the purpose of the dialogue.
[0055] By generating the above questions in the question generation step, the speaker supported by the support device 1 can convey the matters to be conveyed in the dialogue in the natural form of an answer to the generated questions. By repeating such questions, the support device 1 can support obtaining solutions to the initial problems, marketing measures, sales measures, and sales-related education.
[0056] Incidentally, the computational complexity of the large language model usually increases in proportion to the square of the length of the input. Since the question generation step uses the processing unit extracted in the processing unit extraction step as the input, the length of the input is short. Therefore, the support device 1 of the present embodiment reduces the computational complexity related to the large language model as compared with the case where the entire text data is used as the input.
[0057] The support process preferably includes a series of processes (from step S7 to step S8) that generate the degree of achievement for each of a plurality of objectives. At this time, the support device 1 generates questions that prompt the supplementation of information and the degree of achievement for each of the plurality of objectives. Thereby, the support device 1 can prompt the speaker via the questions to supplement the content indicated by the degree of achievement while showing the degree of achievement of the conversation so far. Therefore, the speaker supported by the support device 1 can further enrich the content of the conversation.
[0058] It is preferable that the large language model related to the question generation step has been pre-trained on questions based on the so-called SPIN method. The SPIN method is a framework that incorporates four types of questions, namely Situation (situation question), Problem (problem question), Implication (implication question), and Need-Payoff (solution question), into the conversation. The pre-training is performed, for example, using learning data that associates text indicating the scene of the conversation with any of the above four types of questions corresponding to the scene. Through the pre-training, questions based on the SPIN method are generated. Therefore, the support device 1 can make appropriate proposals for the needs of the person (e.g., customer) with whom the user is having a conversation.
[0059] [Regarding generating examples of utterances other than questions] To further support the conversation, the support process preferably includes an utterance example generation step (not shown) that generates examples of utterances other than questions suitable for achieving the purpose of the conversation by the same procedure as the question generation step. At this time, the input given to the large language model is, for example, a prompt that includes an instruction such as "Please generate an utterance suitable for the conversation on the above objective and theme" and the above processing unit.
[0060] [Step S7: Determine whether to obtain the objective] The control unit 11 executes the objective acquisition unit 115 in cooperation with the storage unit 12 and the communication unit 13. Then, the control unit 11 executes a process of determining whether to acquire a plurality of objectives related to the above-described dialogue by the objective acquisition unit 115 (step S7, objective acquisition step). If it is determined that the objectives are to be acquired, the control unit 11 acquires the objectives. Then, the control unit 11 moves the process to step S8. If it is determined that the objectives are not to be acquired, the control unit 11 moves the process to step S9.
[0061] The procedure for determining whether to acquire objectives in the objective acquisition step is not particularly limited. The procedure may be, for example, a procedure for determining to acquire objectives when the objectives are provided from the source of the objectives. The source of the "objectives" in the objective acquisition step is, for example, the terminal T. The support device 1 may acquire the "objectives" stored in the storage unit 12. In order to enable the acquisition of objectives based on the determination of the type of dialogue, the "objectives" stored in the storage unit 12 may be objectives associated with the type of dialogue.
[0062] The objectives acquired in the objective acquisition step are preferably related to the ongoing dialogue. Thereby, the support device 1 can provide the user with an evaluation of the degree of achievement related to each of the plurality of objectives in real time. Also, the objective acquisition step may be executed prior to step S1. Thereby, the support device 1 can utilize the objectives in the question generation step and the like.
[0063] [Step S8: Evaluate the degree of achievement using a large language model] The control unit 11 executes the achievement evaluation unit 116 in cooperation with the storage unit 12. Then, the control unit 11 executes a process of evaluating the degree of achievement of each of the plurality of objectives acquired in the objective acquisition step by a process using a large language model that includes, as an input, the processing units extracted in the processing unit extraction step by the achievement evaluation unit 116 (step S8, achievement evaluation step). The control unit 11 moves the process to step S9.
[0064] The input given to the large language model is, for example, a prompt that includes any one of a plurality of objectives acquired in the objective acquisition step, an instruction of "Please evaluate whether this objective has been achieved", and the above-described processing unit. In order to output the degree of achievement in a certain format in the output step described later, it is preferable that the above-described instruction includes an additional instruction specifying the mode of evaluation. The additional instruction specifying the mode of evaluation is, for example, "in four levels of ◎○△× from the highest evaluation to the lowest".
[0065] The support device 1 generates the degree of achievement in the degree of achievement evaluation step. Therefore, the speaker supported by the support device 1 can further enrich the content of the dialogue by using the degree of achievement. Since the degree of achievement evaluation step uses the processing unit extracted in the processing unit extraction step as an input, the power consumption can be suppressed in the same manner as in the question generation step.
[0066] The support process preferably includes a series of processes (step S9 to step S10) for generating an advice collection. At this time, the support device 1 generates a question prompting for additional information and an advice collection. Thereby, the support device 1 can prompt the speaker via a question to supplement the advice collection while showing a method for improving the dialogue so far in the advice collection. Therefore, the speaker supported by the support device 1 can further enrich the content of the dialogue.
[0067] [Step S9: Determine whether to generate an advice collection] The control unit 11 executes a process of determining whether to generate an advice collection summarizing the gist of the above-described dialogue in cooperation with the storage unit 12 and the communication unit 13 (step S9, advice collection generation determination step). If it is determined to generate, the control unit 11 moves the process to step S10. If it is not determined to generate, the control unit 11 moves the process to step S11.
[0068] The procedure for determining whether to generate an advice collection in the advice collection generation determination step is not particularly limited. The procedure may be, for example, a procedure for determining to generate an advice collection when there is an instruction from a user via the terminal T or the like.
[0069] [Step S10: Generate an advice collection using a large language model] The control unit 11 executes the advice collection generation unit 117 in cooperation with the storage unit 12. Then, the control unit 11 executes a process of generating an advice collection that summarizes the advice for the above-described dialogue by using a large language model that includes, as an input, the processing unit extracted in the processing unit extraction step from the advice collection generation unit 117 (step S10, advice collection generation step). The control unit 11 moves the process to step S11.
[0070] The input given to the large language model is, for example, a prompt including an instruction such as "Please summarize the improvement suggestions for the following text" and the above-described processing unit. In order to output the advice collection in a certain format in the output step described later, it is preferable that the above-described instruction includes an additional instruction specifying the format of the advice collection. The additional instruction specifying the format of the advice collection is, for example, "using the SPIN method", "within 40 characters", "ending with a noun", "in a desu - masu style", etc.
[0071] The above-described instruction preferably includes an instruction to generate a talk script along the SPIN method. At this time, in order to be able to give appropriate proposals for the needs of the person with whom the user is having a dialogue, it is preferable that the large language model has been subjected to the above-described pre - training related to questions based on the SPIN method.
[0072] In addition, the above instructions preferably include additional instructions for instructing to indicate the theme of the dialogue. The additional instructions include, for example, an enumeration of the theme types such as "Please attach which of the following the theme of this part corresponds to." The theme types enumerated here are, for example, opening, short presentation, probing, probing of problems / issues, suggestion of solutions, other probing, closing, handling of counterarguments, schedule adjustment, summary, etc. Thereby, the support device 1 can compile an advice collection for each theme. Thus, the user can grasp the flow of the dialogue for each theme. Also, when the text data and the speech time are attached to the processing unit, the support device 1 preferably attaches to the advice collection the time range during which the dialogue related to the theme was conducted for each theme. Thereby, the user can grasp the time spent for each theme and utilize it for improving future dialogues.
[0073] In addition, the above instructions preferably include additional instructions for instructing to summarize the numbers mentioned in the dialogue. Numbers are important information for grasping the specific content of the dialogue. With the above additional instructions, the support device 1 can summarize the numbers mentioned in the dialogue in the advice collection. Thus, the user can grasp various numbers mentioned in the dialogue.
[0074] The support device 1 generates an advice collection in the advice collection generation step. Thus, the user supported by the support device 1 can use the advice collection to further enrich the content of future dialogues. Since the advice collection generation step takes the processing units extracted in the processing unit extraction step as input, power consumption can be suppressed in the same way as in the question generation step and the like.
[0075] [Step S11: Output the result] The control unit 11 executes the output unit 118 in cooperation with the storage unit 12. Then, the control unit 11 executes a process of transmitting an instruction to cause the output unit 118 to output various results related to the above-described question generation step, achievement degree evaluation step, advice collection generation step, etc. (step S11, output step). The control unit 11 returns the process to step S1 and repeats the processes from step S1 to step S11.
[0076] The destination of the above-described instruction transmission is not particularly limited. For example, when the support device 1 includes a display unit such as a display, the destination may be the display unit. Also, for example, when dialogue data is provided from the terminal T, the destination may be the terminal T.
[0077] By the output step, various results related to the above-described question generation step, achievement degree evaluation step, advice collection generation step, etc. are output. Thereby, the support device 1 can support so as to be able to convey what should be conveyed in the dialogue in the natural form of an answer to the question generated by the user. Also, thereby, the support device 1 can support so that the user can further enrich the content of the dialogue using the achievement degree. In addition, thereby, the support device 1 can support so that the user can further enrich the content of future dialogue using the advice collection.
[0078] [Big Data Provision Step] The support process preferably includes a big data provision step of providing the data related to the process as big data. The data includes, for example, questions prompting the presentation of information to be provided, the achievement degree of the target, advice collections, and the like. Thereby, the support device 1 can provide support for providing big data related to the dialogue.
[0079] [Usage Example] The following is a usage example of the support device 1.
[0080] [Support Using Real-Time Text] The following is an explanation about support using real-time text.
[0081] [Obtaining the goal] The user transmits, via the terminal T, the purpose of the upcoming conversation to the support device 1. The support device 1 receives the purpose and stores it in the storage unit 12 in association with the above-mentioned conversation.
[0082] [Obtaining the text of the conversation] The user transmits, via the terminal T, the text of the ongoing conversation to the support device 1. The support device 1 receives the text and stores it in the conversation database 121.
[0083] [Providing information contributing to the improvement of the conversation] The support device 1 extracts processing units from the text data. Then, the support device 1 generates, using a large language model, a question prompting the presentation of information to be provided from one speaker to another. In addition, the support device 1 estimates, using a large language model, the degree of goal achievement. Also, in addition, the support device 1 generates, using a large language model, a collection of advice. Thereafter, the support device 1 provides the user with the question, the degree of goal achievement, and the collection of advice.
[0084] The processing related to these large language models uses the processing units extracted in the processing unit extraction step as input, so the input length is short. Therefore, the support device 1 in the present embodiment reduces the computational amount related to the large language model as compared with the case of using the entire text data as input. Therefore, the support device 1 in the present embodiment can suppress the power consumption in the support.
[0085] Thereby, the support device 1 can support the user so that the matters to be conveyed in the conversation are conveyed. Since various information is provided in real time, the support device 1 can support the user so that the matters to be conveyed can be conveyed in an even shorter time. Thereby, the support device 1 enables the user to improve the ongoing conversation.
[0086] [Support using recorded conversation audio] The following is an explanation of support using recorded dialogue voices.
[0087] [Acquisition of Voice, etc.] The user transmits the voice, etc. of the dialogue recorded in the online meeting system to the support device 1 via the terminal T. The voice, etc. may include video corresponding to the voice. The support device 1 receives the voice, etc. and stores it in the dialogue database 121.
[0088] [Acquisition of Purpose] The user transmits the purpose of the above-mentioned dialogue to the support device 1 via the terminal T. The support device 1 receives the purpose and stores the purpose in the storage unit 12 in association with the above-mentioned dialogue.
[0089] [Provision of Questions, etc. Prompting the Presentation of Information to be Provided] The support device 1 generates text data based on the above-mentioned voice. The speaker identification result is attached to this text data. Then, the support device 1 extracts processing units from the text data. The support device 1 provides the user with the question, the degree of goal achievement, and the advice set generated by processing using a large language model that includes the processing unit as an input in real time.
[0090] After that, the support device 1 generates a question prompting the presentation of information to be provided from one speaker to another speaker by the large language model. In addition, the support device 1 estimates the degree of goal achievement by the large language model. Also, in addition, the support device 1 generates an advice set by the large language model. Then, the support device 1 provides the user with the question, the degree of goal achievement, and the advice set.
[0091] The support device 1 of the present embodiment can suppress the power consumption in support even when supporting based on the voice of the dialogue, similar to the case of supporting using real-time text by extracting the processing unit.
[0092] In the scope of the idea of the present invention, those skilled in the art can conceive of various modification examples and correction examples. Therefore, those modification examples and correction examples are understood to belong to the scope of the present invention. For example, with respect to the foregoing embodiments, those in which those skilled in the art appropriately add, delete, or change the design of components, or add, omit, or change the conditions of steps, are also included in the scope of the present invention as long as they have the gist of the present invention.
Explanation of Signs
[0093] S system 1 Support device 11 Control unit 111 Data acquisition unit 112 Voice recognition unit 113 Processing unit extraction unit 114 Question generation unit 115 Purpose acquisition unit 116 Achievement degree evaluation unit 117 Advice collection generation unit 118 Output unit 12 Storage unit 121 Dialogue database 13 Communication unit N Network T Terminal
Claims
1. A processing unit extraction unit that extracts a processing unit of a given length from the text data of the dialogue, A question generation unit that generates a question to one speaker related to the processing unit by processing using a large language model that includes the processing unit as an input, Comprising, The processing unit extraction unit removes a part that deviates from the gist of the utterance from the processing unit, The question generation unit generates a question that prompts the presentation of information to be provided from the one speaker to another speaker regarding the gist of the utterance included in the processing unit, A dialogue progress support device.
2. An objective acquisition unit that acquires a plurality of objectives related to the dialogue, An achievement degree evaluation unit that evaluates the achievement degree of each of the plurality of objectives by processing using a large language model that includes the processing unit as an input, The support device according to claim 1, further comprising.
3. The support device according to claim 1, further comprising an advice collection generation unit that generates an advice collection summarizing the advice to the dialogue by processing using a large language model that includes one or more of the processing units as an input.
4. The processing unit extraction unit extracts a processing unit from the text data of the ongoing dialogue, The question generation unit generates a question to the speaker of the ongoing dialogue, The support device according to claim 1.
Citation Information
Patent Citations
Meeting management device, and meeting management method
JP2011066794A