Dialogue analysis device, dialogue analysis method, and program
The dialogue analysis device uses machine learning models to analyze dialogue text, visualizing the success or failure of dialogue purposes and their reasons, thereby enhancing operational improvements in business settings.
Patent Information
- Application Number
- JP2025001076
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-07
- Estimated Expiration
- 2045-01-06
AI Technical Summary
Existing technologies do not effectively visualize the success or failure of dialogue purposes and the reasons behind them based on dialogue text, limiting their ability to improve business operations.
A dialogue analysis device utilizing multiple machine learning models to estimate the purpose of a speaker, determine the success or failure of that purpose, and identify the reasons for success or failure, with the results being displayed on a user terminal.
Enables visualization of dialogue success or failure and the underlying reasons, allowing users to improve their operations by understanding the effectiveness of their dialogues.
Smart Images

Figure 0007672090000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a dialogue analysis device, a dialogue analysis method, and a program. [Background technology]
[0002] In recent years, a technology has been proposed that analyzes telephone voices and uses them to improve business operations. For example, Patent Document 1 describes a business support device that includes a telephone voice acquisition unit that acquires a recorded telephone voice of a conversation between an inquirer who inquires about a service at an inquiry desk and an answerer who answers the inquiry at the inquiry desk, an analysis unit that analyzes the acquired telephone voices and extracts keywords related to the inquiry, a history information generation unit that generates history information indicating a history of the inquiry based on the extracted keywords, a ranking information generation unit that generates ranking information indicating information ranked on the inquiry based on the generated history information, and an output control unit that transmits the generated ranking information to a device that displays a ranking related to the inquiry on a homepage related to the service based on the generated ranking information. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2024-1473 A Summary of the Invention [Problem to be solved by the invention]
[0004] There is still room for improvement in how voice calls can be used to improve business operations.
[0005] The present invention has been made in consideration of the above problems, and has an object to visualize whether a dialogue goal has been achieved and the reasons for this, based on the dialogue text. [Means for solving the problem]
[0006] According to one embodiment, a dialogue analysis device is provided that includes one or more control units, the control unit utilizing a first machine learning model to estimate a first speaker's purpose from a dialogue text between a first speaker and a second speaker, utilizing a second machine learning model to determine whether the purpose has been achieved from the dialogue text, and utilizing a third machine learning model to obtain reasons for the success or failure of the purpose from the dialogue text, and causes a user terminal to display whether the purpose has been achieved and the reasons therefor. Effect of the Invention
[0007] According to one embodiment, based on the dialogue text, it is possible to visualize whether the dialogue goal was successful and why. [Brief description of the drawings]
[0008] [Figure 1] FIG. 1 illustrates an example of the configuration of a dialogue analysis system 1000. [Diagram 2] 1 is a diagram illustrating an example of a hardware configuration of a dialogue analysis device 1. FIG. [Diagram 3] 2 is a diagram illustrating an example of a hardware configuration of a user terminal 2. FIG. [Figure 4] 1 is a flowchart illustrating an example of a dialogue analysis method. [Diagram 5] FIG. 1 is a schematic diagram illustrating a dialogue analysis method. [Figure 6] FIG. 13 is a diagram showing an example of an analysis result screen sc1. [Figure 7] FIG. 13 is a diagram showing an example of a tally result screen sc2. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] Hereinafter, each embodiment of the present invention will be described with reference to the accompanying drawings. In the description of the specification and drawings of each embodiment, components having substantially the same functional configuration are designated by the same reference numerals, and redundant description will be omitted.
[0010] <System configuration> First, an overview of the dialogue analysis system 1000 according to this embodiment will be described. The dialogue analysis system 1000 is an information processing system that visualizes whether the dialogue purpose has been achieved and the reasons for this from a dialogue text.
[0011] A dialogue is a discussion between two or more speakers. The dialogue may be a voice dialogue such as a phone call, or a text dialogue such as a chat.
[0012] A speaker is a person who has a conversation. A speaker may be a person, or may be a voice assistant or chatbot capable of conversation with a person. Hereinafter, a speaker who is the subject of analysis by the conversation analysis system 1000 is referred to as a first speaker, and another speaker is referred to as a second speaker. A speaker is, for example, a business operator or its customer, but is not limited to this. A business operator is a general term for an organization that conducts business such as a company, and the employees of the business operator.
[0013] When the first speaker is a customer and the second speaker is a business operator, the customer is the subject of analysis, and the purpose of the conversation is analyzed as the purpose of the conversation. The purpose of the customer is, for example, but not limited to, purchase, reservation, inquiry, stock confirmation, or appointment.
[0014] When the first speaker is a business operator (e.g., a salesperson) and the second speaker is a customer, the business operator is the subject of analysis, and the purpose of the business operator is analyzed as the purpose of the dialogue. The purpose of the business operator is, for example, a sale or an appointment, but is not limited to these.
[0015] Fig. 1 is a diagram showing an example of the configuration of a dialogue analysis system 1000. As shown in Fig. 1, the dialogue analysis system 1000 includes a dialogue analysis device 1 and a user terminal 2, which are communicatively connected to each other via a network N. The network N is, for example, a wired LAN (Local Area Network), a wireless LAN, the Internet, a public line network, a mobile data communication network, or a combination of these. In the example of Fig. 1, the dialogue analysis system 1000 includes one dialogue analysis device 1 and one user terminal 2, but may include multiple of each.
[0016] The dialogue analysis device 1 is an information processing device that analyzes a dialogue between a first speaker and a second speaker. The dialogue analysis device 1 is, for example, but not limited to, a PC (Personal Computer), a smartphone, a tablet terminal, a server device, or a microcomputer. In the example of FIG. 1, the dialogue analysis device 1 is one information processing device, but may be realized as a system consisting of multiple information processing devices connected via a network N.
[0017] The user terminal 2 is an information processing device used by a user of the dialogue analysis device 1. The user browses the analysis results by the dialogue analysis device 1 on the user terminal 1. The user is, for example, but is not limited to, a business operator who has a dialogue with a customer. The user terminal 2 is, for example, but is not limited to, a PC, a smartphone, or a tablet terminal.
[0018] <Hardware configuration of the dialogue analysis device 1> Next, a description will be given of the hardware configuration of the dialogue analysis device 1. Fig. 2 is a diagram showing an example of the hardware configuration of the dialogue analysis device 1. As shown in Fig. 2, the dialogue analysis device 1 includes a control unit 11, a storage unit 12, and a communication unit 13, which are connected to each other via a bus B1.
[0019] The control unit 11 executes various programs stored in the storage unit 12, controls the entire dialogue analysis device 1, and realizes the functions of the dialogue analysis device 1. The control unit 11 is, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), an ASIC (Application Specific Integrated Circuit), a DSP (Digital Signal Processor), or a combination of these.
[0020] The storage unit 12 is a computer-readable recording medium that stores various programs and data. The storage unit 12 is, for example, a Read Only Memory (ROM), a Random Access Memory (RAM), a flash memory, a Hard Disk Drive (HDD), a Solid State Drive (SSD), a Storage Class Memories (SCM), or a combination of these.
[0021] The communication unit 13 controls communication between the dialogue analysis device 1 and an external device via the network N. The communication unit 13 is, for example, a Bluetooth (registered trademark) module, a Wi-Fi (registered trademark) module, a ZigBee (registered trademark) module, an Ethernet (registered trademark) module, a mobile communication module, or a combination of these.
[0022] In this embodiment, the program may be written into the storage unit 12 during the manufacturing stage of the dialogue analysis apparatus 1, or may be provided to the dialogue analysis apparatus 1 via the network N, or may be provided to the dialogue analysis apparatus 1 via a non-transitory computer-readable recording medium such as a recording medium. The recording medium is, for example, a CD (Compact Disc), a DVD (Digital Versatile Disc), a FD (Floppy Disk), an MO (Magneto-Optical disk), a BD (Blu-ray (registered trademark) Disc), a USB (registered trademark) memory, an SD card, or a combination of these.
[0023] <Hardware configuration of user terminal 2> Next, a description will be given of the hardware configuration of the user terminal 2. Fig. 3 is a diagram showing an example of the hardware configuration of the user terminal 2. As shown in Fig. 3, the user terminal 2 includes a control unit 21, a storage unit 22, a communication unit 23, an input unit 24, and an output unit 25, which are mutually connected via a bus B2.
[0024] The control unit 21 executes various programs stored in the storage unit 22, controls the entire user terminal 2, and realizes the functions of the user terminal 2. The control unit 21 is, for example, a CPU, an MPU, a GPU, an ASIC, a DSP, or a combination of these.
[0025] The storage unit 22 is a computer-readable recording medium that stores various programs and data. The storage unit 22 is, for example, a ROM, a RAM, a flash memory, a HDD, an SSD, an SCM, or a combination of these.
[0026] The communication unit 23 controls communication between the user terminal 2 and an external device via the network N. The communication unit 23 is, for example, a Bluetooth (registered trademark) module, a Wi-Fi (registered trademark) module, a ZigBee (registered trademark) module, an Ethernet (registered trademark) module, a mobile communication module, or a combination of these.
[0027] The input unit 24 inputs information to the user terminal 2. The input unit 24 is, for example, a mouse, a keyboard, a touch panel, a microphone, a scanner, an imaging device (camera), various sensors, an operation button, or a combination of these.
[0028] The output unit 25 outputs information from the user terminal 2. The output unit 25 is, for example, a display device, a projector, a printer, a speaker, a vibrator, or a combination of these.
[0029] In this embodiment, the program may be written into the storage unit 22 during the manufacturing stage of the user terminal 2, may be provided to the user terminal 2 via the network N, or may be provided to the user terminal 2 via a non-transitory computer-readable recording medium such as a recording medium. The recording medium is, for example, a CD, a DVD, a FD, an MO, a BD, a USB (registered trademark) memory, an SD card, or a combination of these.
[0030] <Dialogue analysis method> Next, a description will be given of a dialogue analysis method executed by the dialogue analysis system 1000. Fig. 4 is a flowchart showing an example of the dialogue analysis method. Fig. 5 is a schematic diagram illustrating the dialogue analysis method.
[0031] (Step S101) First, the control unit 11 of the dialogue analysis device 1 acquires the dialogue voice A, dialogue date and time, and user ID of the first and second speakers, and stores the dialogue voice A, dialogue date and time, and user ID in association with each other in the memory unit 12. The dialogue voice A is audio data obtained by recording a voice dialogue. The dialogue voice A is, for example, a recorded voice of a conversation (telephone call), a meeting, a business negotiation, or an inquiry, but is not limited to this. The dialogue date and time is the date and time when the dialogue voice was recorded. The user ID is an ID that uniquely identifies a user. The control unit 11 may acquire the dialogue voice A and the dialogue date and time from the user terminal 2, or from an external device connected via the network N.
[0032] In the example of Figure 5, conversation voice A is voice data of a call between a customer (first speaker) and a hotel employee (second speaker), who is the business operator. The customer asks, "Can I make a reservation for 12 / 28?", and the hotel employee replies, "I'm sorry, but we are fully booked on 12 / 28."
[0033] (Step S102) Next, the control unit 11 generates a dialogue text T from the dialogue voice A by speech recognition processing using the fourth machine learning model. The dialogue text T is a text (transcription) of the content of the voice dialogue. The control unit 11 can use any machine learning model capable of speech recognition processing.
[0034] Specifically, the control unit 11 inputs the dialogue voice A to a fourth machine learning model that performs voice recognition processing, and acquires the text output by the fourth machine learning model as the dialogue text T. The fourth machine learning model may be stored in the storage unit 12, or may be stored in an external device. In the latter case, the control unit 11 transmits the dialogue voice A to the external device using an API (Application Programming Interface), and receives the dialogue text T from the external device.
[0035] The control unit 11 preferably separates the dialogue voice A into speakers by speech recognition processing, and generates a dialogue text T having speaker information for each sentence. The speaker information is information indicating a speaker. This makes it possible to know which sentence is the utterance of which speaker, thereby improving the accuracy of the subsequent analysis processing.
[0036] The control unit 11 may acquire the dialogue text T from an external device such as the user terminal 2, without acquiring the dialogue voice A. The dialogue text T may be text that the external device has converted the dialogue voice A into text, or may be a text dialogue. As described above, the dialogue text T preferably has speaker information for each sentence.
[0037] In the example of Fig. 5, a dialogue text T having speaker information is generated. This dialogue text T includes the text "Can I make a reservation for 12 / 28?" as a utterance from speaker A, and the text "I'm sorry, we are fully booked on 12 / 28" as a utterance from speaker B.
[0038] (Step S103) Next, the control unit 11 uses the first machine learning model to estimate the purpose of the first speaker from the dialogue text T. The first machine learning model may be a classification model that classifies the purpose of the dialogue text T, or may be a large language model (LLM).
[0039] A plurality of purposes of the first speaker are set in advance by the user or by default. The control unit 11 estimates which purpose the dialogue text T corresponds to from among the plurality of set purposes.
[0040] For example, the first machine learning model may be a classification model capable of multi-class classification that outputs the purpose of the input text from among multiple purposes set in advance. In this case, the control unit 11 inputs the dialogue text T to the first machine learning model, and acquires the purpose output by the first machine learning model as the purpose of the first speaker.
[0041] The first machine learning model may be a classification model capable of binary classification, which outputs an evaluation value indicating whether the purpose of the input text corresponds to one of the previously set purposes. In this case, the control unit 11 inputs the dialogue text T to each of a plurality of first machine learning models prepared for each purpose of the first speaker, compares the evaluation values output by the plurality of first machine learning models, and acquires the purpose with the highest evaluation value as the purpose of the first speaker.
[0042] The first machine learning model may be a large-scale language model. In this case, the control unit 11 inputs a prompt to the first machine learning model to instruct the first machine learning model to estimate the purpose of the first speaker from the dialogue text T, and obtains the purpose output by the first machine learning model as the purpose of the first speaker.
[0043] The purpose of the first speaker, which is a candidate for the estimation result, is set in advance according to the type (situation) of the dialogue corresponding to the dialogue text T. For example, when the dialogue text T corresponds to a dialogue between a customer (first speaker) and a business (second speaker), purchase, reservation, inquiry, stock check, and appointment are set in advance as candidates for the estimation result as the purpose of the first speaker, and the control unit 11 estimates which of these is the customer's purpose from the dialogue text T using the first machine learning model. Also, for example, when the dialogue text T corresponds to a dialogue between a business (first speaker) and a customer (second speaker), sales and appointment are set in advance as candidates for the estimation result as the purpose of the first speaker, and the control unit 11 estimates which of these is the business's purpose from the dialogue text T using the first machine learning model.
[0044] The first machine learning model may be stored in the storage unit 12 or in an external device. In the latter case, the control unit 11 inputs the dialogue text T (and the prompt) to the first machine learning model stored in the external device using an API, and receives the output of the first machine learning model from the external device.
[0045] In the example of FIG. 5, the purpose of the first speaker is estimated from the dialogue text T to be "making a reservation."
[0046] (Step S104) Having estimated the purpose of the first speaker, the control unit 11 uses the second machine learning model to determine whether the purpose has been achieved (success or failure) from the dialogue text T. A successful goal means that the purpose is achieved. A failed goal means that the purpose is not achieved. For example, if the purpose is to make a reservation, being able to make a reservation corresponds to success, and not being able to make a reservation corresponds to failure. The second machine learning model is, for example, but is not limited to, a large-scale language model.
[0047] If the second machine learning model is a large-scale language model, the control unit 11 inputs a prompt to the second machine learning model instructing it to determine the success or failure of the objective estimated in step S103 from the dialogue text T, and obtains the success or failure output by the second machine learning model as the success or failure of the objective.
[0048] The second machine learning model may be stored in the storage unit 12 or in an external device. In the latter case, the control unit 11 inputs the dialogue text T and the prompt to the second machine learning model stored in the external device using an API, and receives the output of the second machine learning model from the external device.
[0049] In the example of FIG. 5, from the dialogue text T, the purpose of the first speaker, "making a reservation," is determined to be "failed."
[0050] (Step S105) When determining whether the first speaker's purpose is successful, the control unit 11 generates a reason text from the dialogue text T using a third machine learning model. The reason text is a text indicating the reason for the success or failure of the purpose. If it is determined in step S104 that the purpose is successful, the reason text becomes a text indicating the reason for success. If it is determined in step S104 that the purpose is unsuccessful, the reason text becomes a text indicating the reason for failure. For example, if the purpose of the customer (first speaker) is to make a hotel reservation and the hotel reservation fails, a reason text such as "the hotel was fully booked" is generated. The third machine learning model is, for example, a large-scale language model, but is not limited to this.
[0051] If the third machine learning model is a large-scale language model, the control unit 11 inputs a prompt to the third machine learning model instructing it to generate text indicating the reason for the success or failure of the objective determined in step S104 from the dialogue text T, and obtains the text output by the third machine learning model as the reason text.
[0052] The third machine learning model may be stored in the storage unit 12 or in an external device. In the latter case, the control unit 11 inputs the dialogue text T and the prompt to the third machine learning model stored in the external device using an API, and receives the output of the third machine learning model from the external device.
[0053] In the example of FIG. 5, the reason text "The hotel was fully booked" is generated from the dialogue text T.
[0054] (Step S106) When the reason text is generated, the control unit 11 determines the type of reason based on the reason text. A plurality of types of reasons may be set in advance for each purpose by the user or by default. In addition, the type of reason may be generated and set for each cluster by clustering a plurality of reason texts and using a large-scale language model. For example, reasons for failure to make a hotel reservation may be set as "full occupancy," "comparison review," "online reservation," "over budget," and the like.
[0055] The control unit 11 determines which type the reason text corresponds to from among a plurality of types of reasons that have been set. The control unit 11 may cluster the reason text and determine the type of reason corresponding to the corresponding cluster as the type of reason corresponding to the reason text.
[0056] The control unit 11 may also determine the type of reason by using a classification model trained to determine the type of reason corresponding to the reason text. In this case, the control unit 11 inputs the reason text to the classification model, and obtains the type of reason output by the classification model as the type of reason corresponding to the reason text.
[0057] Furthermore, the control unit 11 may determine the type of reason corresponding to the reason text by using a third machine learning model. The third machine learning model is, for example, a large-scale language model, but is not limited to this.
[0058] If the third machine learning model is a large-scale language model, the control unit 11 inputs a prompt to the third machine learning model instructing it to determine the type of reason corresponding to the reason text, and obtains the type of reason output by the third machine learning model as the type of reason corresponding to the reason text.
[0059] The third machine learning model may be stored in the storage unit 12 or in an external device. In the latter case, the control unit 11 inputs the dialogue text T and the prompt to the third machine learning model stored in the external device using an API, and receives the output of the third machine learning model from the external device.
[0060] In the example of FIG. 5, the type of reason why the reservation failed is determined to be "full occupancy" from the reason text.
[0061] In the example of FIG. 4, the control unit 11 determines the type of reason based on the reason text, but the type of reason may be determined based on the dialogue text T. In this case, the control unit 11 may generate the reason text, or may not generate the reason text. In any case, the control unit 11 inputs a prompt to the third machine learning model to instruct the third machine learning model to obtain the reason for the success or failure of the objective from the dialogue text (to generate the reason text or determine the type of reason), and can obtain the output of the third machine learning model as the reason for the success or failure of the objective.
[0062] In addition, the first machine learning model, the second machine learning model, and the third machine learning model may be the same large-scale language model, or may be different large-scale language models. When the first machine learning model, the second machine learning model, and the third machine learning model are the same large-scale language model M, the control unit 11 may input to the large-scale language model M each of a "prompt for instructing to estimate the purpose of the first speaker from the dialogue text T", a "prompt for instructing to determine the success or failure of the purpose estimated in step S103 from the dialogue text T", a "prompt for instructing to generate a text indicating the reason for the success or failure of the purpose from the dialogue text T", and a "prompt for instructing to determine the type of reason corresponding to the reason text", or may input one prompt that combines the four prompts to the large-scale language model M.
[0063] (Step S107) The control unit 11 associates the dialogue ID, the user ID, the dialogue date and time, the dialogue voice A, the purpose of the first speaker, the success or failure of the purpose, the reason text, and the type of reason for success or failure, and stores them in the storage unit 12 as an analysis result R. The dialogue ID is an ID that uniquely identifies the dialogue voice.
[0064] 5, the dialogue ID is "T001", the user ID is "U001", the dialogue date and time is "2024 / 12 / 20 11:50", the dialogue voice is "001.mp3", the purpose is "reservation", the success or failure is "failure", the reason text is "hotel was fully booked", and the reason type is "full". A plurality of such analysis results R are stored in the memory unit 12.
[0065] The control unit 11 may perform an analysis other than the above based on the dialogue voice A or the dialogue text T, and include the analysis result r in the analysis result R. For example, the control unit 11 can use a well-known technique to acquire the call duration, speech ratio, speech rate, number of times of speech, emotion, frequently occurring keywords in speech, or speech volume of the first speaker and the second speaker as the analysis result r based on the dialogue voice A or the dialogue text T.
[0066] (Step S108) The control unit 21 of the user terminal 2 transmits a request to acquire the analysis result R to the dialogue analysis device 1. The acquisition request includes a user ID. The control unit 21 may request acquisition of the analysis result R in response to a user operation, or may periodically request acquisition of the analysis result R.
[0067] (Step S109) When the control unit 11 of the dialogue analysis device 1 receives a request from the user terminal 2 to obtain the analysis result R, the control unit 11 refers to the user ID and transmits the analysis result R corresponding to the user of the user terminal 2 stored in the memory unit 12 to the user terminal 2.
[0068] (Step S110) When the control unit 21 of the user terminal 2 receives the analysis result R from the dialogue analysis device 1, it displays the analysis result R on the display device (output unit 25). The control unit 21 may display the analysis result R in a list, or may display the aggregation result of the analysis result R.
[0069] Fig. 6 is a diagram showing an example of the analysis result screen sc1. The analysis result screen sc1 is a screen that displays a list of the analysis results R. In the example of Fig. 6, the analysis results R of a call between a customer (first speaker) and a hotel operator (second speaker) are displayed in a list, and as the analysis results R, the purpose of the call, whether the purpose was successful or not, a reason text indicating the reason for success or failure, and the type of reason for success or failure are displayed in association with the conversation ID and the conversation date and time.
[0070] In this way, the user terminal 2 displays a list of the first speaker's purpose, whether the purpose was successful, and the reason for success or failure (reason text or type of reason), allowing the user to grasp the details of the calls (conversations) between the first and second speakers individually.
[0071] Fig. 7 is a diagram showing an example of the tally result screen sc2. The tally result screen sc2 is a screen that displays the tally result of the analysis result R. In the example of Fig. 7, it is the tally result of the analysis result R of the call between a customer (first speaker) and a hotel operator (second speaker). The tally result screen sc2 of Fig. 7 displays the number of calls sc21, the number of calls with intention to make a reservation sc22, the number of calls with no intention to make a reservation sc23, the number of completed reservations sc24, the number of incomplete reservations sc25, and the reason for incomplete reservations sc26.
[0072] The number of calls sc21 is the number of calls (conversations) that were the subject of analysis. The number of calls with intent to make a reservation sc22 is the number of calls made by customers with the purpose of making a reservation. The number of calls with no intent to make a reservation sc23 is the number of calls made by customers for purposes other than making a reservation. The number of completed reservations sc24 is the number of calls where customers made a reservation (successfully made a reservation). The number of incomplete reservations sc25 is the number of calls where customers did not make a reservation (failed to make a reservation). The reasons for incompleteness sc26 is the aggregated result of the types of reasons why customers did not make a reservation (failed to make a reservation) for calls where customers did not make a reservation (failed to make a reservation).
[0073] According to Figure 7, the analysis targets 1,000 calls, of which 400 were calls for the purpose of making a reservation, and of the calls for the purpose of making a reservation, 180 ended without the customer making a reservation. It can be seen that the reason for calls that ended without the customer making a reservation was that the hotel was "fully booked" in more than 60% of cases.
[0074] In this way, by the user terminal 2 displaying the compiled results of the first speaker's purpose, whether the purpose was successful, and the reason for success or failure (reason text or type of reason), the user can easily grasp the outline of the call (conversation) between the first and second speakers.
[0075] The tabulation result screen sc2 is not limited to the example in Fig. 7. The control unit 21 can tabulate a plurality of analysis results R for each arbitrary attribute and display the number or the ratio. The attributes are, for example, purpose, success or failure, reason for success or failure, time period, day, week, month, day of the week, or analysis results r (call time, speech ratio, speech rate, number of speeches, emotion, frequent keywords in speech, speech volume, etc. of the first speaker and the second speaker), but are not limited thereto.
[0076] <Summary> As described above, according to this embodiment, a dialogue analysis device 1 is provided with one or more control units 11, and the control unit 11 utilizes a first machine learning model to infer the purpose of a first speaker from a dialogue text T between a first speaker and a second speaker, utilizes a second machine learning model to determine whether the purpose has been achieved from the dialogue text T, utilizes a third machine learning model to obtain the reason for the success or failure of the purpose from the dialogue text T, and can realize a dialogue analysis device 1 that displays the success or failure of the purpose and the reason on a user terminal 2.
[0077] 6 and 7, the dialogue analysis device 1 can visualize whether the dialogue goal is achieved and the reasons for that, based on the dialogue text T. A user can improve his / her work by referring to the visualized dialogue goal success / failure and the reasons for that success / failure.
[0078] <Additional Notes> The present embodiment includes the following disclosure.
[0079] (Appendix 1) A dialogue analysis device including one or more control units, The control unit is Utilizing a first machine learning model to estimate a purpose of the first speaker from a dialogue text of the first speaker and a second speaker; Using a second machine learning model, determine whether the objective is achieved from the dialogue text; and Utilizing a third machine learning model to obtain reasons for the success or failure of the objective from the dialogue text; Displaying the success or failure of the objective and the reason for it on the user terminal Dialogue analysis device.
[0080] (Appendix 2) At least one of the first machine learning model, the second machine learning model, and the third machine learning model is a large-scale language model. 2. A dialogue analysis apparatus according to claim 1.
[0081] (Appendix 3) The first machine learning model, the second machine learning model, and the third machine learning model are the same machine learning model. 2. A dialogue analysis apparatus according to claim 1.
[0082] (Appendix 4) The control unit inputs a prompt to the first machine learning model to instruct the first machine learning model to estimate a purpose of the first speaker from the dialogue text. 2. A dialogue analysis apparatus according to claim 1.
[0083] (Appendix 5) The control unit inputs a prompt to the second machine learning model, the prompt instructing the second machine learning model to determine whether the objective is achieved from the dialogue text. 2. A dialogue analysis apparatus according to claim 1.
[0084] (Appendix 6) The control unit inputs a prompt to the third machine learning model to instruct the third machine learning model to obtain a reason for the success or failure of the objective from the dialogue text. 2. A dialogue analysis apparatus according to claim 1.
[0085] (Appendix 7) The control unit generates the dialogue text from the dialogue voices of the first speaker and the second speaker by using a fourth machine learning model. 2. A dialogue analysis apparatus according to claim 1.
[0086] (Appendix 8) The reason for the success or failure of the objective is a text indicating the reason or a type of reason. 2. A dialogue analysis apparatus according to claim 1.
[0087] (Appendix 9) The control unit generates a text indicating a reason for the success or failure of the objective, and determines a type of the reason from the text indicating the reason by clustering. 2. A dialogue analysis apparatus according to claim 1.
[0088] (Appendix 10) The control unit causes the user terminal to display a compilation result of the success or failure of the objective and the reason thereof. 2. A dialogue analysis apparatus according to claim 1.
[0089] (Appendix 11) The purpose includes at least one of a reservation, a purchase, a sale, an inquiry, an inventory check, and an appointment. 2. A dialogue analysis apparatus according to claim 1.
[0090] (Appendix 12) The first speaker is a customer and the second speaker is a business operator. 2. A dialogue analysis apparatus according to claim 1.
[0091] (Appendix 13) The first speaker is a business operator and the second speaker is a customer. 2. A dialogue analysis apparatus according to claim 1.
[0092] (Appendix 14) The dialogue text has speaker information for each sentence. 2. A dialogue analysis apparatus according to claim 1.
[0093] (Appendix 15) A dialogue analysis device including one or more control units, Utilizing a first machine learning model to estimate a purpose of the first speaker from a dialogue text of the first speaker and a second speaker; Using a second machine learning model, determine whether the objective is achieved from the dialogue text; and Utilizing a third machine learning model to obtain reasons for the success or failure of the objective from the dialogue text; Displaying the success or failure of the objective and the reason for it on the user terminal Dialogue analysis methods.
[0094] (Appendix 16) A dialogue analysis device including one or more control units, Utilizing a first machine learning model to estimate a purpose of the first speaker from a dialogue text of the first speaker and a second speaker; Using a second machine learning model, determine whether the objective is achieved from the dialogue text; and Utilizing a third machine learning model to obtain reasons for the success or failure of the objective from the dialogue text; Displaying the success or failure of the objective and the reason for it on the user terminal program.
[0095] The embodiments disclosed herein are illustrative in all respects and should not be considered as limiting. The scope of the present invention is indicated by the claims, not by the above meaning, and is intended to include all modifications within the meaning and scope of the claims. The present invention is not limited to the above-mentioned embodiments, and various modifications are possible within the scope of the claims, and embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention. [Explanation of symbols]
[0096] 1: Dialogue analysis device 2: User terminal 11, 21: Control unit 12,22: Storage part 13,23: Communications Department A: Dialogue voice T: Dialogue text R:Analysis results
Claims
1. A dialogue analysis device including one or more control units, The control unit is Utilizing a first machine learning model to estimate a purpose of the first speaker from a dialogue text of the first speaker and a second speaker; Using a second machine learning model, determine whether the objective is achieved from the dialogue text; and Utilizing a third machine learning model to obtain reasons for the success or failure of the objective from the dialogue text; Displaying the success or failure of the objective and the reason for it on the user terminal Dialogue analysis device.
2. At least one of the first machine learning model, the second machine learning model, and the third machine learning model is a large-scale language model. The conversation analysis device according to claim 1 .
3. The first machine learning model, the second machine learning model, and the third machine learning model are the same machine learning model. The conversation analysis device according to claim 1 .
4. The control unit inputs a prompt to the first machine learning model to instruct the first machine learning model to estimate a purpose of the first speaker from the dialogue text. The conversation analysis device according to claim 1 .
5. The control unit inputs a prompt to the second machine learning model, the prompt instructing the second machine learning model to determine whether the objective is achieved from the dialogue text. The conversation analysis device according to claim 1 .
6. The control unit inputs a prompt to the third machine learning model, the prompt instructing the third machine learning model to obtain a reason for the success or failure of the objective from the dialogue text. The conversation analysis device according to claim 1 .
7. The control unit generates the dialogue text from the dialogue speech of the first speaker and the second speaker by using a fourth machine learning model. The conversation analysis device according to claim 1 .
8. The reason for the success or failure of the objective is a text indicating the reason or a type of reason. The conversation analysis device according to claim 1 .
9. The control unit generates a text indicating a reason for the success or failure of the objective, and determines a type of the reason from the text indicating the reason by clustering. The conversation analysis device according to claim 1 .
10. The control unit causes the user terminal to display a compilation result of the success or failure of the objective and the reason thereof. The conversation analysis device according to claim 1 .
11. The purpose includes at least one of a reservation, a purchase, a sale, an inquiry, an inventory check, and an appointment. The conversation analysis device according to claim 1 .
12. The first speaker is a customer and the second speaker is a business operator. The conversation analysis device according to claim 1 .
13. The first speaker is a business operator and the second speaker is a customer. The conversation analysis device according to claim 1 .
14. The dialogue text has speaker information for each sentence. The conversation analysis device according to claim 1 .
15. A dialogue analysis device including one or more control units, Utilizing a first machine learning model to estimate a purpose of the first speaker from a dialogue text of the first speaker and a second speaker; Using a second machine learning model, determine whether the objective is achieved from the dialogue text; and Utilizing a third machine learning model to obtain reasons for the success or failure of the objective from the dialogue text; Displaying the success or failure of the objective and the reason for it on the user terminal Dialogue analysis methods.
16. A dialogue analysis device including one or more control units, Utilizing a first machine learning model to estimate a purpose of the first speaker from a dialogue text of the first speaker and a second speaker; Using a second machine learning model, determine whether the objective is achieved from the dialogue text; and Utilizing a third machine learning model to obtain reasons for the success or failure of the objective from the dialogue text; Displaying the success or failure of the objective and the reason for it on the user terminal program.
Citation Information
Patent Citations
Information publication device
JP1997081632A
Conversation processing system
JP2002023783A
Interaction device with himself or herself, chatbot, and robot
JP2020154378A
Information processing device, information processing system, information processing method, and program
JP2022012457A
Business support device, business support method and program
JP2024001473A