Program, information processing device, information processing system, information processing method, and information processing terminal
Patent Information
- Application Number
- JP2022169442
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-10-21
- Publication Date
- 2025-05-16
AI Technical Summary
Existing technologies fail to confirm the topics discussed during a dialogue between speakers, making it difficult to understand the content of interactions.
A system that processes dialogue information to extract section audio data, identify topics, and generate summaries based on emotional and impression analysis, allowing users to specify and present the topics discussed.
Enables users to accurately determine the topics communicated in a dialogue, enhancing understanding and management of interaction content.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0004] , ,
[0005]
[0001] The present disclosure relates to a program, an information processing apparatus, an information processing system, an information processing method, and an information processing terminal.
Background Art
[0002] Online interactive services conducted among multiple users are known. Patent Document 1 discloses a method for assisting in realizing more efficient business activities while considering objective indicators.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] There is a problem that it is not possible to confirm what topics the speakers communicated about in the dialogue. Therefore, the present disclosure has been made to solve the above problems, and its object is to provide a technique for confirming what topics the speakers communicated about in the dialogue.
Means for Solving the Problems
[0005] A program comprising a processor and a memory unit, which causes the computer to process information relating to a dialogue between a first user and a second user, the program causing the processor to execute: a reception step of receiving audio data relating to the dialogue; an audio extraction step of extracting multiple segment audio data for each utterance segment from the audio data received in the reception step; a segment identification step of identifying one or more segment audio data from among the multiple segment audio data that are related to a first topic relating to a predetermined topic; and a summarization step of generating a summary text that summarizes the text information contained in one or more segment audio data based on the one or more segment audio data identified in the segment identification step and the first topic. [Effects of the Invention]
[0006] According to this disclosure, it is possible to identify what topics the speakers communicated about in the conversational service. [Brief explanation of the drawing]
[0007] [Figure 1] This is a block diagram showing the functional configuration of System 1. [Figure 2] This block diagram shows the functional configuration of Server 10. [Figure 3] This is a block diagram showing the functional configuration of the first user terminal 20. [Figure 4] This is a block diagram showing the functional configuration of the second user terminal 30. [Figure 5] This is a diagram showing the 50 functional configurations of the CRM system. [Figure 6] This diagram shows the data structure of user table 1012. [Figure 7] This diagram shows the data structure of organization table 1013. [Figure 8] This diagram shows the data structure of dialogue table 1014. [Figure 9] This diagram shows the data structure of label table 1015. [Figure 10] It is a diagram showing the data structure of the voice interval table 1016. [Figure 11] It is a diagram showing the data structure of the topic relevance degree table 1017. [Figure 12] It is a diagram showing the data structure of the emotion condition master 1021. [Figure 13] It is a diagram showing the data structure of the speaker type master 1022. [Figure 14] It is a diagram showing the data structure of the topic master 1023. [Figure 15] It is a diagram showing the data structure of the customer table 5012. [Figure 16] It is a flowchart showing the operation of the emotion analysis process. [Figure 17] It is a flowchart showing the operation of the impression analysis process. [Figure 18] It is a flowchart showing the operation of the topic analysis process. [Figure 19] It is a flowchart showing the operation of the topic presentation process. [Figure 20] It is an example screen showing the operation of the topic presentation process. [Figure 21] It is a block diagram showing the basic hardware configuration of the computer 90.
Mode for Carrying Out the Invention
[0008] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In all the drawings for describing the embodiments, common components are denoted by the same reference numerals, and repeated descriptions are omitted. Note that the following embodiments do not unduly limit the content of the present disclosure described in the claims. Also, not all of the components shown in the embodiments are essential components of the present disclosure. Further, each figure is a schematic diagram and is not necessarily shown exactly.
[0009] <Configuration of System 1> The system 1 in the present disclosure is an information processing system that provides an online interactive service (online interaction service) between a first user who is an operator and a second user who is a customer. Note that the system 1 in the present disclosure may also be capable of providing an interactive service that is conducted online among three or more users, including one or more other users in addition to the first user and the second user. The system 1 includes information processing devices of a server 10, a first user terminal 20, a second user terminal 30, a CRM system 50, and a voice server (PBX) 60, which are connected via a network N. FIG. 1 is a block diagram showing the functional configuration of the system 1. FIG. 2 is a block diagram showing the functional configuration of the server 10. FIG. 3 is a block diagram showing the functional configuration of the first user terminal 20. FIG. 4 is a block diagram showing the functional configuration of the second user terminal 30. FIG. 5 is a block diagram showing the functional configuration of the CRM system 50.
[0010] Each information processing device is composed of a computer having an arithmetic device and a storage device. The basic hardware configuration of the computer and the basic functional configuration of the computer realized by the hardware configuration will be described later. For each of the server 10, the first user terminal 20, the second user terminal 30, the CRM system 50, and the voice server (PBX) 60, descriptions that overlap with the basic hardware configuration of the computer and the basic functional configuration of the computer described later will be omitted.
[0011] <Configuration of Server 10> The server 10 is an information processing device that provides a service for storing and managing data related to an interaction (interaction data) conducted between the first user and the second user. The server 10 includes a storage unit 101 and a control unit 104.
[0012] <Configuration of Storage Unit 101 of Server 10> The storage unit 101 of the server 10 includes an application program 1011, an emotion evaluation model 1031, an impression evaluation model 1032, a first impression evaluation model 1033, a second impression evaluation model 1034, a summary model 1035, a user table 1012, an organization table 1013, a dialogue table 1014, a label table 1015, a speech segment table 1016, a topic relevance table 1017, an emotion condition master 1021, a speaker type master 1022, and a topic master 1023.
[0013] Application program 1011 is a program that causes the control unit 104 of server 10 to function as individual functional units. Application program 1011 includes applications such as a web browser application.
[0014] The emotion evaluation model 1031 is a model that takes audio data, video data, and text data related to the user's statements in the audio or video data as input data, and outputs numerical intensities and values for multiple emotional states.
[0015] The impression evaluation model 1032 is a model that takes audio data, video data, and text data related to the user's statements in the audio or video data as input data, and outputs numerical strengths and values for multiple impressions.
[0016] The first impression evaluation model 1033 is a model that takes audio data, video data, and text data related to the content of user statements in audio or video data as input data, and outputs dialogue features related to the speaker's way of speaking. Dialogue features are features related to at least one of the following aspects of the speaker's way of speaking: speaking speed, intonation, number of polite expressions, number of fillers, and number of grammatical utterances.
[0017] The second impression evaluation model 1034 is a model that takes dialogue features as input data and outputs numerical strengths and values for each of multiple impressions.
[0018] User Table 1012 is a table that stores and manages information about member users (hereinafter referred to as "users") who use the service. When a user registers to use the service, their information is stored in a new record in User Table 1012. This allows the user to use the services related to this disclosure. User table 1012 is a table with User ID as the primary key and containing columns for User ID, CRM ID, Organization ID, Username, and User Attributes. Figure 6 shows the data structure of user table 1012.
[0019] The User ID is an item that stores user identification information to identify a user. User identification information is an item that is set to a unique value for each user. The CRMID is an item in the CRM system 50 that stores user identification information to identify a user. Users can access CRM services by logging into the CRM system 50 using their CRMID. The user ID in server 10 is associated with the CRMID in the CRM system 50. The Organization ID is an item that stores organizational identification information used to identify an organization. The username field is used to store the user's real name. However, the username can also be a nickname or any other string of characters. User attributes are items that store information about the user's attributes, such as age, gender, place of origin, dialect, and occupation (sales, customer support, etc.). In addition to information about the individual user's attributes, user attributes may also include information about the organization, company, or group to which the user belongs, such as industry, business size, and sales volume.
[0020] Organization table 1013 is a table that stores and manages information about the organizations to which a user belongs (organizational information). Organizations include any organization or group, such as companies, corporations, corporate groups, clubs, and various other organizations. Organizations may also be defined for more detailed subgroups, such as departments within a company (sales department, general affairs department, customer support department). Organization table 1013 is a table with Organization ID as the primary key and containing columns for Organization ID, Organization Name, and Organization Attributes. Figure 7 shows the data structure of organization table 1013.
[0021] The Organization ID is an item that stores organizational identification information used to identify an organization. Organizational identification information is an item with a unique value assigned to each piece of organizational information. The organization name is an item that stores the name of the organization. The organization name can be set to any string. Organizational attributes are items that store information about the attributes of an organization, such as the type of organization (company, corporate group, other organization, etc.) and industry (real estate, finance, etc.).
[0022] The dialogue table 1014 is a table for storing and managing information (dialogue information) related to conversations that take place between users and customers. Dialogue table 1014 is a table with dialogue ID as the primary key and containing columns for dialogue ID, user ID, customer ID, dialogue category, call type, audio data, and video data. Figure 8 shows the data structure of the dialogue table 1014.
[0023] The dialogue ID is an item that stores dialogue identification information to identify a dialogue. Dialogue identification information is an item with a unique value assigned to each dialogue. The User ID is an item that stores user identification information used to identify a user in interactions between the user and the customer. Multiple User IDs may be associated with each piece of interaction information. The customer ID is an item that stores user identification information used to identify a customer in interactions between the user and the customer. Multiple customer user IDs may be associated with each interaction. The dialogue category is an item that stores the type (category) of dialogue that took place between the user and the customer. Dialogue data is classified by dialogue category. Depending on the purpose of the dialogue between the user and the customer, values such as telephone operator, telemarketing, customer support, and technical support are stored in the dialogue category. The "Incoming / Outgoing Call Type" field stores information to distinguish whether a conversation between a user and a customer was initiated by the user (outbound) or received by the user (inbound). Additionally, when a conversation involves three or more users, a "Room" type is also stored. The audio data item stores audio data collected by the microphone. It may also store reference information (paths) to audio data files located elsewhere. The audio data format can be any data format, such as AAC, ATRAC, mp3, or mp4. The audio data may be in a format in which identifiers are set that allow the user's voice and the customer's voice to be independently identified. In this case, the control unit 104 of the server 10 can perform independent analysis processing on the user's voice and the customer's voice. Furthermore, the user IDs of the user and the customer can be identified based on the user and customer's audio data. In this disclosure, video data containing audio information may be used instead of audio data. Furthermore, audio data in this disclosure also includes audio data contained within video data. The video data item stores video data captured by a camera or similar device. It may also store reference information (paths) to video data files located elsewhere. The video data format can be any format, such as MP4, MOV, WMV, AVI, or AVCHD. The video data may be in a format in which user videos and customer videos are each assigned identifiers that allow them to be independently identified. In this case, the control unit 104 of the server 10 can perform independent analysis processing on user videos and customer videos. Furthermore, it can identify the user IDs of users and customers based on the user and customer video data.
[0024] Label table 1015 is a table for storing and managing information about labels (label information). Label table 1015 is a table that has columns for dialogue ID and label data. Figure 9 shows the data structure of label table 1015.
[0025] The dialogue ID is an item that stores dialogue identification information to identify a dialogue. Label data is an item that stores label information for managing interactions. Label information includes additional information for managing interaction information, such as classification names, labels, classification labels, and tags. Label data can be a string representing the name of the label information, or it can be a label ID or similar that references the name of the label information stored in another table. Label data includes classification information corresponding to the emotional state of the speaker in a particular dialogue. Classification data includes classification information for classifying the quality of the speaker's response in a particular dialogue.
[0026] The speech segment table 1016 is a table for storing and managing information about multiple speech segments included in dialogue information (speech segment information). The audio segment table 1016 is a table with segment ID as the primary key and containing the following columns: segment ID, dialogue ID, speaker ID, start date and time, end date and time, segment audio data, segment video data, segment spoken text, emotion data, impression data, and topic ID. Figure 10 shows the data structure of the audio interval table 1016.
[0027] The section ID is an item that stores section identification information used to identify audio sections. Each section of audio information has a unique value assigned to it. The dialogue ID is an item that stores dialogue identification information to identify the dialogue to which the audio segment information is associated. The speaker ID is an item that stores speaker identification information to identify the speaker to whom the speech segment information is associated. Specifically, the speaker ID is an item that stores the user IDs of multiple users who participated in the dialogue. The start date and time field stores the start date and time of the audio and video sections. The end date and time field stores the end date and time of the audio and video segments. Section audio data is an item that stores audio data contained within an audio section. It may also store reference information (path) to audio data files located elsewhere. Alternatively, it may store references to audio data for the period from the start date and time to the end date and time of the audio data in dialogue table 1014, based on the start and end dates and times. Furthermore, section audio data may also include audio data contained within section video data. The audio data format can be any format, such as AAC, ATRAC, mp3, or mp4. Section video data is an item that stores video data included in an audio section. It may also store reference information (path) to video data files located elsewhere. Alternatively, it may store references to video data for the period from the start date and time to the end date and time of the video data in the dialogue table 1014, based on the start date and time and end date and time. The video data format can be any format you like, such as MP4, MOV, WMV, AVI, or AVCHD. Section reading text is an item that stores text information of what the speaker said in the section audio data contained within an audio section. Specifically, section reading text may be generated based on section audio data and section video data, either manually or using a learning model such as machine learning or deep learning. Emotional data is an item that stores the speaker's emotional state within a speech interval. Emotional data is a multidimensional scale (emotional vector) relating to multiple emotional states of the speaker, such as interest / excitement, joy, surprise, anxiety, anger, disgust, contempt, fear, shame, and guilt. Emotional data quantitatively represents the speaker's emotional state within a dialogue interval, expressing the intensity and numerical value of each of the multiple emotional states (dimensions). Emotional data can also be structured to calculate and store an emotional scalar representing the intensity of a one-dimensional emotion based on the emotional vector. Impression data consists of items that record the impression a speaker makes during a speech segment. It is a multidimensional scale (vector) relating to multiple different impressions a speaker gives, including like, dislike, loud, difficult to understand, polite, unclear, timid, nervous, intimidating, violent, and sexual. It quantitatively represents the impression a speaker gives during a dialogue segment, expressing the intensity and numerical value of each of these multiple impressions (dimensions). The topic ID is an item that stores topic identification information associated with a particular audio segment.
[0028] Topic relevance table 1017 is a table for storing and managing information about the topic relevance of each audio segment (topic relevance information). Topic relevance table 1017 is a table that has columns for interval ID, topic ID, and relevance. Figure 11 shows the data structure of topic relevance table 1017.
[0029] The section ID is an item that stores the section identification information of the target audio section. The topic ID is an item that stores topic identification information used to identify a topic. The relevance level is an item that stores information about the relevance of each topic identification information, which is specified by the topic ID, within the audio segments included in the dialogue information. For each audio segment, it stores a numerical value indicating the degree of relevance with the topic specified by the topic ID. A higher relevance level indicates a stronger connection between the dialogue information and the topic.
[0030] The Emotional Condition Master 1021 is a table for storing and managing information related to emotional conditions (emotional condition information). The emotion condition master 1021 is a table that has columns for emotion condition and label data. Figure 12 shows the data structure of the emotion condition master 1021.
[0031] The emotional conditions section stores conditions related to emotional data. Specifically, it stores conditions for emotional data thresholds, mean values, and regression coefficients when regression analysis is performed. Label data is an item that stores label information associated with emotional conditions.
[0032] Speaker type master 1022 is a table for storing and managing information related to impression conditions (impression condition information). Speaker type master 1022 is a table that has columns for impression conditions and speaker type. Figure 13 shows the data structure of the speaker type master 1022.
[0033] The impression conditions section stores the conditions related to impression data. Specifically, it stores conditions such as the threshold, mean, and regression coefficients when regression analysis is performed on impression data. Speaker type is an item that stores information about speaker types associated with impression conditions. Speaker type is a classification of the impression a speaker gives to the listener, such as assertive, reserved, serious, friendly, assertive, and emotional.
[0034] Topic Master 1023 is a table for storing and managing information about topics (topic information). Topic Master 1023 is a table with Topic ID as the primary key and containing columns for Topic ID and Keyword. Figure 14 shows the data structure of topic master 1023.
[0035] The Topic ID is an item that stores topic identification information to identify a topic. The topic identification information is an item with a unique value assigned to each topic. A keyword is an item that stores multiple keywords associated with a topic. Specifically, multiple keywords are associated with one topic.
[0036] <Configuration of the control unit 104 of server 10> The control unit 104 of the server 10 comprises a user registration control unit 1041, an emotion analysis unit 1042, an impression analysis unit 1043, a topic processing unit 1044, and a learning unit 1051. The control unit 104 realizes each functional unit by executing the application program 1011 stored in the storage unit 101.
[0037] The user registration control unit 1041 processes information of users who wish to use the services related to this disclosure and stores it in the user table 1012. Information stored in the user table 1012 is obtained when a user opens a web page operated by the service provider from any information processing terminal, enters information into a designated input form, and sends it to the server 10. The user registration control unit 1041 stores the received information in a new record in the user table 1012, and user registration is completed. As a result, users stored in the user table 1012 can use the service. Prior to the registration of user information in the user table 1012 by the user registration control unit 1041, the service provider may perform a prescribed review and restrict whether or not the user can use the service. The user ID can be any string or number that can identify the user, and may be any string or number that the user wishes, or the user registration control unit 1041 may automatically set any string or number.
[0038] The emotion analysis unit 1042 performs emotion analysis processing. Details will be described later.
[0039] The impression analysis unit 1043 performs impression analysis processing. Details will be described later.
[0040] The topic processing unit 1044 performs topic definition processing, topic analysis processing, and topic presentation processing. Details will be described later.
[0041] The learning unit 1051 executes the learning process.
[0042] <Configuration of the first user terminal 20> The first user terminal 20 is an information processing device operated by the first user who uses the service. The first user terminal 20 may be, for example, a stationary PC (Personal Computer), a laptop PC, or a mobile device such as a smartphone or tablet. It may also be a wearable device such as an HMD (Head Mount Display) or a smartwatch. The first user terminal 20 includes a storage unit 201, a control unit 204, an input device 206, and an output device 208.
[0043] <Configuration of the storage unit 201 of the first user terminal 20> The storage unit 201 of the first user terminal 20 includes a first user ID 2011 and an application program 2012.
[0044] The first user ID 2011 stores the user identification information of the first user. The user transmits the first user ID 2011 from the first user terminal 20 to the voice server (PBX) 60. The voice server (PBX) 60 identifies the first user based on the first user ID 2011 and provides the services related to this disclosure to the first user. The first user ID 2011 includes information such as a session ID that is temporarily assigned by the voice server (PBX) 60 to identify the user using the first user terminal 20.
[0045] The application program 2012 may be pre-stored in the memory unit 201, or it may be configured to be downloaded from a web server operated by the service provider via a communication interface. Application Program 2012 includes applications such as web browser applications. Application program 2012 includes an interpreted programming language such as JavaScript (registered trademark) that runs on a web browser application stored on the first user terminal 20.
[0046] <Configuration of the control unit 204 of the first user terminal 20> The control unit 204 of the first user terminal 20 comprises an input control unit 2041 and an output control unit 2042. The control unit 204 realizes each functional unit by executing an application program 2012 stored in the storage unit 201.
[0047] <Configuration of the input device 206 of the first user terminal 20> The input device 206 of the first user terminal 20 includes a camera 2061, a microphone 2062, a position information sensor 2063, a motion sensor 2064, and a keyboard 2065.
[0048] <Configuration of output device 208 of the first user terminal 20> The output device 208 of the first user terminal 20 includes a display 2081 and a speaker 2082.
[0049] <Configuration of the second user terminal 30> The second user terminal 30 is an information processing device operated by a second user who uses the service. The second user terminal 30 may be, for example, a mobile device such as a smartphone or tablet, or a stationary PC (Personal Computer) or laptop PC. It may also be a wearable device such as an HMD (Head Mount Display) or a smartwatch. The second user terminal 30 includes a storage unit 301, a control unit 304, an input device 306, and an output device 308.
[0050] <Configuration of the storage unit 301 of the second user terminal 30> The storage unit 301 of the second user terminal 30 includes an application program 3012 and a phone number 3013.
[0051] The application program 3012 may be stored in the storage unit 301 in advance, or may be configured to be downloaded from a web server or the like operated by a service provider via a communication IF. The application program 3012 includes applications such as a web browser application. The application program 3012 includes an interpreter-type programming language such as JavaScript (registered trademark) that is executed on the web browser application stored in the second user terminal 30.
[0052] <Configuration of the control unit 304 of the second user terminal 30> The control unit 304 of the second user terminal 30 includes an input control unit 3041 and an output control unit 3042. The control unit 304 realizes each functional unit by executing the application program 3012 stored in the storage unit 301.
[0053] <Configuration of the input device 306 of the second user terminal 30> The input device 306 of the second user terminal 30 includes a camera 3061, a microphone 3062, a position information sensor 3063, a motion sensor 3064, and a touch device 3065.
[0054] <Configuration of the output device 308 of the second user terminal 30> The output device 308 of the second user terminal 30 includes a display 3081 and a speaker 3082.
[0055] <Configuration of the CRM system 50> The CRM system 50 is an information processing device managed and operated by an operator (CRM operator) that provides CRM (Customer Relationship Management, second user relationship management) services. Examples of CRM services include SalesForce, HubSpot, Zoho CRM, kintone, etc. The CRM system 50 includes a storage unit 501 and a control unit 504.
[0056] <Configuration of the storage unit 501 of the CRM system 50> The storage unit 501 of the CRM system 50 includes an application program 5011 and a customer table 5012.
[0057] The application program 5011 is a program for causing the control unit 504 of the CRM system 50 to function as each functional unit. The application program 5011 includes applications such as a web browser application.
[0058] The customer table 5012 is a table for storing and managing user information (customer information) related to customers. The customer table 5012 is a table having columns for customer ID, user ID, name, phone number, and speaker type, with the customer ID as the primary key. Figure 15 is a diagram showing the data structure of the customer table 5012.
[0059] The customer ID is an item for storing the user identification information of the customer. The user identification information is an item for which a unique value is set for each customer. The user ID is an item for storing the user identification information of the user who manages the customer. The name is an item for storing the name of the customer. The phone number is an item for storing the phone number of the customer. The user can make a call to the customer's phone number from the first user terminal 20 by accessing the website provided by the CRM system, selecting the customer to whom the call is to be made, and performing a predetermined operation such as "making a call". The speaker type is an item for storing the speaker type of the user identified by the customer ID.
[0060] <Configuration of the control unit 504 of the CRM system 50> The control unit 504 of the CRM system 50 includes a user registration control unit 5041. The control unit 504 realizes each functional unit by executing the application program 5011 stored in the storage unit 501.
[0061] The user registration control unit 5041 performs a process of storing customer information in the customer table 5012 in the service according to the present disclosure. The information stored in the customer table 5012 is that the user opens a web page or the like operated by the service provider from an arbitrary information processing terminal, inputs information into a predetermined input form, and transmits it to the CRM system 50. The user registration control unit 5041 stores the received information in a new record of the customer table 5012, and the registration of the customer is completed. As a result, the customer information is stored in association with the user ID of the user who manages the customer. The customer ID may be any character string or number that can identify the user, any character string or number desired by the user, or the user registration control unit 5041 may automatically set any character string or number.
[0062] <Configuration of the voice server (PBX) 60> The voice server (PBX) 60 is an information processing device that functions as a switch that enables a conversation between the first user terminal 20 and the second user terminal 30 by connecting the network N and the telephone network T to each other. The voice server (PBX) 60 includes a storage unit 601.
[0063] <Configuration of the storage unit 601 of the voice server (PBX) 60> The storage unit 601 of the voice server (PBX) 60 includes an application program 6011.
[0064] Application program 6011 is a program that causes the control unit 604 of the voice server (PBX) 60 to function as a functional unit. Application program 6011 includes applications such as web browser applications.
[0065] <System 1 operation> The following describes each process in System 1. Figure 16 is a flowchart showing the operation of the sentiment analysis process. Figure 17 is a flowchart showing the operation of the impression analysis process. Figure 18 is a flowchart showing the operation of the topic analysis process. Figure 19 is a flowchart showing the operation of the topic presentation process. Figure 20 shows an example screen illustrating the operation of the topic presentation process.
[0066] <Outgoing call processing> Outgoing call processing is the process of making an outgoing call from a user (first user) to a customer (second user).
[0067] <Overview of outgoing call processing> The outgoing call process is a series of operations in which the user selects a customer to whom they wish to make a call from among multiple customers displayed on the screen of the first user terminal 20, and then makes a call to that customer. In this disclosure, the case in which the second user is selected as the customer will be explained as an example.
[0068] <Details of the outgoing call process> This section describes the outgoing call processing of System 1 when a user makes a call to a customer.
[0069] When a user makes a call to a customer, the following process is executed in System 1.
[0070] The user operates the first user terminal 20 to launch a web browser and access the website of the CRM service provided by the CRM system 50. By opening the customer management screen provided by the CRM service, the user can view a list of their customers on the display 2081 of the first user terminal 20. Specifically, the first user terminal 20 sends a request to the CRM system 50 to display a list of CRMID2013 and customers. Upon receiving the request, the CRM system 50 searches the customer table 5012 and sends information about the user's customers, such as customer ID, name, phone number, customer attributes, customer organization name, and customer organization attributes, to the first user terminal 20. The first user terminal 20 displays the received customer information on its display 2081.
[0071] The user selects the customer they wish to call (the second user) from the list of customers displayed on the display 2081 of the first user terminal 20. With the customer selected, the user sends a request including the phone number to the CRM system 50 by pressing the "Call" button or the phone number button displayed on the display 2081 of the first user terminal 20. Upon receiving the request, the CRM system 50 sends the request including the phone number to the server 10. Upon receiving the request, the server 10 sends a call request to the voice server (PBX) 60. When the voice server (PBX) 60 receives the call request, it makes a call to the second user terminal 30 based on the received phone number.
[0072] Accordingly, the first user terminal 20 controls the speaker 2082 and other components to make a sound indicating that a call is being made by the voice server (PBX) 60. In addition, the display 2081 of the first user terminal 20 displays information indicating that a call is being made to the customer by the voice server (PBX) 60. For example, the display 2081 of the first user terminal 20 may display the words "Calling".
[0073] The customer can make the second user terminal 30 ready for conversation by lifting the handset (not shown) on the second user terminal 30 or by pressing the "Receive" button displayed on the input device 306 of the second user terminal 30 when an incoming call is received. Accordingly, the voice server (PBX) 60 sends information indicating that a response has been made by the second user terminal 30 (hereinafter referred to as a "response event") to the first user terminal 20 via the server 10, CRM system 50, etc. This enables the user and the customer to communicate using the first user terminal 20 and the second user terminal 30, respectively, allowing for dialogue between the user and the customer. Specifically, the user's voice, picked up by the microphone 2062 of the first user terminal 20, is output from the speaker 3082 of the second user terminal 30. Similarly, the customer's voice, picked up by the microphone 3062 of the second user terminal 30, is output from the speaker 2082 of the first user terminal 20.
[0074] When the display 2081 of the first user terminal 20 becomes interactive, it receives a response event and displays information indicating that an interaction is taking place. For example, the display 2081 of the first user terminal 20 may display the words "Responding".
[0075] <Incoming Call Processing> Incoming call processing is the process by which a user receives an incoming call from a customer.
[0076] <Overview of incoming call processing> Incoming call processing is a series of processes that occur when a customer makes a call to a user while the user has launched an application on the first user terminal 20, and the user receives the call.
[0077] <Details of incoming call processing> This section describes the incoming call processing of System 1 when a user receives a call from a customer.
[0078] When a user receives a call from a customer, the following process is executed in System 1.
[0079] The user operates the first user terminal 20 to launch a web browser and access the website for the CRM service provided by the CRM system 50. At this time, the user is assumed to be logged into the CRM system 50 with their account in the web browser and waiting. The user only needs to be logged into the CRM system 50 and may be performing other tasks related to the CRM service.
[0080] The customer operates the second user terminal 30, enters a predetermined telephone number assigned to the voice server (PBX) 60, and makes a call to the voice server (PBX) 60. The voice server (PBX) 60 receives the call from the second user terminal 30 as an incoming event.
[0081] The voice server (PBX) 60 sends an incoming call event to the server 10. Specifically, the voice server (PBX) 60 sends an incoming call request to the server 10, including the customer's telephone number 3011. The server 10 sends the incoming call request to the first user terminal 20 via the CRM system 50. Accordingly, the first user terminal 20 controls the speaker 2082 and other components to make a sound indicating that an incoming call is being received by the voice server (PBX) 60. The display 2081 of the first user terminal 20 displays information indicating that an incoming call is being received from a customer by the voice server (PBX) 60. For example, the display 2081 of the first user terminal 20 may display the words "Incoming Call".
[0082] The first user terminal 20 accepts responses from the user. Responses are performed, for example, by the user lifting the receiver (not shown) on the first user terminal 20, or by the user using the mouse 2066 to press a button labeled "Answer Call" on the display 2081 of the first user terminal 20. When the first user terminal 20 receives a response operation, it sends a response request to the voice server (PBX) 60 via the CRM system 50 and server 10. The voice server (PBX) 60 receives the transmitted response request and establishes voice communication. As a result, the first user terminal 20 becomes ready to communicate with the second user terminal 30. The display 2081 of the first user terminal 20 displays information indicating that a conversation is taking place. For example, the display 2081 of the first user terminal 20 may display the words "Conversation in progress".
[0083] <Variations of outgoing and incoming call processing> The method by which the first user can become interactive with the second user is not limited to outgoing or incoming call processing; any method is permitted to enable interaction between the first and second users. For example, a virtual interaction space called a "room" can be created on the server 10 for interaction between the first and second users, and the first and second users can become interactive by accessing this room via a web browser or application program stored on the first user terminal 20 and the second user terminal 30. In this case, the voice server (PBX) 50 is not required. Specifically, the first user, who is the host of the dialogue, operates the input device 206 of the first user terminal 20 and sends a request to the server 10 to hold a dialogue. Upon receiving the request, the control unit 104 of the server 10 issues room identification information, such as a unique room ID, and sends a response to the first user terminal 20. The first user sends the received room identification information to the second user, who is the dialogue partner, via email, chat, or any other means of communication. The first user can enter the room by operating the input device 206 of the first user terminal 20, accessing the URL that provides the room service on the server 10 using a web browser, and entering the room identification information. Similarly, the second user can enter the room by operating the input device 306 of the second user terminal 30, accessing the URL that provides the room service on the server 10 using a web browser, and entering the room identification information. As a result, the first user and the second user can communicate with each other via the first user terminal 20 and the second user terminal 30, respectively, within a virtual dialogue space called a room, which is associated with the room identification information. By entering room identification information, one or more users, in addition to the first and second users, can enter a single room. This allows three or more users to communicate through their respective user terminals within a virtual dialogue space called a room, which is associated with the room identification information.
[0084] <Video Dialogue> System 1 in this disclosure may provide an online conversation service (video conversation service) that includes video data. For example, the control unit 204 of the first user terminal 20 and the control unit 304 of the second user terminal 30 transmit video data captured by the camera 2061 of the first user terminal 20 and the camera 3061 of the second user terminal 30 to the server 10. Based on the received video data, server 10 transmits video data captured by the camera 2061 of the first user terminal 20 to the second user terminal 30, and video data captured by the camera 3061 of the second user terminal 30 to the first user terminal 20. The control unit 204 of the first user terminal 20 displays the video data captured by the camera 3061 of the second user terminal 30 on the display 2081. The control unit 304 of the second user terminal 30 displays the video data captured by the camera 2061 of the first user terminal 20 on the display 3081. Server 10 may transmit video data of some or all of the users participating in the online dialogue to the first user terminal 20 and the second user terminal 30. In this case, the control unit 204 of the first user terminal 20 displays the received video data of some or all of the users participating in the online dialogue on the display 2081 of the first user terminal 20, arranged on a single screen. This allows the user to check the dialogue status of multiple users participating in the online dialogue. The second user terminal 30 may perform a similar process.
[0085] <Dialogue Memory Processing> Dialogue memory processing is the process of storing data related to conversations that take place between users and customers.
[0086] <Overview of Dialogue Memory Processing> The dialogue memory processing is a series of processes that store data related to a dialogue in the dialogue table 1014 when a dialogue is initiated between a user and a customer.
[0087] <Details of Dialogue Memory Processing> When a conversation between a user and a customer begins, the voice server (PBX) 60 records audio data related to the conversation between the user and the customer and sends it to the server 10. When the control unit 104 of the server 10 receives the audio data, it creates a new record in the conversation table 1014 and stores the data related to the conversation between the user and the customer. Specifically, the control unit 104 of the server 10 stores the user ID, customer ID, conversation category, call type, and the content of the audio data in the new record in the conversation table 1014.
[0088] The control unit 104 of the server 10 obtains the first user ID 2011 of the first user from the first user terminal 20 during outgoing or incoming call processing, and stores it in the user ID field of a new record in the dialogue table 1014. The control unit 104 of server 10 queries the CRM system 50 based on the telephone number during outgoing or incoming call processing. The CRM system 50 retrieves the customer ID by searching the customer table 5012 using the telephone number and sends it to server 10. The control unit 104 of server 10 stores the retrieved customer ID in the customer ID field of a new record in the dialogue table 1014. The control unit 104 of the server 10 stores the dialogue category value, which has been set in advance for each user or customer, in the dialogue category field of a new record in the dialogue table 1014. Alternatively, the dialogue category may be stored by the user selecting and entering a value for each dialogue. The control unit 104 of the server 10 identifies whether the ongoing conversation was initiated by the user or the customer, and stores either an outbound (initiated by the user) or inbound (initiated by the customer) value in the "Initiated / Initiated Type" field of the new record in the conversation table 1014.
[0089] The control unit 104 of server 10 stores the voice data received from the voice server (PBX) 60 in the voice data field of a new record in the dialogue table 1014. Alternatively, the voice data may be stored as a voice data file in another location, and reference information (path) to the voice data file may be stored after the dialogue ends. Furthermore, the control unit 104 of server 10 may be configured to store the voice data after the dialogue ends.
[0090] Furthermore, in the video dialogue service, the control unit 104 of the server 10 stores the video data received from the first user terminal 20 and the second user terminal 30 in the video data field of a new record in the dialogue table 1014. Alternatively, the video data may be stored as a video data file in another location, and reference information (path) to the video data file may be stored after the dialogue ends. Alternatively, the control unit 104 of the server 10 may be configured to store the video data after the dialogue ends.
[0091] <emotional analysis processing> Emotion analysis processing involves analyzing dialogue information such as audio and video from online conversations conducted by multiple users, identifying the emotional states of the users participating in the conversation, identifying label information based on those emotional states, and storing it in association with the dialogue information.
[0092] <Overview of emotion analysis processing> The sentiment analysis process involves detecting online conversations between users, storing conversation information, dividing the audio and video data contained in the conversation information into segment data such as segment audio data and segment video data for each utterance segment, calculating sentiment features for each segment data, identifying label information based on the sentiment features, and storing the label information in association with the conversation information.
[0093] <Details of the sentiment analysis process> The details of the emotion analysis process are described below.
[0094] In step S101, an online conversation between the user and the customer is initiated via the outgoing call processing, incoming call processing, room, etc., as previously described.
[0095] In step S102, the emotion analysis unit 1042 of the server 10 performs a reception step to receive voice data related to the dialogue. Specifically, through dialogue memory processing, the first user terminal 20 transmits the first user ID 2011, audio data collected from the microphone 2062, and video data captured by the camera 2061 to the server 10. The control unit 104 of the server 10 stores the received first user ID 2011, audio data, and video data in the user ID, audio data, and video data fields of a new record in the dialogue table 1014, respectively. Similarly, the second user terminal 30 transmits the second user ID 3011, audio data collected from the microphone 3062, and video data captured by the camera 3061 to the server 10. The control unit 104 of the server 10 stores the received second user ID 3011, audio data, and video data in the user ID, audio data, and video data fields of a new record in the dialogue table 1014, respectively. Accordingly, a new dialogue ID is assigned and stored in the dialogue ID field of the new record in dialogue table 1014.
[0096] In step S103, the emotion analysis unit 1042 of the server 10 executes a speech extraction step in which it extracts multiple segment speech data for each speech segment from the speech data received in the reception step. Specifically, in step S102, the emotion analysis unit 1042 of server 10 acquires (receives) the dialogue ID, audio data, and video data stored in the dialogue table 1014. From the acquired (received) audio data and video data, the emotion analysis unit 1042 of server 10 detects the sections in which audio exists (speech sections), and extracts the audio data and video data for each speech section as section audio data and section video data, respectively. The section audio data and section video data are associated with the speaker's user ID, the start date and time of the speech section, and the end date and time of the speech section for each speech section. The emotion analysis unit 1042 of server 10 performs text recognition on the extracted segment audio data and segment video data to convert the segment audio data and segment video data into segment spoken text, which is written as text. The specific method of text recognition is not particularly limited. For example, it may be converted using signal processing technology, machine learning or deep learning using AI (artificial intelligence), etc.
[0097] The emotion analysis unit 1042 of server 10 stores the dialogue ID to be processed, the speaker's user ID (first user ID 2011 or second user ID 3011), start date and time, end date and time, section audio data, section video data, and section spoken text in the fields for dialogue ID, speaker ID, start date and time, end date and time, section audio data, section video data, and section spoken text, respectively, in the new record of the audio section table 1016.
[0098] The audio segment table 1016 stores the segment text for each utterance segment of the audio data, associated with the start date and time and the speaker, as continuous time-series data. By checking the segment text stored in the audio segment table 1016, the user can confirm the content of the dialogue as text information without having to check the content of the audio data.
[0099] Furthermore, during the text recognition process, it is also possible to configure the system to remove information that is meaningless for understanding the conversation between the user and the customer, such as fillers included in the text, and store the speech recognition information in the speech interval table 1016.
[0100] In step S104, the emotion analysis unit 1042 of the server 10 executes an emotion calculation step, which calculates multiple emotion features related to the speaker's emotional state in each of the multiple segment audio data extracted in the speech extraction step. The emotion calculation step calculates emotion features as output data by applying the segment audio data extracted in the speech extraction step as input data to a learning model. Specifically, in S103, the emotion analysis unit 1042 of server 10 acquires the section audio data, section video data, and section spoken text stored in the audio section table 1016, and applies them as input data to the emotion evaluation model 1031. The emotion evaluation model 1031 outputs emotion features corresponding to the input data as output data.
[0101] In step S104, the emotion calculation step performs the step of calculating an emotion vector that indicates the intensity of multidimensional emotions, corresponding to each of the multiple interval audio data extracted in the audio extraction step. Specifically, in S103, the emotion analysis unit 1042 of server 10 acquires the section audio data, section video data, and section spoken text stored in the audio section table 1016, and applies them as input data to the emotion evaluation model 1031. The emotion evaluation model 1031 outputs the intensity of multiple emotional states (dimensions) corresponding to the input data, and an emotion vector that is quantitatively expressed as a numerical value, as output data.
[0102] The emotion calculation step performs the step of calculating an emotion scalar that represents the intensity of one-dimensional emotion, corresponding to each of the multiple interval audio data extracted in the audio extraction step, based on the calculated emotion vector. The emotion analysis unit 1042 of server 10 calculates an emotion scalar, which represents the intensity of a one-dimensional emotion, by applying principal component analysis, a learning model such as a deep learning model, and calculations on each component of the emotion vector to the emotion vector. For example, the emotion scalar is an index that quantitatively expresses the degree of positivity or negativity of the speaker's emotional state in the speech interval information, and may be numerical data normalized to a range of values from +1 (positive) to -1 (negative).
[0103] The emotion analysis unit 1042 of server 10 stores the calculated emotion features, namely the emotion vector and emotion scalar, in the emotion data field of the record to be analyzed in the speech interval table 1016. The emotion data field may be configured to store either the emotion vector or the emotion scalar.
[0104] In step S104, the emotion analysis unit 1042 of the server 10 searches for the user ID in the user table 1012 based on the speaker ID of the record to be analyzed in the speech interval table 1016, and obtains the user attribute.
[0105] In step S105, the emotion analysis unit 1042 of the server 10 performs a label identification step to identify label information for the dialogue based on the multiple emotion features calculated in the emotion calculation step. Specifically, the emotion analysis unit 1042 of server 10 searches for the dialogue ID in the voice section table 1016 based on the dialogue ID and obtains the emotion data items. Based on the emotion data, the emotion analysis unit 1042 of server 10 searches for the existence of records that match the emotion conditions in the emotion condition master 1021 and obtains the label data items of the corresponding records. In this disclosure, the emotion analysis unit 1042 of the server 10 may be configured to calculate multiple emotion features corresponding to multiple stored emotion data for each of the multiple speech segment pieces extracted from a single conversation piece of information, and to identify and acquire label data as emotion conditions.
[0106] In step S105, the label identification step performs the step of identifying label information for the dialogue based on the multiple emotion scalars calculated in the emotion calculation step. Specifically, the emotion analysis unit 1042 of the server 10 may calculate for each of the multiple speech segment information extracted from a single conversational information, and use the emotion scalar included in the multiple stored emotion data as an emotion condition to identify label data.
[0107] In step S105, the label identification step performs the step of identifying label information for the dialogue based on the multiple emotion vectors calculated in the emotion calculation step. Specifically, the emotion analysis unit 1042 of server 10 may calculate for each of the multiple speech segment information extracted from a single conversational information, and use the emotion vector contained in the multiple stored emotion data as an emotion condition to identify label data. For example, the emotion condition may be defined by the range of each element component of the emotion vector.
[0108] In step S105, the label identification step performs the step of identifying label information for the dialogue based on the number of emotion features above or below a predetermined threshold among the multiple emotion features calculated in the emotion calculation step. Specifically, the emotion condition master 1021 stores information in the emotion condition field that includes a predetermined threshold and a predetermined number of items that are equal to or greater than the threshold. The emotion analysis unit 1042 of the server 10 compares the emotion scalar values corresponding to each of the multiple speech segment information extracted from a single conversation information with the predetermined threshold and counts the number of speech segment information (emotion scalars) that are equal to or greater than the predetermined threshold. It is also acceptable to count the number of items that are less than or equal to the predetermined threshold. The emotion analysis unit 1042 of server 10 determines that the emotion condition is met if the number of counted audio segment information items is greater than a predetermined number, and retrieves and identifies the label data item associated with the emotion condition in the emotion condition master 1021. For example, if the number of speech interval information (emotion scalars) above a predetermined threshold is greater than a predetermined number, label information indicating a positive emotional state in the dialogue is identified. Similarly, if the number of speech interval information (emotion scalars) below a predetermined threshold is greater than a predetermined number, label information indicating a negative emotional state in the dialogue is identified.
[0109] In step S105, the label identification step performs the step of identifying label information for the dialogue based on the proportion of emotion features that are above or below a predetermined threshold among the multiple emotion features calculated in the emotion calculation step. Specifically, the emotion condition master 1021 stores information about an emotion condition, including a predetermined threshold and a percentage (predetermined percentage) of values above or above that threshold. The emotion analysis unit 1042 of the server 10 compares the emotion scalar values corresponding to each of the multiple speech segment information extracted from a single conversational data set with the predetermined threshold and counts the number of speech segment information (emotion scalars) that are above or above the predetermined threshold. It is also acceptable to count the number of values below or below the predetermined threshold. The emotion analysis unit 1042 of server 10 determines that the emotion condition is met if the ratio of the number of counted speech segment information to the total number of speech segment information extracted for one dialogue information is greater than a predetermined ratio, and retrieves and identifies the label data item associated with the emotion condition in the emotion condition master 1021. For example, if the proportion of speech segment information (emotion scalar) above a predetermined threshold is greater than a predetermined proportion, label information indicating a positive emotional state in the dialogue is identified. Similarly, if the proportion of speech segment information (emotion scalar) below a predetermined threshold is greater than a predetermined proportion, label information indicating a negative emotional state in the dialogue is identified.
[0110] Alternatively, instead of using an emotion scalar, you may use a single element component included in the emotion vector, or an index calculated based on one or more elements included in the emotion vector, as the emotion feature and perform the same processing.
[0111] In step S105, the label identification step performs the step of identifying label information for the dialogue based on the statistical values of multiple sentiment features calculated in the sentiment calculation step. Specifically, assume that the emotion condition items in the emotion condition master 1021 store information for a predetermined threshold. The emotion analysis unit 1042 of the server 10 calculates statistical values such as the mean, median, mode, etc., of the emotion scalar values corresponding to each of the multiple speech segment information extracted from one conversation information, compares them with the predetermined threshold, and determines that the emotion condition applies if the value is greater than or equal to the predetermined threshold, and retrieves and identifies the label data item associated with the emotion condition in the emotion condition master 1021. Note that the condition may also be set to be less than or equal to the predetermined threshold.
[0112] In step S105, the label identification step performs the step of identifying label information for the dialogue based on the time-series changes of multiple sentiment features calculated in the sentiment calculation step. The label identification step includes the step of performing regression analysis on the time-series changes of multiple sentiment features calculated in the sentiment calculation step, and the step of identifying label information for the dialogue based on the regression coefficients obtained from the regression analysis. Specifically, assume that the emotional condition item in the emotional condition master 1021 stores the range of the regression coefficient. For each of the multiple audio segment pieces associated with the dialogue data, a regression analysis of Y=f(X) is performed, where the X-axis represents the start date and time, end date and time, or any date and time between the start and end dates and times of the audio segment piece, and the Y-axis represents the emotional scalar value included in the emotional data of that audio segment piece. Any regression analysis, such as linear regression or quadratic regression, may be applied. By performing the regression analysis, the regression coefficient is calculated, compared with the range of the regression coefficient, and if it falls within the range of the regression coefficient, it is determined that it corresponds to that emotional condition, and the label data item associated with the emotional condition in the emotional condition master 1021 is retrieved and identified. For example, in the case of linear regression (first-order regression), if the intercept is negative and the slope is positive, label information is identified that indicates an improvement in the emotional state during the conversation. Alternatively, instead of using an emotion scalar, you may use a single element component included in the emotion vector, or an index calculated based on one or more elements included in the emotion vector, as the emotion feature and perform the same processing.
[0113] In step S105, the emotion analysis unit 1042 of the server 10 performs the step of identifying a first emotion group, which is a set of multiple emotion features corresponding to multiple time-series consecutive interval audio data extracted in the audio extraction step. The emotion analysis unit 1042 of the server 10 then performs the step of identifying a second emotion group, which is a set of multiple emotion features corresponding to multiple time-series consecutive interval audio data extracted in the audio extraction step. Specifically, the emotion analysis unit 1042 of server 10 may divide the multiple speech segment information extracted from one conversational information into a group of segments consisting of multiple speech segment information, and then perform the label identification step described earlier for each group of segments. This identifies the label information corresponding to each of the multiple group of segments. For example, the emotion analysis unit 1042 of server 10 calculates an emotion scalar for each of the extracted speech segment information included in the segment group and stores it in emotion data. The emotion scalars included in the stored emotion data may be used as emotion conditions to identify label data. For example, the emotion analysis unit 1042 of server 10 calculates an emotion vector for each of the extracted audio segment information included in the segment group and stores it in emotion data. The emotion vectors included in the stored emotion data may be used as emotion conditions to identify label data.
[0114] In step S105, the label identification step includes identifying first label information for the dialogue based on multiple emotion features included in the first emotion group, and identifying second label information for the dialogue based on multiple emotion features included in the second emotion group. Specifically, the emotion analysis unit 1042 of server 10 divides the multiple speech segment information extracted from one conversational information into a group of segments, each consisting of multiple speech segment information, and performs the label identification step described above for each group of segments, thereby identifying the label information corresponding to each of the multiple group of segments.
[0115] In step S105, the emotion analysis unit 1042 of the server 10 performs a label presentation step in which it presents the first label information and the second label information to the first user. Specifically, the emotion analysis unit 1042 of the server 10 transmits the identified first label information and second label information to the first user terminal 20. The control unit 204 of the first user terminal 20 displays the received first label information and second label information on the display 2081 of the first user terminal 20 and presents it to the first user. The first label information and second label information may also be presented to any other user, such as the second user, other administrators, or other users.
[0116] In step S105, the emotion analysis unit 1042 of the server 10 executes a selection reception step in which it receives a selection instruction from the first user to select at least one of the first label information and the second label information presented in the label presentation step. Specifically, the first user selects either the first label information or the second label information displayed on the display 2081 of the first user terminal 20 by operating the input device 206 of the first user terminal 20. The first user may choose not to select either label information. The control unit 204 of the first user terminal 20 transmits the selected label information to the server 10. The emotion analysis unit 1042 of the server 10 identifies the received label information.
[0117] In step S105, the label identification step performs the step of identifying label information for the dialogue based on multiple emotion features calculated in the emotion calculation step and the user attributes of the first or second user who spoke the interval speech data corresponding to the multiple emotion features. Specifically, when the emotion analysis unit 1042 of the server 10 identifies label information, it may consider the user attributes of the first user and the second user identified in step S104 when identifying label information. For example, the user attributes of the first user and the second user may be included as conditions in the emotion conditions of the emotion condition master 1021.
[0118] In step S105, the label identification step performs the step of identifying label information for the dialogue based on multiple emotion features corresponding to the segment audio data of the second user's utterance, which were calculated in the emotion calculation step, without considering multiple emotion features corresponding to the segment audio data of the first user's utterance. Specifically, the emotion analysis unit 1042 of server 10 may exclude the voice segment information where the speaker ID is the first user ID 2011 from the multiple voice segment information extracted for the one dialogue information, and perform the label identification step described above based only on the voice segment information where the speaker ID is the second user ID 3011. This allows for the identification of label information that takes only the customer's emotional state into consideration. Typically, the first user, such as an operator, is generally more interested in the customer's emotional state than their own. By configuring it in this way, label information that specifically considers the customer's emotional state can be identified.
[0119] The emotion analysis unit 1042 of server 10 may exclude the speech segment information in which the speaker ID is the second user ID 3011 from the multiple speech segment information extracted for the 1 dialogue information, and perform the label identification step described above based only on the speech segment information in which the speaker ID is the first user ID 2011.
[0120] The emotion analysis unit 1042 of server 10 may perform the label identification step described above for each of the speech segment information where the speaker ID is the first user ID 2011 and the speech segment information where the speaker ID is the second user ID 3011, and identify multiple label information for the first label information and the second label information, respectively.
[0121] Furthermore, the emotion analysis unit 1042 of the server 10 may exclude the audio segment information from the multiple audio segment information extracted for the 1 dialogue information where the user identified by the speaker ID is the host user who initiated the dialogue, and perform the label identification step described above based only on the audio segment information where the user identified by the speaker ID is not the host user. This allows for the identification of label information without considering the emotional state of the dialogue facilitator. Typically, the dialogue facilitator is more interested in the emotional state of the dialogue partner than their own. By using this configuration, label information that takes the emotional state of the dialogue partner into account can be identified.
[0122] In step S106, the emotion analysis unit 1042 of the server 10 performs a storage step in which it stores the label information identified in the label identification step in association with the dialogue. Specifically, the emotion analysis unit 1042 of the server 10 stores the label information identified in step S105 in the label data items of the label table 1015, associating it with the dialogue ID assigned in step S101. In step S105, the identified label information may be presented to the first user, and the label information selected by the first user may be stored as label data in the label table 1015.
[0123] In step S106, the memory step performs the step of storing the first label information or the second label information identified in the label identification step in association with the dialogue. The memory step performs the step of storing at least one of the first label information and the second label information in association with the dialogue based on the selection instruction received from the first user in the selection acceptance step. Specifically, the system may be configured to store the label information selected by the first user as label data in the label table 1015.
[0124] Furthermore, the first user can display the label information stored in the label table 1015 from the server 10 on the display 2081 of the first user terminal 20 by operating the input device 206 of the first user terminal 20.
[0125] <Regarding the timing of sentiment analysis processing> Steps S103-S106 of the emotion analysis process may be configured to be executed after the completion of an online conversation involving multiple users. This allows label information corresponding to the users' emotional states during the conversation to be identified and stored in association with the conversation information after the online conversation has ended and the conversation content has been finalized.
[0126] Furthermore, the sentiment analysis process may be configured to run after the start of an online conversation involving multiple users and before the end of the conversation. In other words, the system can be configured to be executed at any time during an online conversation involving multiple users. Alternatively, steps S103 to S106 can be configured to be executed periodically in real time during the online conversation. This allows for the identification of label information corresponding to the user's emotional state in the previous conversation, even in the middle of the online conversation, and to be stored in association with the conversation information. This allows users to see the emotional state of other users participating in an online conversation in real time, and to organize and manage conversation information based on the latest emotional state.
[0127] <Impression Analysis Processing> The impression analysis process analyzes dialogue information such as audio and video from online conversations conducted by multiple users, identifies the impression state of each user participating in the conversation, and presents the impression state and speaker type to the user.
[0128] <Overview of Impression Analysis Processing> The impression analysis process involves detecting online conversations between users, storing conversation information, dividing the audio and video data contained in the conversation information into segment data such as segment audio data and segment video data for each utterance segment, calculating impression features for each segment data, identifying the speaker type based on the impression features, and presenting the identified speaker type to the user.
[0129] <Details of the impression analysis process> The details of the impression analysis process are explained below.
[0130] In step S301, an online conversation between the user and the customer is initiated via the outgoing call processing, incoming call processing, room, etc., as previously described.
[0131] In step S302, the impression analysis unit 1043 of the server 10 performs a dialogue acquisition step to acquire dialogue information related to the dialogue interaction between the second user and the first user. Step S302 is the same as step S102 in the emotion analysis process, so its explanation is omitted.
[0132] In step S303, the impression analysis unit 1043 of the server 10 performs a speech extraction step in which it extracts multiple segment speech data for each speech segment from the second user's speech data received in step S302. Step S303 is the same as step S103 in the emotion analysis process, so its explanation is omitted.
[0133] In step S304, the impression analysis unit 1043 of the server 10 performs an impression calculation step, which calculates impression features related to the impression the second user gives to other users in a dialogue, based on the dialogue information of the second user acquired in the dialogue acquisition step. The impression calculation step performs a step of calculating impression features that indicate the intensity of at least one impression from the following: like, dislike, noisy, hard to listen to, polite, hard to understand, timid, nervous, intimidating, violent, and sexual, based on the dialogue information acquired from the second user in the dialogue acquisition step. The impression calculation step takes the dialogue information obtained from the second user in the dialogue acquisition step as input data and applies it to a learning model to calculate impression features related to the impression the second user gives to other users in the dialogue, which are then output as data. Specifically, in S303, the impression analysis unit 1043 of server 10 acquires the section audio data, section video data, and section spoken text stored in the audio section table 1016. It then excludes the audio section information where the speaker ID is the first user ID 2011, and applies only the audio section information where the speaker ID is the second user ID 3011 as input data to the impression evaluation model 1032. The impression evaluation model 1032 outputs impression features corresponding to the input data as output data. This allows the impression given by the second user to be evaluated using impression features. Furthermore, the input data applied to the impression evaluation model 1032 may exclude the speech segment information where the speaker ID is the second user ID 3011, and instead use only the speech segment information where the speaker ID is the first user ID 2011. In this case, the impression given by the first user can be evaluated using impression features.
[0134] In step S304, the impression calculation step includes a step of calculating dialogue features related to the second user's way of speaking in the dialogue based on the dialogue information of the second user obtained in the dialogue acquisition step, and a step of calculating impression features based on the calculated dialogue features. The impression calculation step includes the steps of: applying the dialogue information of the second user obtained in the dialogue acquisition step as input data to the first learning model to calculate dialogue features related to the second user's way of speaking in the dialogue as output data; and applying the calculated dialogue features as input data to the second learning model to calculate impression features. The impression calculation step includes calculating dialogue features related to at least one of the following aspects of the second user's speaking style in the dialogue, based on the dialogue information of the second user obtained in the dialogue acquisition step: speaking speed, intonation, number of polite expressions, number of fillers, and number of grammatical utterances.
[0135] Specifically, in S303, the impression analysis unit 1043 of server 10 acquires the section audio data, section video data, and section spoken text stored in the audio section table 1016. It then excludes the audio section information where the speaker ID is the first user ID 2011, and applies only the audio section information where the speaker ID is the second user ID 3011 as input data to the first impression evaluation model 1033. The first impression evaluation model 1033 outputs dialogue features corresponding to the input data as output data. The impression analysis unit 1043 of server 10 applies dialogue features as input data to the second impression evaluation model 1034, and the second impression evaluation model 1034 outputs impression features corresponding to the input data as output data. This makes it possible to evaluate the impression given by the second user using impression features. Furthermore, the input data applied to the impression evaluation model 1032 may exclude the speech segment information where the speaker ID is the second user ID 3011, and instead use only the speech segment information where the speaker ID is the first user ID 2011. In this case, the impression given by the first user can be evaluated using impression features.
[0136] In step S304, the impression analysis unit 1043 of the server 10 performs a storage step in which it stores the impression features calculated in the impression calculation step in association with the second user. Specifically, the impression analysis unit 1043 of server 10 stores the calculated impression features in the impression data items of the records to be analyzed in the speech section table 1016. This allows the impression features to be stored in association with the second user via the speaker ID (second user ID) in the speech section table 1016. Alternatively, the impression features may be stored in association with the second user ID by providing a column for storing impression data (not shown) in the customer table 5012 of the CRM system 50. Furthermore, the impression features may also be stored in association with the second user ID by providing a column for storing impression data (not shown) in the user table 1012 of server 10. By storing the user impression characteristics identified in the relevant conversation in the customer table 5012 of the CRM system 50, the impression characteristics of the user can be shared with members of other departments within the company. For example, efficient work can be performed according to the impression of the conversation partner identified by the impression characteristics.
[0137] In step S305, the impression analysis unit 1043 of the server 10 performs an identification step to identify speaker types that label the impression a second user gives to other users, based on the impression features calculated in the impression calculation step. Specifically, the impression analysis unit 1043 of server 10 searches for the dialogue ID in the speech section table 1016 based on the dialogue ID and obtains the impression data items. Based on the impression data, the impression analysis unit 1043 of server 10 searches for the existence of records that match the impression conditions in the speaker type master 1022 and obtains the speaker type item of the corresponding record. In this disclosure, the impression analysis unit 1043 of the server 10 may be configured to calculate and acquire impression features for each of the multiple speech segment information extracted from one conversational information, and to identify the speaker type as impression conditions.
[0138] In step S305, the impression analysis unit 1043 of the server 10 performs a storage step in which it stores the speaker type identified in a specific step in association with a second user. Specifically, the impression analysis unit 1043 of server 10 transmits the identified speaker type and second user ID to the CRM system 50. The control unit 504 of the CRM system 50 stores the received speaker type and second user ID in the speaker type and user ID fields of the customer table 5012, respectively. In other words, it stores the identified speaker type in association with the user ID of the user who spoke in that conversation. By storing this information in the customer table 5012 of the CRM system 50, the speaker type of the user identified in the relevant conversation can be shared with members of other departments within the company. For example, this allows for more efficient customer service depending on the speaker type of the person being spoken to. In this disclosure, the user's speaker type is stored in the customer table 5012 of the CRM system 50, but it may also be stored in the user table 1012 of the server 10 in association with a second user.
[0139] In step S306, the impression analysis unit 1043 of the server 10 performs a presentation step in which it presents to the first user the impression features that were stored in association with the second user during the storage step. Specifically, the impression analysis unit 1043 of the server 10 transmits the impression features identified in step S305 to the first user terminal 20. The control unit 204 of the first user terminal 20 displays the received impression features on the display 2081 of the first user terminal 20 and presents them to the first user. The impression features may also be presented to any other user, such as the second user, other administrators, or other users.
[0140] In step S306, the impression analysis unit 1043 of the server 10 performs a presentation step in which it presents to the first user the impression features stored in association with the second user during the storage step, prior to the dialogue between the first user and the second user. For example, when a first user or another user initiates an online conversation with a second user via outgoing call processing, incoming call processing, a room, etc., the impression features of the second user, which were associated with the second user and stored in step S305, may be displayed on the outgoing call screen for making a call to the second user, the incoming call screen for receiving a call from the second user, the room screen before the conversation begins, etc., which are displayed on the display 2081 of the first user terminal 20, and presented to the first user. This allows the first user to prepare a response tailored to their impression of the second user before the conversation begins.
[0141] Furthermore, the impression analysis unit 1043 of the server 10 may perform a presentation step in which it presents to the first user the speaker type that was stored in association with the second user during the storage step, prior to the dialogue between the first user and the second user. For example, when a first user or another user initiates an online conversation with a second user via outgoing call processing, incoming call processing, a room, etc., the speaker type of the second user, which was associated with the second user and stored in step S305, may be displayed on the outgoing call screen for making a call to the second user, the incoming call screen for receiving a call from the second user, the room screen before the conversation begins, etc., which are displayed on the display 2081 of the first user terminal 20, and presented to the first user. This allows the first user to prepare responses tailored to the second user's speaker type before the conversation begins.
[0142] The impression analysis unit 1043 of the server 10 may perform a presentation step in which it presents to the first user the impression features that were stored in association with the second user during the storage step, before the end of the dialogue between the first user and the second user. For example, while the first user or another user is having an online conversation with the second user, the impression features of the second user, which were associated with the second user and stored in step S305, may be displayed on the conversation screen, room screen, etc., shown on the display 2081 of the first user terminal 20, and presented to the first user. The impression features may also be presented to any user, such as the second user, other administrators, or other users. This allows the first user to prepare responses during the conversation that are tailored to the second user's impression.
[0143] The impression analysis unit 1043 of the server 10 may perform a presentation step before the end of the dialogue between the first user and the second user, in which it presents to the first user the speaker type that was stored in association with the second user during the storage step. For example, while the first user or another user is having an online conversation with the second user, the speaker type of the second user, which was associated with the second user and stored in step S305, may be displayed on the conversation screen, room screen, etc., on the display 2081 of the first user terminal 20 and presented to the first user. The impression features may also be presented to any user, such as the second user, other administrators, or other users. This allows the first user to prepare responses during the conversation that are tailored to the second user's speaker type.
[0144] The impression analysis unit 1043 of the server 10 may perform a presentation step in the impression calculation step in which it presents one or more of the dialogue features that have a large influence on the impression features from among the multiple dialogue features. Specifically, the impression analysis unit 1043 of the server 10 may apply multiple dialogue features as input data to the second impression evaluation model 1034. When the second impression evaluation model 1034 outputs impression features corresponding to the input data as output data, it may identify one or more dialogue features that have a significant impact on the output impression features, and transmit them to the first user terminal 20, the second user terminal 30, and other user terminals, etc., for presentation to the user. For example, the second impression evaluation model 1034 may output one or more dialogue features as output data that have a significant impact on the output impression features. This allows for the rapid acquisition of dialogue features that have a significant impact on the impression features.
[0145] <Variations of impression analysis processing> The impression analysis process may also be configured to identify the impression state of the first user, who is the operator, rather than the second user, who is the customer. Furthermore, the system may accept target impression features and target speaker types that the first user wants to convey to other users, calculate dialogue features that the first user should improve, and present them to the first user. In other words, it may include a step to suggest a preferred way of speaking to the first user. In this case, the processing content is the same in steps S301 to S305 of the impression analysis process, except that the second user is read as the first user, so the explanation is omitted.
[0146] In step S306, the impression analysis unit 1043 of the server 10 performs a target reception step in which it receives a target speaker type that the first user should set as a target for other users in the dialogue. Specifically, the first user accesses a predetermined web page provided by the server 10 by operating the input device 206 of the first user terminal 20, and selects a target speaker type (target speaker type) from a list of speaker types. The control unit 204 of the first user terminal 20 identifies the selected target speaker type and transmits it to the server 10. The server 10 receives and accepts the target speaker type. The target speaker type is a speaker type related to the impression state that the first user desires to convey to other users, and the first user may select it themselves, or the first user's administrator or other person may select it according to the first user's duties, etc.
[0147] In step S306, the impression analysis unit 1043 of the server 10 performs a target reception step in which it receives target impression features that the first user should give to other users in the dialogue. Specifically, the impression analysis unit 1043 of server 10 searches the speaker type item in the speaker type master 1022 based on the received target speaker type and obtains impression conditions. Based on the obtained impression conditions, the impression analysis unit 1043 of server 10 identifies the impression features included within the range of those impression conditions as target impression features and accepts them. The impression analysis unit 1043 of server 10 may also be configured to obtain and accept target impression features output by applying the target speaker type as input data to a learning model (not shown), etc. Alternatively, the target impression features may be received from the first user via the input device 206 of the first user terminal 20.
[0148] In step S306, the impression analysis unit 1043 of the server 10 performs an improvement step in which it calculates the dialogue features that the first user should improve, based on the impression features calculated in the impression calculation step and the target impression features received in the target acceptance step. Specifically, the impression analysis unit 1043 of server 10 identifies and accepts the dialogue features necessary to obtain the identified target impression features as target dialogue features, based on the identified target impression features. The impression analysis unit 1043 of server 10 may also be configured to acquire and accept the target dialogue features by applying the target impression features as input data to a learning model (not shown), etc. Examples of dialogue features that the first user should improve include "speak faster," "speak slower," "use more intonation," and "use less intonation." These dialogue features that the first user should improve can also be used as target dialogue features (target dialogue features).
[0149] The impression analysis unit 1043 of server 10 compares the dialogue features calculated in step S304 with the target dialogue features. The impression analysis unit 1043 of server 10 calculates the difference between the dialogue features and the target dialogue features as the dialogue features that the first user should improve. The impression analysis unit 1043 of server 10 also compares the dialogue features with the target dialogue features and identifies dialogue features with a large degree of deviation as the dialogue features that the first user should improve. The impression analysis unit 1043 of the server 10 transmits the dialogue features that the first user should improve to the first user terminal 20. The control unit 204 of the first user terminal 20 displays the received dialogue features that should be improved on the display 2081 of the first user terminal 20 and presents them to the first user. For example, among the dialogue features of the first user in a conversation, such as speaking speed, intonation, number of polite expressions, number of fillers, and number of grammatical utterances, the system identifies the dialogue features that the first user should improve and suggests to the first user to what extent they should improve their speaking speed, intonation, number of polite expressions, number of fillers, etc. This allows operators and others to improve the impression they give to others by specifically improving their speaking style. Furthermore, the dialogue features may be presented to a second user and other users.
[0150] As a result, the impression analysis unit 1043 of the server 10 can execute an improvement step that calculates the dialogue features that the first user should improve, based on the speaker type calculated in the impression calculation step and the target speaker type received in the target reception step. In other words, users can understand the dialogue features that need improvement according to the target speaker type they receive, and by improving their speaking style based on these features, they can bring the impression they give to others closer to that of the target speaker type.
[0151] <Topic definition processing> The topic definition process is the process by which a user registers and stores topics related to a given subject, associated with multiple keywords.
[0152] <Overview of Topic Definition Processing> Users can define and remember new topics based on keywords such as multiple words, nouns, and adjectives. Furthermore, for topics already remembered, the system expands the range of keywords associated with a topic by receiving suggestions of keywords highly relevant to that topic based on previously remembered conversational information, adding these keywords to the existing keyword list, and remembering them.
[0153] <Details of topic definition processing> The following describes the details of the topic definition process.
[0154] The topic processing unit 1044 of the server 10 executes a keyword presentation step in which it presents to the first user one or more new keywords to newly associate with the first topic, based on the voice data stored in the voice storage step and the multiple keywords received in the keyword reception step. Specifically, the first user executes the application program 2012 and runs a browser application by operating the input device 206 of the first user terminal 20. In the browser application, the first user sends a request to the server 10 to request a page for defining a topic by entering a predetermined URL (Uniform Resource Locator) that specifies a predetermined web server provided by the server 10.
[0155] The topic processing unit 1044 of server 10 searches the speaker ID entry in the speech segment table 1016 based on the first user ID 2011 included in the received request and retrieves the segment reading text. The topic processing unit 1044 of server 10 extracts strings such as nouns, adjectives, and keywords contained in the section-reading text by performing morphological analysis and other processing on the section-reading text. At this time, the importance of the strings may be calculated based on the frequency of occurrence of strings for each dialogue information and audio section information. A method for calculating importance may be tf-idf. The topic processing unit 1044 of server 10 identifies a predetermined number of strings with high importance as keyword candidates.
[0156] The topic processing unit 1044 of server 10 may obtain topic IDs and keywords from the topic master 1023, and identify as keyword candidates strings that are co-occurring with multiple keywords associated with each of multiple topic IDs in one or more dialogue information or voice segment information, but are not associated with topic IDs. In calculating the co-occurrence relationship, the importance of each keyword and string may be considered. In identifying keyword candidates, a predetermined number of strings may be identified as keyword candidates, taking into account the importance calculated based on the frequency of occurrence, etc.
[0157] The topic processing unit 1044 of server 10 sends the identified keyword candidates to the first user terminal 20. The control unit 204 of the first user terminal 20 displays the received keyword candidates on the display 2081 of the first user terminal 20 and presents them to the first user.
[0158] The topic processing unit 1044 of server 10 executes a keyword reception step that accepts one or more keywords from the first user. Specifically, the first user selects a keyword to associate with a new topic from the keyword candidates displayed on the display 2081 of the first user terminal 20 by operating the input device 206 of the first user terminal 20. The control unit 204 of the first user terminal 20 sends one or more keyword candidates selected by the first user to the server 10.
[0159] The keyword reception step performs the step of receiving one or more keywords selected by the first user from among the multiple new keywords presented to the first user in the keyword presentation step. Specifically, the topic processing unit 1044 of server 10 receives and accepts one or more keyword candidates from the first user terminal 20.
[0160] The topic processing unit 1044 of the server 10 executes a topic storage step in which it stores one or more keywords received in the keyword reception step, associating them with a first topic relating to a predetermined topic. Specifically, the topic processing unit 1044 of server 10 stores the received keyword candidates in the topic master 1023, associating them with topic IDs. Note that one or more keyword candidates selected by the first user may be associated with topic IDs already stored in the topic master 1023, or a new topic ID may be generated and associated with the newly generated topic ID. If the topic ID is to be associated with a topic ID already stored in the topic master 1023, the first user performs a selection operation to select the topic ID to be associated by operating the input device 206 of the first user terminal 20.
[0161] <Topic analysis processing> Topic analysis processing involves analyzing dialogue information such as audio and video from online conversations conducted by multiple users, calculating the degree of relevance between the dialogue information and one or more topics, and then associating and storing the topics with the dialogue information based on the degree of relevance.
[0162] <Overview of Topic Analysis Processing> The topic analysis process involves detecting online conversations between users, storing conversation information, dividing the audio and video data contained in the conversation information into segment data such as segment audio data and segment video data for each utterance segment, calculating the degree of relevance of each segment data to multiple topics, identifying the topic for each segment data, and storing representative topics as label information for the conversation information.
[0163] <Details of topic analysis processing> The following describes the details of the topic analysis process.
[0164] In step S511, an online conversation between the user and the customer is initiated via the outgoing call processing, incoming call processing, room, etc., as previously described.
[0165] In step S512, the topic processing unit 1044 of the server 10 performs a reception step to receive audio data related to the dialogue. The topic processing unit 1044 of the server 10 performs an audio storage step to store the audio data received in the reception step. Step S512 is the same as step S102 in the emotion analysis process, so its explanation is omitted.
[0166] In step S513, the topic processing unit 1044 of the server 10 executes a speech extraction step in which it extracts multiple segment speech data for each speech segment from the speech data received in the reception step. Step S513 is the same as step S103 in the emotion analysis process, so its explanation is omitted.
[0167] In step S513, the speech extraction step may, before the dialogue ends, perform a step of extracting multiple segment speech data for each utterance segment from the speech data received in the reception step. In other words, the voice extraction step can be configured to be executed at any point during an online conversation involving multiple users.
[0168] In step S514, the topic processing unit 1044 of the server 10 performs a topic identification step to identify a first topic relating to a predetermined topic, which is associated with multiple keywords. Specifically, the topic processing unit 1044 of server 10 refers to the topic master 1023 to obtain and identify the topic ID and one or more keywords associated with the topic ID, which have been pre-registered through the topic definition process.
[0169] The relevance calculation step performs a step to calculate the degree of relevance for each of the multiple topics identified in the topic identification step, for each of the multiple audio data segments. In this disclosure, for the sake of simplicity, we will mainly describe a first topic and one or more keywords associated with the first topic, but the topic is not limited to one, and similar processing may be performed on multiple topics (second topic, third topic, etc.).
[0170] In step S514, the topic processing unit 1044 of the server 10 performs a relevance calculation step in which it calculates a first relevance for each of the multiple interval audio data, which indicates the degree of relevance with the first topic identified in the topic identification step. Specifically, the topic processing unit 1044 of server 10 calculates a first relevance score, which indicates the degree of relevance to the first topic, based on the relationship between the audio segment information acquired in S513 and the keywords associated with the first topic.
[0171] An example of how to calculate the first relevance is described below. The topic processing unit 1044 of server 10 creates a high-dimensional vector (topic vector) as a distributed representation (embedded representation) based on keywords associated with the first topic. The topic processing unit 1044 of server 10 also performs morphological analysis and other processing on the segment reading text contained in multiple speech segment information to extract strings such as nouns, adjectives, and keywords contained in the segment reading text, and creates a high-dimensional vector (speech segment vector) as a distributed representation based on the extracted strings. A method called Word2vec is known for creating distributed representations. The topic processing unit 1044 of server 10 calculates the first relevance by calculating the cosine similarity between the topic vector and the speech segment vector. Note that the first relevance can also be calculated using algorithms that calculate the distance between arbitrary multidimensional vectors, such as the Euclidean distance, Mahalanobis distance, Manhattan distance, Chebyshev distance, and Minkowski distance. The first relevance calculated in this way reflects the overall similarity trend between multiple keywords associated with the first topic and the strings contained in multiple audio segment information. As a result, the strings contained in the audio segment information are not treated as different words due to paraphrasing or differences in notation of keywords contained in the topic, and a higher relevance is obtained for audio segment information that has a high semantic relevance to the keywords contained in the first topic. In this disclosure, we have explained how to calculate the first relevance score, which indicates the degree of relevance with the first topic. The calculation of the relevance score between any topic and the audio segment information is similar.
[0172] The relevance calculation step may include, before the dialogue ends, a step of calculating a first relevance for each segment of audio data included in multiple segment audio data, which indicates the degree of relevance to the first topic identified in the topic identification step. In other words, it can be configured to run at any point during an online conversation involving multiple users. This allows for the calculation of the degree of relevance between each topic and the audio segment information from the previous conversation, even in the middle of an online conversation.
[0173] In the relevance calculation step, keywords that are frequently included in the multiple segment audio data extracted in the audio extraction step are given less weight in the relevance calculation, and the degree of agreement, which takes into account the weighting of the multiple keywords associated with the first topic for each of the multiple segment audio data, may be calculated as the first relevance, which indicates the degree of relevance to the first topic. Specifically, when calculating relevance, different weights may be assigned to the importance of each of the multiple keywords associated with the first topic. For example, for multiple audio segment information extracted from one dialogue piece, the importance and weight of keywords that frequently appear in many audio segment pieces may be set to smaller values compared to other keywords, so that their influence on relevance is reduced. This prevents the relevance of topics associated with common keywords that frequently appear in many audio segment pieces from being overestimated. In this disclosure, we have described the calculation of the first relevance score, which indicates the degree of relevance with the first topic. However, the calculation of the relevance score between any topic and the audio segment information may be performed in the same manner.
[0174] The relevance calculation step may involve assigning less weight to keywords that are frequently found in multiple audio data segments up to a predetermined number of time-series entries prior to the target audio data segment for which the first relevance is to be calculated, among the multiple keywords associated with the first topic. The degree of agreement, which takes into account the weighting of multiple keywords associated with the first topic for each of the multiple audio data segments, may then be calculated as the first relevance, indicating the degree of relevance with the first topic. For example, instead of using all of the multiple audio segment information extracted from a single dialogue data point, the importance and weight of keywords that frequently appear in many audio segment pieces can be set to smaller values compared to other keywords, so that their influence on the relevance is reduced, for multiple audio segment pieces up to a predetermined number of points prior in time from the target audio segment information being calculated. This allows for a more accurate calculation of the relevance between the most recent audio segment information and the topic, even at any point during the dialogue before the dialogue ends. In this disclosure, we have described the calculation of the first relevance score, which indicates the degree of relevance with the first topic. However, the calculation of the relevance score between any topic and the audio segment information may be performed in the same manner.
[0175] The topic processing unit 1044 of server 10 stores the relevance calculated for each of the multiple topics for each of the multiple speech segment pieces extracted for a single conversation piece, in the section ID that identifies the speech segment piece, the topic ID that identifies the topic, and the calculated relevance in the section ID, topic ID, and relevance fields of a new record in the topic relevance table 1017, respectively.
[0176] In step S515, among one or more topics with a relevance of a predetermined value or higher in each audio segment, the topic with the highest relevance is identified as the topic relating to the predetermined topic mentioned in the audio segment information. Note that the topic does not necessarily need to be identified. The topic processing unit 1044 of the server 10 stores the topic ID of the identified topic in the audio segment table 1016 in the topic ID field of the record identified by the segment ID of the audio segment information subject to relevance calculation. As a result, the audio segment information is stored in association with the topic with the highest relevance.
[0177] In step S516, the topic processing unit 1044 of the server 10 executes a label identification step to identify label information for the dialogue based on the relevance of each of the multiple topics calculated in the relevance calculation step. The topic processing unit 1044 of the server 10 executes a storage step to store the label information identified in the label identification step in association with the dialogue. Specifically, in step S515, the topic processing unit 1044 of server 10 aggregates the topic IDs stored for each of the multiple audio segment information extracted for one dialogue information, and identifies one or more topic IDs in descending order of the aggregated topic IDs as the topics that characterize the one dialogue information. Alternatively, one or more topic IDs with a predetermined number or more aggregated topic IDs may be identified as the topics that characterize the one dialogue information. The topic processing unit 1044 of server 10 identifies the topic name, label, and other information of the identified topic ID as label information. Alternatively, the system may be configured to identify arbitrary label information based on the identified topic ID by referring to a table (not shown). The identified label information and the dialogue ID of the relevant dialogue information are stored in the label data and dialogue ID fields of the new record in label table 1015. This associates the dialogue information with the topics that characterize it as label information, making it convenient to use when searching for dialogue information.
[0178] <Regarding the timing of topic analysis processing> Steps S513 to S516 of the topic analysis process may be configured to be executed after the online dialogue between multiple users has ended. This allows topics related to the dialogue to be identified, associated with the dialogue information, and stored after the online dialogue has ended and the dialogue content has been finalized.
[0179] Alternatively, the topic analysis process may be configured to run after the start of an online conversation involving multiple users and before the end of the conversation. In other words, it can be configured to be executed at any time during an online conversation involving multiple users. Alternatively, steps S513 to S516 can be configured to be executed periodically in real time during the online conversation. This allows for a configuration where topics corresponding to previous conversations are identified and stored in association with the conversation information, even in the middle of an online conversation. This allows users to see in real time what topics other users are mentioning during online conversations, and to organize and manage conversation information based on the latest topics.
[0180] <Topic presentation process> The topic presentation process visually visualizes and presents dialogue information, such as audio and video, from online conversations conducted by multiple users, and also presents topics associated with that dialogue information. Users can see the dialogue information and related topics at a glance, allowing them to intuitively grasp the gist of the conversation.
[0181] <Overview of Topic Presentation Process> This process involves receiving a specification of the dialogue information to be presented from the user, acquiring the dialogue information, acquiring segment data and topics for each segment, analyzing the dialogue information, presenting the user with a speech graph that allows them to visually confirm the speech status of each speaker, and then overlaying the topics for each speech segment onto the speech graph and presenting them to the user.
[0182] <Details of Topic Presentation Process> The following describes the details of the topic presentation process.
[0183] In step S521, the first user selects the dialogue information for which they want to review the topic. Specifically, the first user executes the application program 2012 and runs the browser application by operating the input device 206 of the first user terminal 20. In the browser application, the first user sends a request to the server 10 to request a page to present a topic by entering a predetermined URL (Uniform Resource Locator) that specifies a predetermined web server provided by the server 10. The topic processing unit 1044 of server 10 searches the user ID entry in the dialogue table 1014 based on the first user ID 2011 included in the received request and obtains the dialogue ID. The topic processing unit 1044 of server 10 sends one or more obtained dialogue IDs to the first user terminal 20. The control unit 204 of the first user terminal 20 presents the received one or more dialogue IDs to the first user by displaying them on the display 2081 of the first user terminal 20. The first user selects a predetermined dialogue ID from the presented dialogue IDs by operating the input device 206 of the first user terminal 20. The control unit 204 of the first user terminal 20 transmits the selected predetermined dialogue ID to the server 10. The server 10 receives and accepts the dialogue ID.
[0184] Furthermore, if the first user is currently engaged in a conversation using the online conversation service related to this disclosure, the conversation information for that conversation may be selected. In other words, the configuration may be such that the topic presentation process is executed on the conversation screen displayed on the display 2081 of the first user terminal 20 during the conversation.
[0185] In step S522, the topic processing unit 1044 of the server 10 searches the dialogue ID item in the dialogue table 1014 based on the received dialogue ID and obtains dialogue information such as user ID, customer ID, dialogue category, send / receive type, audio data, and video data.
[0186] In step S523, the topic processing unit 1044 of server 10 searches the dialogue ID column in the voice interval table 1016 based on the received dialogue ID and obtains the interval ID, start date and time, end date and time, and topic ID. Based on the obtained interval ID, the topic processing unit 1044 of server 10 searches the interval ID column in the topic relevance table 1017 and obtains the topic ID and relevance. In other words, the topic processing unit 1044 of server 10 obtains multiple audio segment information associated with the dialogue ID, as well as the topic ID and relevance level for each audio segment information.
[0187] In step S524, the topic processing unit 1044 of the server 10 outputs a speech graph showing the time-series progression of the speaker's utterances based on the dialogue information acquired in step S522, and transmits it to the first user terminal 20. The control unit 204 of the first user terminal 20 displays the received speech graph on the display 2081 of the first user terminal 20 and presents it to the first user. Figure 20 shows an example screen 70 including the speech graph presented to the first user. Furthermore, the audio graph may be shown to any user, including the second user, other administrators, and other users.
[0188] The audio graph uses the horizontal axis to represent dialogue time, the vertical axis (upper) to represent the output volume of the first user's voice, and the vertical axis (lower) to represent the output volume of the second user's voice. The solid line L1 represents the first user's voice, and the dashed line L2 represents the second user's voice. Looking at the solid line L1 and the dashed line L2, we can see that, basically, while the first user is speaking, the second user is silent (listening), and while the second user is speaking, the first user is silent (listening). Here, the area indicated by Z3 is a state where both are speaking simultaneously (overlapping), and it is possible that the first user started speaking before the second user finished speaking. The areas indicated by Z1 and Z2 are periods of silence where neither is speaking. The areas indicated by P1 and P2 are the locations where the specified keywords appear.
[0189] In step S525, the topic processing unit 1044 of the server 10 performs a segment group identification step to identify a first segment group that includes one or more segment audio data from among multiple segment audio data whose first relevance, calculated in the relevance calculation step, is equal to or greater than a predetermined value. Specifically, in the topic analysis process, if the topic processing unit 1044 of server 10 determines that one or more audio segment pieces, each of which has a first relevance calculated to a predetermined value or higher for one of the multiple audio segment pieces extracted for one dialogue piece of information, refer to a topic related to the first topic, it identifies one or more audio segment pieces, including the one or more audio segment pieces, as the first segment group. For example, if the association of multiple time-series consecutive audio segment pieces with topics is as follows: segment 1: topic A, segment 2: topic A, segment 3: no topic, segment 4: topic A, segment 5: no topic, segment 6: topic B, segment 7: topic B, segment 8: topic B, then segments 1 to 4 are identified as the segment group related to topic A, and segments 6 to 8 are identified as the segment group related to topic B. Even if a segment of topic A, such as segment 3, contains an audio segment related to another topic, if segments 1 to 4 as a whole are considered to refer to the topic of topic A, then segments 1 to 4 may be identified together as the segment group related to topic A.
[0190] In this disclosure, the first group of intervals is identified, but it may also be used to identify one or more interval audio data from among multiple interval audio data that are related to a first topic concerning a predetermined topic. Alternatively, the first group of intervals may be identified by selecting one or more interval audio data and the first group of intervals through input operations by the first or second user.
[0191] In step S525, the topic processing unit 1044 of the server 10 executes a presentation step of presenting the first section group identified in the section group identification step in association with the first topic to the first user or the second user. The presentation step presents the first section group identified in the section group identification step on the same time series axis as the voice graph in a voice graph showing the time series transition of the speaker's utterance situation obtained by analyzing the voice data received in the reception step, and executes a step of presenting the first topic in association with the first section group to the first user or the second user. Specifically, in the voice graph of FIG. 20, the topic processing unit 1044 of the server 10 overlays and presents the first section group T1 associated with the first topic, the second section group T2 associated with the second topic, and the third section group T3 associated with the third topic on the voice graph as drawing objects. For example, the first section group T1, the second section group T2, and the third section group T3 may be configured to be drawn as drawing objects with different colors assigned to each topic. Thereby, the first user can visually recognize the section group in association with the related topic and overlaid on the voice graph. Thereby, the first user can visually confirm at a glance in the voice graph which part is about which topic. Note that the topic processing unit 1044 of the server 10 may be configured to present the first section group identified in the section group identification step to any user such as an administrator other than the first user and the second user, or other users.
[0192] In step S525, the section group identification step may include a step of calculating a moving average based on the first correlation degree calculated for each of a plurality of section voice data arranged in time series, and a step of identifying the section voice data whose calculated moving average is greater than or equal to a predetermined value as the first section group. Specifically, when identifying the group of intervals, the topic processing unit 1044 of the server 10 arranges the voice interval information obtained from the topic relevance table in chronological order based on the start date and time of the voice interval information, etc. The topic processing unit 1044 of the server 10 calculates the average of the relevance levels of the N closest relevance levels to the relevance level of a predetermined voice interval information as a moving average. N is an arbitrary integer. The calculated moving average is regarded as the new relevance level for the predetermined voice interval information, and the voice interval information with a relevance level equal to or higher than a predetermined value is identified as the first group of intervals associated with the first topic. In the present disclosure, mainly for simplicity, the moving average for the relevance level of one first topic has been described, but the topics are not limited to one, and the same processing may be executed for a plurality of topics. Thereby, even when the topic with a high relevance level switches in a short period for each utterance interval, by smoothing the relevance level of the topic, the group of intervals referring to the topic can be collectively identified. In an online dialogue service, it becomes easier for the user to confirm what topics the speaker has spoken about.
[0193] In step S525, the interval group identification step may execute a step of identifying, as the first group of intervals, a plurality of consecutive interval voice data whose calculated first relevance level is equal to or higher than a predetermined value among the plurality of interval voice data arranged in chronological order. Specifically, when identifying the group of intervals, the topic processing unit 1044 of the server 10 arranges the voice interval information obtained from the topic relevance table in chronological order based on the start date and time of the voice interval information, etc. The topic processing unit 1044 of the server 10 identifies a plurality of consecutive voice interval information with a relevance level equal to or higher than a predetermined value as the first group of intervals associated with the first topic. In the present disclosure, mainly for simplicity, the moving average for the relevance level of one first topic has been described, but the topics are not limited to one, and the same processing may be executed for a plurality of topics. This allows for the identification of consecutive, highly relevant audio segments related to a specific topic, grouped together as segments that mention that topic. In online conversation services, this makes it easier for users to see what topics the speaker was discussing.
[0194] In step S525, the topic processing unit 1044 of the server 10 executes a summarization step to generate a summary text that summarizes the text information contained in one or more section audio data, based on one or more section audio data and the first topic identified in the topic identification step. The summarization step executes a step to generate a summary text that summarizes the text information contained in one or more section audio data by extracting only the parts of the text information contained in one or more section audio data that are highly relevant to the first topic identified in the topic identification step.
[0195] In step S525, the summarization step performs the step of generating a summary text by applying text information contained in one or more segment audio data and multiple keywords associated with the first topic as input data to a learning model. Specifically, section data, which includes at least one of section audio data, section video data, and section text, along with multiple keywords associated with the topic of the section data, are applied to the summarization model 1035 as input data, and a summary text, which is text information that summarizes the text information contained in the section data, is obtained as output data. This makes it possible to extract only the parts of the text information contained in the section data that are particularly relevant to the topic, and to obtain a summary text that summarizes the text information contained in the section data.
[0196] In step S525, the summarization step performs the step of generating a summary text that summarizes the text information contained in one or more interval audio data, based on one or more interval audio data included in the first interval group identified in the interval group identification step and the first topic identified in the topic identification step. Specifically, one or more interval data points included in an interval group, along with multiple keywords associated with the topic of that interval group, are used as input data and applied to the summarization model 1035. The output data is a summary text, which is text information that summarizes the text information included in the interval group. This allows for the extraction of sections of the text information included in the interval data that are particularly relevant to the topic, and enables the acquisition of a summary text that summarizes the text information included in the interval data.
[0197] In step S525, the topic processing unit 1044 of the server 10 performs a presentation step in which it presents the summary text generated in the summarization step in association with one or more segment audio data. In step S525, the topic processing unit 1044 of the server 10 performs a presentation step in which it presents the summary text generated in the summarization step in association with the single interval group identified in the interval group identification step. Specifically, in the audio graph of Figure 20, the topic processing unit 1044 of the server 10 presents the summary text 701 relating to the first topic of the first interval group T1, in association with the first interval group T1. The topic processing unit 1044 of the server 10 may also present the summary text 701 in association with any one or more audio intervals, rather than an interval group. Furthermore, the topic processing unit 1044 of server 10 may be configured to present the first group of intervals identified in the interval group identification step to any user, such as the first user, the second user, other administrators, or other users.
[0198] <Learning Process> The learning processes for emotion evaluation model 1031, impression evaluation model 1032, first impression evaluation model 1033, and second impression evaluation model 1034 are described below.
[0199] <Training process of emotion evaluation model 1031> The training process for the emotion evaluation model 1031 involves training the training parameters of the deep neural network included in the emotion evaluation model 1031 using deep learning.
[0200] <Overview of the learning process for emotion evaluation model 1031> The training process for the emotion evaluation model 1031 involves using deep learning to train the training parameters of the deep neural network included in the emotion evaluation model 1031, with interval audio data, interval video data, and interval spoken text as input data (input vectors), so that emotion vectors or emotion scalars, which are emotion features, become the output data (training data). For the emotion evaluation model 1031, you may omit any of the following from the input data: section audio data, section video data, or section text reading.
[0201] <Details of the learning process for emotion evaluation model 1031> The learning unit 1051 of server 10 takes interval audio data, interval video data, interval text reading data, etc. as input data (input vectors) and creates training data so that predetermined emotion features become output data (training data). The learning unit 1051 of server 10 creates datasets such as training data, test data, and validation data for training the deep neural network of the emotion evaluation model 1031, based on the training data. The learning unit 1051 of server 10 trains the learning parameters of the deep neural network included in the emotion evaluation model 1031 using deep learning, based on the created dataset.
[0202] <Training process of impression evaluation model 1032> The learning process for impression evaluation model 1032 involves training the learning parameters of the deep neural network included in impression evaluation model 1032 using deep learning.
[0203] <Overview of the learning process for impression evaluation model 1032> The learning process for the impression evaluation model 1032 involves using interval audio data, interval video data, and interval spoken text as input data (input vectors) to train the learning parameters of the deep neural network included in the impression evaluation model 1032 using deep learning, so that impression features become output data (training data). You may omit any of the following from the input data of the impression evaluation model 1032: section audio data, section video data, or section read-aloud text.
[0204] <Details of the learning process for impression evaluation model 1032> The learning unit 1051 of server 10 takes interval audio data, interval video data, interval spoken text, etc. as input data (input vectors) and creates training data so that predetermined impression features become output data (training data). The learning unit 1051 of server 10 creates datasets such as training data, test data, and validation data for training the deep neural network of the impression evaluation model 1032, based on the training data. The learning unit 1051 of server 10 trains the learning parameters of the deep neural network included in the impression evaluation model 1032 using deep learning, based on the created dataset.
[0205] <Training process of the first impression evaluation model 1033> The training process for the first impression evaluation model 1033 involves training the training parameters of the deep neural network included in the first impression evaluation model 1033 using deep learning.
[0206] <Overview of the learning process for the first impression evaluation model 1033> The learning process of the first impression evaluation model 1033 is a process of learning the learning parameters of the deep neural network included in the first impression evaluation model 1033 through deep learning, with the segment voice data, segment video data, and segment read text as input data (input vectors) so that the dialogue feature amounts become output data (teacher data). Any one of the segment voice data, segment video data, and segment read text may be omitted from the input data of the first impression evaluation model 1033.
[0207] <Details of the learning process of the first impression evaluation model 1033> The learning unit 1051 of the server 10 creates learning data with the segment voice data, segment video data, and segment read text as input data (input vectors) so that a predetermined dialogue feature amount becomes output data (teacher data). Based on the learning data, the learning unit 1051 of the server 10 creates data sets such as training data, test data, and verification data for learning the deep neural network of the first impression evaluation model 1033. Based on the created data set, the learning unit 1051 of the server 10 learns the learning parameters of the deep neural network included in the first impression evaluation model 1033 through deep learning.
[0208] <Learning process of the second impression evaluation model 1034> The learning process of the second impression evaluation model 1034 is a process of learning the learning parameters of the deep neural network included in the second impression evaluation model 1034 through deep learning.
[0209] <Overview of the learning process of the second impression evaluation model 1034> The learning process of the second impression evaluation model 1034 is a process of learning the learning parameters of the deep neural network included in the second impression evaluation model 1034 through deep learning, with the dialogue feature amount as input data (input vector) so that the impression feature amount becomes output data (teacher data).
[0210] <Details of the learning process for the second impression evaluation model 1034> The learning unit 1051 of server 10 takes dialogue features and other data as input data (input vectors) and creates training data so that predetermined impression features become output data (training data). The learning unit 1051 of server 10 creates datasets such as training data, test data, and validation data for training the deep neural network of the second impression evaluation model 1034, based on the training data. The learning unit 1051 of server 10 trains the learning parameters of the deep neural network included in the second impression evaluation model 1034 using deep learning, based on the created dataset.
[0211] <Details of the training process for summary model 1035> The learning unit 1051 of the server 10 takes as input data (input vectors) section data, which includes at least one of section audio data, section video data, and section text, and a plurality of keywords associated with a topic related to a predetermined topic, and creates training data such that the output data (training data) is a summarized text, which is text information that summarizes the text information contained in the section data. The learning unit 1051 of server 10 creates datasets such as training data, test data, and validation data for training the deep neural network of the summarization model 1035, based on the training data. The learning unit 1051 of server 10 trains the learning parameters of the deep neural network included in the summarization model 1035 using deep learning, based on the created dataset.
[0212] <Basic Computer Hardware Configuration> Figure 21 is a block diagram showing the basic hardware configuration of computer 90. Computer 90 comprises at least a processor 901, main memory 902, auxiliary memory 903, and a communication interface IF991. These are electrically connected to each other by a communication bus 921.
[0213] The processor 901 is hardware for executing the instruction set written in a program. The processor 901 consists of an arithmetic unit, registers, peripheral circuits, etc.
[0214] Main memory 902 is used to temporarily store programs and data processed by programs, etc. For example, it is a volatile memory such as DRAM (Dynamic Random Access Memory).
[0215] Auxiliary storage device 903 refers to a storage device for saving data and programs. Examples include flash memory, HDD (Hard Disc Drive), magneto-optical disk, CD-ROM, DVD-ROM, and semiconductor memory.
[0216] The IF991 communication interface is an interface for inputting and outputting signals for communication with other computers via a network using wired or wireless communication standards. A network consists of various mobile communication systems built on the internet, LANs, wireless base stations, etc. For example, a network includes 3G, 4G, and 5G mobile communication systems, LTE (Long Term Evolution), and wireless networks that can connect to the internet via designated access points (e.g., Wi-Fi®). When connecting wirelessly, communication protocols include, for example, Z-Wave®, ZigBee®, and Bluetooth®. When connecting via a wired connection, the network also includes connections made directly via USB (Universal Serial Bus) cables, etc.
[0217] Furthermore, by distributing all or part of each hardware configuration across multiple computers 90 and connecting them to each other via a network, a computer 90 can be virtually realized. Thus, the concept of computer 90 includes not only a computer 90 housed in a single enclosure or case, but also a virtualized computer system.
[0218] <Basic Functional Configuration of Computer 90> The functional configuration of the computer realized by the basic hardware configuration of computer 90 (Figure 21) will be explained. The computer comprises at least one functional unit: a control unit, a memory unit, and a communication unit.
[0219] Furthermore, the functional units of computer 90 can also be realized by distributing all or part of each functional unit across multiple computers 90 interconnected via a network. The concept of computer 90 includes not only a single computer 90 but also a virtualized computer system.
[0220] The control unit is realized when the processor 901 reads various programs stored in the auxiliary storage device 903, loads them into the main memory device 902, and executes processing according to those programs. The control unit can realize various functional units that perform information processing depending on the type of program. In this way, the computer is realized as an information processing device that performs information processing.
[0221] The memory unit is implemented by the main memory 902 and the auxiliary memory 903. The memory unit stores data, various programs, and various databases. The processor 901 can also reserve memory areas corresponding to the memory unit in the main memory 902 or the auxiliary memory 903 according to the program. The control unit can also cause the processor 901 to perform operations such as adding, updating, and deleting data stored in the memory unit according to the various programs.
[0222] A database, specifically a relational database, is used to manage and link together tabular data sets called masters, which are structurally defined by rows and columns. In a database, tables are called tables, masters are called masters, the columns of tables are called columns, and the rows of tables are called records. In a relational database, relationships can be established and linked between tables and masters. Typically, each table and master has a primary key column to uniquely identify records, but setting a primary key column is not mandatory. The control unit can instruct the processor 901 to add, delete, or update records in specific tables and masters stored in the memory unit, according to various programs.
[0223] Furthermore, the databases and masters in this disclosure may include any data structures (lists, dictionaries, associative arrays, objects, etc.) in which information is structurally defined. Data structures also include data that can be considered as data structures by combining data with functions, classes, methods, etc., written in any programming language.
[0224] The communication unit is implemented by the communication IF991. The communication unit provides the functionality to communicate with other computers 90 via the network. The communication unit can receive information transmitted from other computers 90 and input it to the control unit. The control unit can cause the processor 901 to perform information processing on the received information according to various programs. The communication unit can also transmit information output from the control unit to other computers 90.
[0225] <Note> The details described in each of the above embodiments are noted below.
[0226] (Note 1) A program comprising a processor and a memory unit, which causes the computer to process information relating to a dialogue between a first user and a second user, the program causing the processor to execute: a reception step (S512) for receiving audio data relating to the dialogue; an audio extraction step (S513) for extracting multiple section audio data for each utterance section from the audio data received in the reception step; a section identification step (S525) for identifying one or more section audio data from among the multiple section audio data that are related to a first topic relating to a predetermined topic; and a summarization step (S525) for generating a summary text that summarizes the text information contained in one or more section audio data based on the one or more section audio data identified in the section identification step and the first topic. This allows users to see at a glance what topics the speakers discussed during the conversation.
[0227] (Note 2) The program as described in Appendix 1, wherein the summarization step (S525) is a step of generating a summary text by extracting portions of the text information contained in one or more audio segments that are highly relevant to the first topic. This allows users to see at a glance what topics the speakers discussed during the conversation.
[0228] (Note 3) The program described in Appendix 1 is a step in which the summarization step (S525) generates a summary text by applying text information contained in one or more interval audio data and multiple keywords associated with the first topic as input data to a learning model. This allows users to see at a glance what topics the speakers discussed during the conversation.
[0229] (Note 4) The program, as described in Appendix 1, causes the processor to perform a relevance calculation step (S514) which calculates a first relevance indicating the degree of relevance with a first topic for each of a plurality of interval audio data; an interval identification step (S525) which identifies a first interval group from among the plurality of interval audio data that includes one or more interval audio data for which the first relevance calculated in the relevance calculation step is equal to or greater than a predetermined value; and a summarization step (S525) which generates a summary text that summarizes the text information contained in one or more interval audio data based on one or more interval audio data included in the first interval group identified in the interval identification step and the first topic. This allows users to see at a glance what topics the speakers discussed during the conversation.
[0230] (Note 5) The program is the program described in Appendix 1, which causes the processor to perform a presentation step (S525) which presents the summarized text generated in the summarization step in association with one or more segment audio data. This allows users to see at a glance what topics the speakers discussed during the conversation.
[0231] (Note 6) The program is the program described in Appendix 4, which causes the processor to perform a presentation step (S525) which presents the summary text generated in the summarization step in association with the single group of intervals identified in the interval identification step. This allows users to see at a glance what topics the speakers discussed during the conversation.
[0232] (Note 7) The program is the program described in Appendix 4, which causes the processor to perform a presentation step (S525) in which it presents the first group of intervals identified in the interval identification step in relation to the first topic. This allows users to see at a glance what topics the speakers discussed during the conversation.
[0233] (Note 8) The presentation step (S525) is a step in the program described in Appendix 7, in which the first group of segments identified in the segment identification step is presented on the same time-series axis as the audio graph, which shows the time-series progression of the speaker's utterances obtained by analyzing the audio data received in the reception step, and the first topic is presented in relation to the first group of segments. This allows users to see at a glance what topics the speaker discussed during a conversation, by overlaying the audio graph, which shows the speaker's utterances in chronological order, with the audio graph.
[0234] (Note 9) The program is the program described in Appendix 1, which causes the processor to execute a keyword reception step (S502) for receiving one or more keywords from a first user, and a topic storage step (S503) for storing the one or more keywords received in the keyword reception step in association with a first topic relating to a predetermined subject. This allows users to see at a glance what topics the speaker discussed in a conversation, based on topics they have pre-associated with keywords and memorized themselves.
[0235] (Note 10) The program, as described in Appendix 9, causes the processor to perform a voice storage step (S512) for storing voice data received in the reception step, and a keyword presentation step (S501) for presenting one or more new keywords to the first topic to the first user based on the voice data stored in the voice storage step, and the keyword reception step (S502) is a step for receiving one or more keywords selected by the first user from among the multiple new keywords presented to the first user in the keyword presentation step. This allows the user to receive suggestions of one or more new keywords that are preferable to associate with a topic, based on keywords used in past conversations. The user can easily define and remember topics.
[0236] (Note 11) The program as described in Appendix 4, wherein the speech extraction step (S513) is a step in which multiple segment speech data are extracted for each utterance segment from the speech data received in the reception step before the dialogue ends, and the relevance calculation step (S514) is a step in which a first relevance indicating the degree of relevance with the first topic is calculated for each segment speech data included in the multiple segment speech data before the dialogue ends. This allows for the calculation of the degree of relevance between segment audio data and topics in real time during a conversation. For example, during a business negotiation, it can be used to see what topics the speakers are communicating about.
[0237] (Note 12) The relevance calculation step (S514) is a step in which the degree of relevance for each of the multiple topics associated with each of the multiple interval audio data is calculated, and the program is the program described in Appendix 4 which causes the processor to execute a memo identification step (S516) which identifies a response memo for a dialogue based on the degree of relevance for each of the multiple topics calculated in the relevance calculation step, and a storage step (S516) which stores the response memo identified in the memo identification step in association with the dialogue. This allows for the identification of topics that characterize the entire conversation and the management of conversational information by attaching response notes related to those topics to the conversation.
[0238] (Note 13) The relevance calculation step (S514) is a program described in Appendix 4, in which, among multiple keywords associated with the first topic, the weight given to the relevance is reduced for keywords that are frequently included in the multiple section audio data extracted in the audio extraction step, and the degree of agreement, taking into account the weighting of multiple keywords associated with the first topic for each of the multiple section audio data, is calculated as the first relevance, which indicates the degree of relevance with the first topic. This allows for a reduction in the weight of common keywords that appear in many audio segments, among the keywords associated with a topic. By increasing the importance of keywords that appear in specific audio segments, the degree of relevance between the audio segments and the topic can be calculated more accurately.
[0239] (Note 14) The program described in Appendix 13 calculates the degree of relevance to the first topic by considering the weighting of multiple keywords associated with the first topic, such that keywords that appear frequently in multiple audio data segments up to a predetermined number of time-series entries prior to the target audio data segment for which the first degree of relevance is to be calculated are given a smaller weight to the degree of relevance, and the degree of agreement considering the weighting of multiple keywords associated with the first topic for each of the multiple audio data segments is calculated as the first degree of relevance, which indicates the degree of relevance to the first topic. This allows for a more accurate calculation of the relevance between segment audio data and a topic, using less computation, by considering only multiple past segment audio data points near the target segment audio data among the keywords associated with the topic. Furthermore, the relevance to the topic can be calculated in real time.
[0240] (Note 15) The interval identification step (S525) is a program as described in Appendix 4, which includes the step of calculating a moving average based on the first correlation calculated for each of a plurality of interval audio data arranged in time series, and the step of identifying interval audio data whose calculated moving average is equal to or greater than a predetermined value as a first interval group. This allows for the identification of groups of segments that mention a particular topic, even when highly relevant topics switch rapidly between utterance segments, by smoothing the topic relevance. This makes it easier for users to understand what topics the speaker is discussing in a dialogue.
[0241] (Note 16) The program described in Appendix 4 is a step in which, from among multiple interval audio data arranged in chronological order, a group of consecutive interval audio data for which the calculated first relevance is equal to or greater than a predetermined value is identified as the first interval group. This allows for the identification of consecutive, highly relevant audio segments related to a specific topic, grouped together as segments that mention that topic. This makes it easier for users to understand what topics the speaker is discussing in a dialogue.
[0242] (Note 17) An information processing apparatus comprising a processor and a memory unit, wherein the processor executes a program described in any of the appendices 1 to 16. This allows users to see at a glance what topics the speakers discussed during the conversation.
[0243] (Note 18) An information processing system including an information processing device comprising a processor and a memory unit, wherein the processor executes a program described in any of the appendices 1 to 16. This allows users to see at a glance what topics the speakers discussed during the conversation.
[0244] (Note 19) An information processing method performed by a computer comprising a processor and a memory unit, wherein the processor is made to execute a program described in any of the appendices 1 to 16. This allows users to see at a glance what topics the speakers discussed during the conversation.
[0245] (Note 20) An information processing terminal comprising a processor and a display device, wherein the processor is capable of displaying information presented by a presentation step performed in an information processing device capable of executing any of the programs described in Appendix 5 to 8 on the display device. This allows users to see at a glance what topics the speakers discussed during the conversation. [Explanation of Symbols]
[0246] 1 System, 10 Servers, 101 Storage Unit, 104 Control Unit, 106 Input Device, 108 Output Device, 20 First User Terminal, 201 Storage Unit, 204 Control Unit, 206 Input Device, 208 Output Device, 30 Second User Terminal, 301 Storage Unit, 304 Control Unit, 306 Input Device, 308 Output Device, 50 CRM System, 501 Storage Unit, 504 Control Unit, 506 Input Device, 508 Output Device, 60 Voice Server (PBX), 601 Storage Unit, 604 Control Unit, 606 Input Device, 608 Output Device
Claims
1. A program for causing a computer to process data including a dialogue between a first user and a second user, the program comprising: a processor; and a storage unit, The program causes the processor to: obtaining one or more topics and one or more strings defining the topics; a step of identifying a topic related to at least a portion of a plurality of speech segments obtained by dividing the data in a time series manner based on the data, the topic, and the character string; A program that executes the following.
2. The program causes the processor to: a segmentation step of time-series segmenting the data into a plurality of speech segments; Run the command, the identifying step is a step of identifying a topic related to the speech section based on the speech section, the topic, and the character string. The program according to claim 1.
3. the identifying step includes a calculating step of calculating a degree of relevance between the voice section and a topic related to the voice section identified in the identifying step. The program according to claim 2.
4. The program causes the processor to: a presentation step of presenting the speech section in association with a topic related to the speech section identified in the identification step; Execute the The program according to claim 2.
5. The program causes the processor to: a keyword receiving step of receiving one or more keywords; a storage step of storing the keyword received in the keyword receiving step in association with a predetermined topic; Execute the The program according to claim 1.
6. The program causes the processor to: An extraction step of extracting one or more keywords included in the conversation; a storage step of storing the keywords extracted in the extraction step in association with a predetermined topic; Execute the The program according to claim 1.
7. A method executed by an information processing device including a processor and a storage unit, the processor executing all of the steps executed in any one of the inventions according to claims 1 to 6.
8. 10. An information processing device comprising a processor and a storage unit, the processor executing all the steps executed in the invention according to any one of claims 1 to 6.
9. A system comprising means for executing all the steps performed in any one of claims 1 to 6.