Program, information processing device, information processing system, information processing method, and information processing terminal

JP2023169092A5Pending Publication Date: 2025-05-16REVCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022169219
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-10-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

There is a lack of effective management of dialogue information between speakers based on their emotional states in online interactions.

Method used

A system that processes dialogue information by extracting emotional and impression features from voice and video data, calculating emotional states, identifying labels, and storing this information for managing dialogue interactions.

Benefits of technology

Enables the management of dialogue information based on the emotional states of speakers, allowing for improved interaction analysis and response preparation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To solve the problem that the dialogue information between speakers in the dialogue cannot be managed on the basis of an emotional state of the speakers.SOLUTION: A program causes a processor to execute a reception step of receiving the voice date related to the dialogue, a voice extraction step of extracting a plurality of section voice data for each utterance section from the voice data received in the reception step, an emotion calculation step of calculating, a plurality of emotion feature amount related to an emotional state of a speaker in the section voice data corresponding to each of the plurality of section voice data extracted in the voice extraction step, a label specifying step of specifying the label information on the dialogue based on the plurality of emotion feature amount calculated in the emotion calculation step, and a storage step of storing the label information specified in the label specifying step in association with the dialogue.SELECTED DRAWING: Figure 16
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a program, an information processing device, an information processing system, an information processing method, and an information processing terminal. [Background technology]

[0002] Online interactive services that are conducted among multiple users are known. Patent Document 1 discloses a technique for evaluating the sales activities of a person who is engaged in sales activities. Patent Document 2 discloses a technology that automatically evaluates the response of an operator in a call handling task, thereby reducing the burden of training the operators. Patent Document 3 discloses a learning support device that evaluates a learner or the learner's speech in consideration of the liveliness of the exchange of opinions. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-182390 [Patent Document 2] Japanese Patent Application Laid-Open No. 2007-286377 [Patent Document 3] Patent Publication No. 2020-091609 Summary of the Invention [Problem to be solved by the invention]

[0004] There is a problem in that the dialogue information between speakers in a dialogue cannot be managed. Therefore, the present disclosure has been made to solve the above problem, and its purpose is to provide a technology for managing dialogue information between speakers in a dialogue based on the emotional state of the speakers. [Means for solving the problem]

[0005] a processor and a storage unit, and a program for causing a computer to process information relating to a dialogue between a first user and a second user, the program causing the processor to execute: a reception step of receiving voice data relating to the dialogue; a voice extraction step of extracting a plurality of section voice data for each speech section from the voice data received in the reception step; an emotion calculation step of calculating a plurality of emotion features relating to the emotional state of a speaker in the section voice data, corresponding to each of the plurality of section voice data extracted in the voice extraction step; a label identification step of identifying label information for the dialogue based on the plurality of emotion features calculated in the emotion calculation step; and a storage step of storing the label information identified in the label identification step in association with the dialogue. [Effects of the Invention]

[0006] According to the present disclosure, dialogue information between speakers in a dialogue can be managed based on the emotional states of the speakers. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 2 is a block diagram showing the functional configuration of the system 1. [Figure 2] FIG. 2 is a block diagram showing the functional configuration of the server 10. [Figure 3] 2 is a block diagram showing the functional configuration of a first user terminal 20. FIG. [Figure 4] 3 is a block diagram showing the functional configuration of a second user terminal 30. FIG. [Figure 5] FIG. 2 is a block diagram showing the functional configuration of a CRM system 50. [Figure 6] FIG. 10 is a diagram showing the data structure of a user table 1012. [Figure 7] FIG. 10 is a diagram showing the data structure of an organization table 1013. [Figure 8] FIG. 10 is a diagram showing the data structure of a dialogue table 1014. [Figure 9] FIG. 10 is a diagram showing the data structure of a label table 1015. [Figure 10] FIG. 10 is a diagram showing the data structure of a speech segment table 1016. [Figure 11] FIG. 10 is a diagram showing the data structure of a topic relevance table 1017. [Figure 12] FIG. 10 is a diagram showing the data structure of the emotion condition master 1021. [Figure 13] FIG. 10 is a diagram showing the data structure of a speaker type master 1022. [Figure 14] FIG. 10 is a diagram showing the data structure of a topic master 1023. [Figure 15] FIG. 10 is a diagram showing the data structure of a customer table 5012. [Figure 16] 10 is a flowchart showing the operation of emotion analysis processing. [Figure 17] 10 is a flowchart showing the operation of impression analysis processing. [Figure 18] 10 is a flowchart showing the operation of a topic analysis process. [Figure 19] 10 is a flowchart showing the operation of a topic presentation process. [Figure 20] 10 is a screen example showing the operation of a topic presentation process. [Figure 21] FIG. 2 is a block diagram showing the basic hardware configuration of a computer 90. DETAILED DESCRIPTION OF THE INVENTION

[0008] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In all drawings describing the embodiments, common components are designated by the same reference numerals, and repeated description will be omitted. Note that the following embodiments do not unduly limit the content of the present disclosure described in the claims. Furthermore, not all components shown in the embodiments are necessarily essential components of the present disclosure. Furthermore, each drawing is a schematic diagram and is not necessarily a precise illustration.

[0009] <System 1 Configuration> The system 1 in the present disclosure is an information processing system that provides an online interactive service (online interactive service) between a first user who is an operator and a second user who is a customer. Note that the system 1 in the present disclosure may also be capable of providing an online interactive service between three or more users including the first user, the second user, and one or more other users. The system 1 includes information processing devices, namely, a server 10, a first user terminal 20, a second user terminal 30, a CRM system 50, and a voice server (PBX) 60, which are connected via a network N. FIG. 1 is a block diagram showing the functional configuration of the system 1. As shown in FIG. FIG. 2 is a block diagram showing the functional configuration of the server 10. As shown in FIG. FIG. 3 is a block diagram showing the functional configuration of the first user terminal 20. As shown in FIG. FIG. 4 is a block diagram showing the functional configuration of the second user terminal 30. As shown in FIG. FIG. 5 is a block diagram showing the functional configuration of the CRM system 50. As shown in FIG.

[0010] Each information processing device is configured by a computer equipped with an arithmetic unit and a storage device. The basic hardware configuration of the computer and the basic functional configuration of the computer realized by the hardware configuration will be described later. For each of the server 10, the first user terminal 20, the second user terminal 30, the CRM system 50, and the voice server (PBX) 60, descriptions that overlap with the basic hardware configuration and basic functional configuration of the computer will be omitted.

[0011] <Server 10 configuration> The server 10 is an information processing device that provides a service of storing and managing data (dialogue data) related to a dialogue between a first user and a second user. The server 10 includes a storage unit 101 and a control unit 104 .

[0012] <Configuration of the storage unit 101 of the server 10> The memory unit 101 of the server 10 includes an application program 1011, an emotion evaluation model 1031, an impression evaluation model 1032, a first impression evaluation model 1033, a second impression evaluation model 1034, a summary model 1035, a user table 1012, an organization table 1013, a dialogue table 1014, a label table 1015, a speech segment table 1016, a topic relevance table 1017, an emotion condition master 1021, a speaker type master 1022, and a topic master 1023.

[0013] The application program 1011 is a program for causing the control unit 104 of the server 10 to function as each functional unit. Application programs 1011 include applications such as a web browser application.

[0014] The emotion evaluation model 1031 is a model for outputting numerical intensities and values ​​for a plurality of emotional states, using as input data audio data, video data, or text data relating to the content of user comments in the audio data or video data.

[0015] The impression evaluation model 1032 is a model for outputting numerical intensities and values ​​for each of a plurality of impressions, using as input data audio data, video data, or text data relating to the content of user comments in the audio data or video data.

[0016] The first impression evaluation model 1033 is a model for outputting dialogue features related to a speaker's speaking style using audio data, video data, or text data related to the content of a user's utterances in the audio data or video data as input data. The dialogue features are features related to at least one of the speaker's speaking style, including speech rate, intonation, number of polite expressions, number of fillers, and number of grammatical utterances.

[0017] The second impression evaluation model 1034 is a model for outputting numerical intensities and values ​​for each of a plurality of impressions, using conversation features as input data.

[0018] User table 1012 is a table that stores and manages information about member users (hereinafter, "users") who use the service. When a user registers to use the service, the user's information is stored in a new record in user table 1012. This allows the user to use the service according to the present disclosure. The user table 1012 is a table having a user ID as a primary key and columns of a user ID, a CRM ID, an organization ID, a user name, and a user attribute. FIG. 6 is a diagram showing the data structure of the user table 1012.

[0019] The user ID is an item that stores user identification information for identifying a user. The user identification information is an item that is set with a unique value for each user. The CRMID is an item that stores user identification information for identifying a user in the CRM system 50. A user can receive CRM services by logging in to the CRM system 50 using the CRMID. The user ID in the server 10 is associated with the CRMID in the CRM system 50. The organization ID is an item for storing organization identification information for identifying an organization. The user name is an item for storing the name of the user. The user name may be set to any character string such as a nickname instead of a name. The user attributes are items that store information about the user's attributes, such as the user's age, gender, hometown, dialect, occupation (sales, customer support, etc.), etc. In addition to information about the user's personal attributes, the user attributes may also include information about company attributes, such as the industry, business size, and sales volume, of the organization, company, group, etc. to which the user belongs.

[0020] The organization table 1013 is a table that stores and manages information (organization information) about organizations to which users belong. Organizations include any organization or group, such as a company, a corporation, a corporate group, a club, or various associations. Organizations may also be defined for more detailed subgroups, such as company departments (sales department, general affairs department, customer support department). The organization table 1013 is a table having the organization ID as a primary key, and columns of organization ID, organization name, and organization attribute. FIG. 7 is a diagram showing the data structure of the organization table 1013.

[0021] The organization ID is an item for storing organization identification information for identifying an organization. The organization identification information is an item for which a unique value is set for each piece of organization information. The organization name is an item for storing the name of the organization. Any character string can be set as the organization name. The organizational attribute is an item for storing information about the attributes of an organization, such as the type of organization (company, corporate group, other organization, etc.) and the type of industry (real estate, finance, etc.).

[0022] The dialogue table 1014 is a table for storing and managing information (dialogue information) relating to a dialogue between a user and a customer. The dialogue table 1014 is a table having a dialogue ID as a primary key and columns of dialogue ID, user ID, customer ID, dialogue category, sending / receiving type, audio data, and video data. FIG. 8 is a diagram showing the data structure of the dialogue table 1014.

[0023] The dialogue ID is an item for storing dialogue identification information for identifying a dialogue. The dialogue identification information is an item for which a unique value is set for each piece of dialogue information. The user ID is an item for storing user identification information for identifying a user in a dialogue between a user and a customer. Multiple user IDs may be associated with each piece of dialogue information. The customer ID is an item for storing user identification information for identifying a customer in a dialogue between a user and a customer. The user IDs of multiple customers may be associated with each piece of dialogue information. The dialogue category is an item that stores the type (category) of dialogue between the user and the customer. The dialogue data is classified by the dialogue category. The dialogue category stores values ​​such as telephone operator, telemarketing, customer support, and technical support depending on the purpose of the dialogue between the user and the customer. The sending / receiving type is an item that stores information to distinguish whether the conversation between the user and the customer was sent by the user (outbound) or received by the user (inbound). In addition, when a conversation involves three or more users, the sending / receiving type "room" is stored. The audio data field stores audio data collected by a microphone. It may also store reference information (path) to an audio data file located elsewhere. The audio data may be in any format, such as AAC, ATRAC, mp3, or mp4. The voice data may be in a format in which identifiers are set that allow the user's voice and the customer's voice to be independently identifiable. In this case, the control unit 104 of the server 10 can perform independent analysis processing on the user's voice and the customer's voice. Furthermore, the user IDs of the user and the customer can be identified based on the voice data of the user and the customer. In the present disclosure, moving image data including audio information may be used instead of audio data. Also, audio data in the present disclosure includes audio data included in moving image data. Video data is an item that stores video data captured by a camera or other device. It may also store reference information (path) to video data files located elsewhere. The video data may be in any format, such as MP4, MOV, WMV, AVI, or AVCHD. The video data may be in a format in which identifiers are set that allow the user's video and the customer's video to be independently identifiable. In this case, the control unit 104 of the server 10 can perform independent analysis processing on the user's video and the customer's video. Furthermore, the user IDs of the user and the customer can be identified based on the user's and customer's video data.

[0024] The label table 1015 is a table for storing and managing information about labels (label information). The label table 1015 is a table having columns for a conversation ID and label data. FIG. 9 is a diagram showing the data structure of the label table 1015.

[0025] The dialogue ID is an item for storing dialogue identification information for identifying a dialogue. The label data is an item for storing label information for managing dialogues. The label information is additional information for managing dialogue information, such as a classification name, label, classification label, tag, etc. The label data may be a character string indicating the name of the label information, or may be a label ID or the like for referencing the name of the label information stored in another table. The label data includes classification information according to the emotional state of the speaker in a particular dialogue. The classification data includes classification information for classifying the speaker's response in a particular dialogue as good or bad.

[0026] The voice section table 1016 is a table for storing and managing information (voice section information) relating to a plurality of voice sections included in the dialogue information. The speech segment table 1016 is a table having the segment ID as the primary key and columns of segment ID, dialogue ID, speaker ID, start date and time, end date and time, segment audio data, segment video data, segment reading text, emotion data, impression data, and topic ID. FIG. 10 is a diagram showing the data structure of the speech segment table 1016.

[0027] The section ID is an item for storing section identification information for identifying a speech section, and a unique value is set for each piece of speech section information. The dialogue ID is an item for storing dialogue identification information for identifying a dialogue to which the voice section information is associated. The speaker ID is an item for storing speaker identification information for identifying a speaker to which the voice section information is associated. Specifically, the speaker ID is an item for storing the user IDs of multiple users who participated in the dialogue. The start date and time is an item for storing the start date and time of an audio section or a video section. The end date and time is an item for storing the end date and time of the audio section and the video section. The section audio data is an item that stores audio data included in an audio section. It may store reference information (path) for an audio data file located in another location. It may also store a reference to audio data for the period from the start date / time to the end date / time of the audio data in the dialogue table 1014 based on the start date / time and end date / time. The section audio data may also include audio data included in the section video data. The audio data may be in any format such as AAC, ATRAC, mp3, or mp4. The section video data is an item that stores video data included in the audio section. Reference information (path) for a video data file located in another location may also be stored. Furthermore, a reference to video data for the period from the start date / time to the end date / time of the video data in the dialogue table 1014 may also be stored based on the start date / time and the end date / time. The video data format can be any data format such as MP4, MOV, WMV, AVI, or AVCHD. The section reading text is an item that stores text information about the content spoken by the speaker in the section audio data included in the audio section.Specifically, the section reading text may be generated based on the section audio data and section video data by hand, or by using a learning model such as any machine learning or deep learning. Emotion data is an item that stores the emotional state of a speaker during a speech segment. Emotion data is a multidimensional scale (emotion vector) of multiple emotional states of a speaker, such as interest / excitement, joy, surprise, anxiety, anger, disgust, contempt, fear, shame, and guilt. Emotion data quantitatively expresses the emotional state of a speaker during a dialogue segment, as a numerical value representing the intensity of each of multiple emotional states (dimensions). Emotion data may also be configured to calculate and store an emotion scalar that indicates the intensity of one-dimensional emotions based on the emotion vector. Impression data is an item that stores the impression of a speaker during a speech section. The impression data is a multidimensional scale (vector) of multiple different impressions given by a speaker, such as like, dislike, noisy, difficult to listen to, polite, difficult to understand, timid, nervous, intimidating, violent, and sexual. It quantitatively expresses the impression given by a speaker during a dialogue section, as a numerical value representing the strength of each of multiple impressions (dimensions). The topic ID is an item for storing topic identification information associated with a speech section in a speech section.

[0028] The topic relevance table 1017 is a table for storing and managing information (topic relevance information) relating to the topic relevance for each speech section. The topic relevance table 1017 is a table having columns for section ID, topic ID, and relevance. FIG. 11 is a diagram showing the data structure of the topic relevance table 1017. As shown in FIG.

[0029] The section ID is an item for storing section identification information of a target voice section. The topic ID is an item for storing topic identification information for identifying a topic. The relevance is an item that stores information about the relevance of each topic identification information identified by a topic ID in a voice section included in the dialogue information. For one voice section, this item stores a numerical value that indicates the relevance with the topic identified by the topic ID. The higher the relevance, the stronger the relevance between the dialogue information and the topic.

[0030] The emotional condition master 1021 is a table for storing and managing information relating to emotional conditions (emotional condition information). The emotion condition master 1021 is a table having columns of emotion conditions and label data. FIG. 12 is a diagram showing the data structure of the emotion condition master 1021.

[0031] The emotion conditions are items for storing conditions related to emotion data. Specifically, conditions related to the threshold value, average value, regression coefficients when performing regression analysis, etc. of emotion data are stored. The label data is an item for storing label information associated with an emotional condition.

[0032] The speaker type master 1022 is a table for storing and managing information relating to impression conditions (impression condition information). The speaker type master 1022 is a table having columns for impression conditions and speaker types. FIG. 13 is a diagram showing the data structure of the speaker type master 1022.

[0033] The impression conditions are items for storing conditions related to impression data, such as thresholds, average values, and regression coefficients when performing regression analysis of the impression data. The speaker type is an item for storing the speaker type associated with the impression condition. The speaker type is a classification of the impression that a speaker gives to a conversation partner, such as forceful, modest, serious, friendly, active, emotional, etc.

[0034] The topic master 1023 is a table for storing and managing information related to topics (topic information). The topic master 1023 is a table having a topic ID as a primary key and columns of topic ID and keyword. FIG. 14 is a diagram showing the data structure of the topic master 1023.

[0035] The topic ID is an item that stores topic identification information for identifying a topic. The topic identification information is an item that has a unique value set for each piece of topic information. The keyword field stores multiple keywords associated with a topic. Specifically, multiple keywords are associated with one topic.

[0036] <Configuration of the control unit 104 of the server 10> The control unit 104 of the server 10 includes a user registration control unit 1041, an emotion analysis unit 1042, an impression analysis unit 1043, a topic processing unit 1044, and a learning unit 1051. The control unit 104 executes an application program 1011 stored in the storage unit 101, thereby realizing each functional unit.

[0037] The user registration control unit 1041 performs processing to store information about users who wish to use the service according to the present disclosure in the user table 1012. The information stored in the user table 1012 is generated when a user opens a web page operated by a service provider from any information processing terminal, enters information into a predetermined input form, and transmits the information to the server 10. The user registration control unit 1041 stores the received information in a new record in the user table 1012, completing the user registration. This allows the user stored in the user table 1012 to use the service. Before the user registration control unit 1041 registers the user information in the user table 1012, the service provider may conduct a predetermined examination to restrict whether or not the user is permitted to use the service. The user ID may be any character string or number that can identify the user, any character string or number desired by the user, or may be automatically set by the user registration control unit 1041.

[0038] The emotion analysis unit 1042 executes emotion analysis processing, the details of which will be described later.

[0039] The impression analysis unit 1043 executes impression analysis processing, the details of which will be described later.

[0040] The topic processing unit 1044 executes topic definition processing, topic analysis processing, and topic presentation processing, which will be described in detail later.

[0041] The learning unit 1051 executes the learning process.

[0042] <Configuration of First User Terminal 20> The first user terminal 20 is an information processing device operated by a first user who uses a service. The first user terminal 20 may be, for example, a desktop personal computer (PC) or laptop PC, or may be a mobile terminal such as a smartphone or tablet. It may also be a wearable terminal such as an HMD (Head Mount Display) or a wristwatch terminal. The first user terminal 20 includes a storage unit 201 , a control unit 204 , an input device 206 , and an output device 208 .

[0043] <Configuration of the storage unit 201 of the first user terminal 20> The storage unit 201 of the first user terminal 20 includes a first user ID 2011 and an application program 2012 .

[0044] The first user ID 2011 stores user identification information of the first user. The user transmits the first user ID 2011 from the first user terminal 20 to the voice server (PBX) 60. The voice server (PBX) 60 identifies the first user based on the first user ID 2011 and provides the first user with the service according to the present disclosure. The first user ID 2011 includes information such as a session ID temporarily assigned by the voice server (PBX) 60 to identify the user using the first user terminal 20.

[0045] The application program 2012 may be stored in advance in the storage unit 201, or may be configured to be downloaded from a web server or the like operated by a service provider via a communication IF. The application programs 2012 include applications such as a web browser application. The application program 2012 includes an interpreted programming language such as JavaScript (registered trademark) that runs on a web browser application stored on the first user terminal 20.

[0046] <Configuration of the control unit 204 of the first user terminal 20> The control unit 204 of the first user terminal 20 includes an input control unit 2041 and an output control unit 2042. The control unit 204 executes an application program 2012 stored in the storage unit 201, thereby realizing each functional unit.

[0047] <Configuration of the input device 206 of the first user terminal 20> The input device 206 of the first user terminal 20 includes a camera 2061 , a microphone 2062 , a position information sensor 2063 , a motion sensor 2064 , and a keyboard 2065 .

[0048] <Configuration of the output device 208 of the first user terminal 20> The output device 208 of the first user terminal 20 includes a display 2081 and a speaker 2082 .

[0049] <Configuration of second user terminal 30> The second user terminal 30 is an information processing device operated by a second user who uses the service. The second user terminal 30 may be, for example, a mobile terminal such as a smartphone or tablet, a stationary personal computer (PC) or a laptop PC, or a wearable terminal such as a head mounted display (HMD) or a wristwatch terminal. The second user terminal 30 includes a storage unit 301 , a control unit 304 , an input device 306 , and an output device 308 .

[0050] <Configuration of the storage unit 301 of the second user terminal 30> The storage unit 301 of the second user terminal 30 includes an application program 3012 and a telephone number 3013.

[0051] The application program 3012 may be pre-stored in the storage unit 301, or may be configured to be downloaded from a web server or the like operated by a service provider via a communication IF. The application program 3012 includes applications such as a web browser application. The application program 3012 includes an interpreter-type programming language such as JavaScript (registered trademark) that is executed on the web browser application stored in the second user terminal 30.

[0052] <Configuration of the control unit 304 of the second user terminal 30> The control unit 304 of the second user terminal 30 includes an input control unit 3041 and an output control unit 3042. The control unit 304 realizes each functional unit by executing the application program 3012 stored in the storage unit 301.

[0053] <Configuration of the input device 306 of the second user terminal 30> The input device 306 of the second user terminal 30 includes a camera 3061, a microphone 3062, a position information sensor 3063, a motion sensor 3064, and a touch device 3065.

[0054] <Configuration of the output device 308 of the second user terminal 30> The output device 308 of the second user terminal 30 includes a display 3081 and a speaker 3082.

[0055] <Configuration of the CRM system 50> The CRM system 50 is an information processing device managed and operated by an operator (CRM operator) that provides CRM (Customer Relationship Management) services. Examples of CRM services include SalesForce, HubSpot, Zoho CRM, and kintone. The CRM system 50 includes a storage unit 501 and a control unit 504.

[0056] <Configuration of the storage unit 501 of the CRM system 50> The storage unit 501 of the CRM system 50 includes an application program 5011 and a customer table 5012.

[0057] The application program 5011 is a program for causing the control unit 504 of the CRM system 50 to function as each functional unit. The application program 5011 includes applications such as a web browser application.

[0058] The customer table 5012 is a table for storing and managing user information (customer information) related to customers. The customer table 5012 is a table having columns for customer ID, user ID, name, phone number, and speaker type, with the customer ID as the primary key. FIG. 15 is a diagram showing the data structure of the customer table 5012.

[0059] The customer ID is an item for storing the user identification information of the customer. The user identification information is an item for which a unique value is set for each customer. The user ID is an item for storing the user identification information of the user who manages the customer. The name is an item for storing the name of the customer. The phone number is an item for storing the phone number of the customer. The user can make a call to the customer's phone number from the first user terminal 20 by accessing the website provided by the CRM system, selecting the customer to whom the call is to be made, and performing a predetermined operation such as "initiate call". The speaker type is an item that stores the speaker type of the user identified by the customer ID.

[0060] <Configuration of the control unit 504 of the CRM system 50> The control unit 504 of the CRM system 50 includes a user registration control unit 5041. The control unit 504 realizes each functional unit by executing the application program 5011 stored in the storage unit 501.

[0061] The user registration control unit 5041 performs a process of storing customer information in the customer table 5012 in the service according to the present disclosure. The information stored in the customer table 5012 is that the user opens a web page or the like operated by the service provider from an arbitrary information processing terminal, inputs information into a predetermined input form, and transmits it to the CRM system 50. The user registration control unit 5041 stores the received information in a new record of the customer table 5012, and the registration of the customer is completed. Thereby, the customer information is stored in association with the user ID of the user who manages the customer. The customer ID may be any string or number that can identify the user, any string or number desired by the user, or the user registration control unit may automatically set any string or number.

[0062] <Configuration of the voice server (PBX) 60> The voice server (PBX) 60 is an information processing device that functions as a switch that enables a conversation between the first user terminal 20 and the second user terminal 30 by connecting the network N and the telephone network T to each other. The voice server (PBX) 60 includes a storage unit 601.

[0063] <Configuration of the storage unit 601 of the voice server (PBX) 60> The storage unit 601 of the voice server (PBX) 60 includes an application program 6011 .

[0064] The application program 6011 is a program for causing the control unit 604 of the voice server (PBX) 60 to function as each functional unit. The application programs 6011 include applications such as a web browser application.

[0065] <System 1 Operation> Each process of the system 1 will be explained below. FIG. 16 is a flowchart showing the operation of the emotion analysis process. FIG. 17 is a flowchart showing the operation of the impression analysis process. FIG. 18 is a flowchart showing the operation of the topic analysis process. FIG. 19 is a flowchart showing the operation of the topic presentation process. FIG. 20 is an example of a screen showing the operation of the topic presentation process.

[0066] <Call processing> The outgoing call process is a process in which a user (first user) makes an outgoing call (call) to a customer (second user).

[0067] <Outline of outgoing call processing> The call processing is a series of processes in which the user selects a customer to whom he / she wishes to make a call from among multiple customers displayed on the screen of the first user terminal 20, and performs a call operation to make a call to the customer. In the present disclosure, a case in which a second user is selected as a customer will be described as an example.

[0068] <Details of outgoing call processing> The call processing of the system 1 when a user calls a customer will be described.

[0069] When a user makes a call to a customer, the following process is executed in the system 1.

[0070] The user operates the first user terminal 20 to launch a web browser and access the website of the CRM service provided by the CRM system 50. The user can display a list of his / her own customers on the display 2081 of the first user terminal 20 by opening a customer management screen provided by the CRM service. Specifically, the first user terminal 20 transmits a request to display a list of CRMIDs 2013 and customers to the CRM system 50. Upon receiving the request, the CRM system 50 searches the customer table 5012 and transmits information about the user's customers, such as customer IDs, names, telephone numbers, customer attributes, customer organization names, and customer organization attributes, to the first user terminal 20. The first user terminal 20 displays the received information about the customers on the display 2081 of the first user terminal 20.

[0071] The user presses and selects a customer (second user) to whom they wish to make a call from the list of customers displayed on the display 2081 of the first user terminal 20. With the customer selected, the user presses the "Call" button or the phone number button displayed on the display 2081 of the first user terminal 20 to send a request including the phone number to the CRM system 50. The CRM system 50, which receives the request, sends the request including the phone number to the server 10. The server 10, which receives the request, sends a call request to the voice server (PBX) 60. When the voice server (PBX) 60 receives the call request, it makes a call (call) to the second user terminal 30 based on the received phone number.

[0072] In response to this, the first user terminal 20 controls the speaker 2082 etc. to make a sound indicating that a call is being made (a call) by the voice server (PBX) 60. In addition, the display 2081 of the first user terminal 20 displays information indicating that a call is being made (a call) to the customer by the voice server (PBX) 60. For example, the display 2081 of the first user terminal 20 may display the words "Calling".

[0073] The customer can place the second user terminal 30 in a conversation-enabled state by lifting the receiver (not shown) of the second user terminal 30 or by pressing an "answer" button or the like that is displayed on the input device 306 of the second user terminal 30 when a call arrives. In response to this, the voice server (PBX) 60 transmits information indicating that a response has been made by the second user terminal 30 (hereinafter referred to as a "response event") to the first user terminal 20 via the server 10, the CRM system 50, etc. As a result, the user and the customer are able to interact using the first user terminal 20 and the second user terminal 30, respectively, and can converse with each other. Specifically, the user's voice collected by the microphone 2062 of the first user terminal 20 is output from the speaker 3082 of the second user terminal 30. Similarly, the customer's voice collected by the microphone 3062 of the second user terminal 30 is output from the speaker 2082 of the first user terminal 20.

[0074] When the display 2081 of the first user terminal 20 becomes ready for interaction, it receives the response event and displays information indicating that an interaction is taking place. For example, the display 2081 of the first user terminal 20 may display the words "Responding."

[0075] <Incoming call processing> The incoming call process is a process in which the user receives an incoming call (a call) from a customer.

[0076] <Outline of incoming call processing> The incoming call processing is a series of processes in which, when a user has an application running on the first user terminal 20, the user receives a call when a customer makes a call to the user.

[0077] <Details of incoming call processing> The following describes the incoming call processing of the system 1 when the user receives a call from a customer.

[0078] When the user receives a call from a customer, the following process is executed in the system 1.

[0079] The user operates the first user terminal 20 to launch a web browser and access the website of the CRM service provided by the CRM system 50. At this time, the user is assumed to be logged in to the CRM system 50 using his or her own account in the web browser and is on standby. Note that the user only needs to be logged in to the CRM system 50, and may also be performing other tasks related to the CRM service.

[0080] The customer operates the second user terminal 30, inputs a predetermined phone number assigned to the voice server (PBX) 60, and makes a call to the voice server (PBX) 60. The voice server (PBX) 60 receives the call made by the second user terminal 30 as an incoming call event.

[0081] The voice server (PBX) 60 transmits an incoming call event to the server 10. Specifically, the voice server (PBX) 60 transmits an incoming call request including the customer's telephone number 3011 to the server 10. The server 10 transmits the incoming call request to the first user terminal 20 via the CRM system 50. In response to this, the first user terminal 20 controls the speaker 2082 etc. to make a sound indicating that an incoming call is being received by the voice server (PBX) 60. The display 2081 of the first user terminal 20 displays information indicating that an incoming call is being received from the customer by the voice server (PBX) 60. For example, the display 2081 of the first user terminal 20 may display the words "Incoming call".

[0082] The first user terminal 20 accepts a response operation by the user. The response operation is realized, for example, by lifting a receiver (not shown) on the first user terminal 20, or by the user operating the mouse 2066 to press a button labeled "Answer the call" on the display 2081 of the first user terminal 20. When the first user terminal 20 receives the response operation, it transmits a response request to the voice server (PBX) 60 via the CRM system 50 and the server 10. The voice server (PBX) 60 receives the transmitted response request and establishes voice communication. This enables the first user terminal 20 to interact with the second user terminal 30. The display 2081 of the first user terminal 20 displays information indicating that a conversation is taking place. For example, the display 2081 of the first user terminal 20 may display the words "dialogue in progress."

[0083] <Modifications of outgoing call processing and incoming call processing> The method by which the first user becomes ready to interact with the second user is not limited to outgoing call processing and incoming call processing, and any method for realizing an interaction between the first user and the second user may be used. For example, a virtual interaction space called a room for the first user and the second user to interact may be created on the server 10, and the first user and the second user may access the room via a web browser or application program stored in the first user terminal 20 and the second user terminal 30, thereby enabling the interaction. In this case, the voice server (PBX) 50 is not required. Specifically, a first user who will be the host of the conversation operates the input device 206 of the first user terminal 20 to send a request to the server 10 to hold a conversation. Upon receiving the request, the control unit 104 of the server 10 issues room identification information such as a unique room ID and sends a response to the first user terminal 20. The first user then sends the received room identification information to the second user, the conversation partner, via any communication means such as email or chat. The first user can enter the room by operating the input device 206 of the first user terminal 20, accessing a URL that provides a room-related service on the server 10 using a web browser or the like, and entering the room identification information. Similarly, the second user can enter the room by operating the input device 306 of the second user terminal 30, accessing a URL that provides a room-related service on the server 10 using a web browser or the like, and entering the room identification information. This allows the first user and the second user to engage in a conversation via the first user terminal 20 and the second user terminal 30, respectively, in a virtual conversation space called a room, which is associated by the room identification information. By inputting room identification information, one or more other users can enter a single room in addition to the first and second users. This allows three or more users to converse via their respective user terminals in a virtual conversation space called a room, which is associated by the room identification information.

[0084] <Video Dialogue> The system 1 in the present disclosure may provide an online interactive service (video interactive service) including video data. For example, the control unit 204 of the first user terminal 20 and the control unit 304 of the second user terminal 30 transmit video data captured by the camera 2061 of the first user terminal 20 and the camera 3061 of the second user terminal 30, respectively, to the server 10. Based on the received video data, the server 10 transmits video data captured by the camera 2061 of the first user terminal 20 to the second user terminal 30, and transmits video data captured by the camera 3061 of the second user terminal 30 to the first user terminal 20. The control unit 204 of the first user terminal 20 displays the received video data captured by the camera 3061 of the second user terminal 30 on the display 2081. The control unit 304 of the second user terminal 30 displays the received video data captured by the camera 2061 of the first user terminal 20 on the display 3081. The server 10 may transmit video data of some or all of the multiple users participating in the online dialogue to the first user terminal 20 and the second user terminal 30. In this case, the control unit 204 of the first user terminal 20 displays the received video data of some or all of the multiple users participating in the online dialogue on a single screen on the display 2081 of the first user terminal 20. This allows the dialogue status of the multiple users participating in the online dialogue to be confirmed. A similar process may also be performed in the second user terminal 30.

[0085] <Interactive Amnestic Processing> The dialogue storage process is a process for storing data relating to a dialogue between a user and a customer.

[0086] <Outline of interactive memory processing> The dialogue storage process is a series of processes for storing data relating to a dialogue in the dialogue table 1014 when a dialogue is started between a user and a customer.

[0087] <Details of interactive memory processing> When a conversation between a user and a customer begins, the voice server (PBX) 60 records voice data relating to the conversation between the user and the customer and transmits it to the server 10. When the control unit 104 of the server 10 receives the voice data, it creates a new record in the conversation table 1014 and stores the data relating to the conversation between the user and the customer. Specifically, the control unit 104 of the server 10 stores the user ID, customer ID, conversation category, call reception / transmission type, and the content of the voice data in the new record in the conversation table 1014.

[0088] The control unit 104 of the server 10 acquires the first user ID 2011 of the first user from the first user terminal 20 during the outgoing call processing or incoming call processing, and stores it in the user ID field of the new record in the dialogue table 1014. The control unit 104 of the server 10 queries the CRM system 50 based on the telephone number during outgoing call processing or incoming call processing. The CRM system 50 searches the customer table 5012 by telephone number to obtain the customer ID and transmits it to the server 10. The control unit 104 of the server 10 stores the obtained customer ID in the customer ID field of a new record in the dialogue table 1014. The control unit 104 of the server 10 stores the dialogue category value set in advance for each user or customer in the dialogue category field of the new record in the dialogue table 1014. Note that the dialogue category may also be stored by the user selecting and inputting a value for each dialogue. The control unit 104 of the server 10 identifies whether the ongoing conversation was initiated by the user or the customer, and stores either the value of outbound (initiated by the user) or inbound (initiated by the customer) in the incoming / outgoing call type field of the new record in the conversation table 1014.

[0089] The control unit 104 of the server 10 stores the voice data received from the voice server (PBX) 60 in the voice data field of a new record in the dialogue table 1014. Note that the voice data may be stored as a voice data file in another location, and reference information (path) for the voice data file may be stored after the dialogue ends. The control unit 104 of the server 10 may also be configured to store the voice data after the dialogue ends.

[0090] Furthermore, in the video dialogue service, the control unit 104 of the server 10 stores video data received from the first user terminal 20 and the second user terminal 30 in the video data item of a new record in the dialogue table 1014. Note that the video data may be stored as a video data file in another location, and reference information (path) for the video data file may be stored after the dialogue ends. The control unit 104 of the server 10 may also be configured to store the video data after the dialogue ends.

[0091] <Emotion analysis processing> Emotion analysis processing is a process that analyzes dialogue information such as audio and video of online dialogues conducted by multiple users, identifies the emotional states of the users participating in the dialogue, identifies label information based on the emotional states, and stores the information in association with the dialogue information.

[0092] <Outline of emotion analysis processing> The emotion analysis process is a series of processes that, when an online dialogue between users is detected, stores dialogue information about the dialogue, divides the audio data and video data contained in the dialogue information into section data such as section audio data and section video data for each speech section, calculates emotional features for each section data, identifies label information based on the emotional features, and stores the label information in association with the dialogue information.

[0093] <Details of emotion analysis processing> The emotion analysis process will be described in detail below.

[0094] In step S101, online conversation between the user and the customer is started via the outgoing call processing, incoming call processing, room, etc., which have already been described.

[0095] In step S102, the emotion analysis unit 1042 of the server 10 executes a receiving step of receiving voice data related to the dialogue. Specifically, through the dialogue storage process, the first user terminal 20 transmits the first user ID 2011, the audio data collected by the microphone 2062, and the video data captured by the camera 2061 to the server 10. The control unit 104 of the server 10 stores the received first user ID 2011, the audio data, and the video data in the user ID, audio data, and video data items of a new record in the dialogue table 1014, respectively. Similarly, the second user terminal 30 transmits the second user ID 3011, audio data collected by the microphone 3062, and video data captured by the camera 3061 to the server 10. The control unit 104 of the server 10 stores the received second user ID 3011, audio data, and video data in the user ID, audio data, and video data items of a new record in the dialogue table 1014, respectively. Accordingly, a new dialogue ID is assigned and stored in the dialogue ID field of the new record in the dialogue table 1014.

[0096] In step S103, the emotion analysis unit 1042 of the server 10 executes a voice extraction step of extracting a plurality of section voice data for each utterance section from the voice data received in the receiving step. Specifically, the emotion analysis unit 1042 of the server 10 acquires (accepts) the dialogue ID, audio data, and video data stored in the dialogue table 1014 in step S102. The emotion analysis unit 1042 of the server 10 detects sections in which audio exists (utterance sections) from the acquired (accepted) audio data and video data, and extracts the audio data and video data for each utterance section as section audio data and section video data, respectively. The section audio data and section video data are associated with the speaker's user ID, the start date and time of the utterance section, and the end date and time of the utterance section for each speech section. The emotion analysis unit 1042 of the server 10 performs text recognition on the extracted speech content of the section audio data and section video data, converting the section audio data and section video data into section read-aloud text, which is characters (text), and transcribing it. Note that the specific method of text recognition is not particularly limited. For example, conversion may be performed using signal processing technology, machine learning using AI (artificial intelligence), deep learning, or the like.

[0097] The emotion analysis unit 1042 of the server 10 stores the dialogue ID to be processed, the speaker's user ID (first user ID 2011 or second user ID 3011), start date and time, end date and time, section audio data, section video data, and section reading text in the dialogue ID, speaker ID, start date and time, end date and time, section audio data, section video data, and section reading text fields of a new record in the audio section table 1016, respectively.

[0098] The section reading text for each speech section of the voice data is stored as continuous time-series data in association with the start date and time and the speaker in the voice section table 1016. By checking the section reading text stored in the voice section table 1016, the user can check the dialogue content as text information without checking the content of the voice data.

[0099] In addition, during the text recognition process, information that is meaningless in terms of understanding the conversation between the user and the customer, such as fillers contained in the text, may be excluded from the text in advance, and the speech recognition information may be stored in the speech segment table 1016.

[0100] In step S104, the emotion analysis unit 1042 of the server 10 executes an emotion calculation step of calculating a plurality of emotion features relating to the emotional state of the speaker in the section audio data, corresponding to each of the plurality of section audio data extracted in the audio extraction step. The emotion calculation step calculates the emotion features as output data by applying the section audio data extracted in the audio extraction step as input data to a learning model. Specifically, the emotion analysis unit 1042 of the server 10 acquires the section audio data, section video data, and section reading text stored in the audio section table 1016 in S103, and applies these as input data to the emotion evaluation model 1031, and the emotion evaluation model 1031 outputs emotion features according to the input data as output data.

[0101] In step S104, an emotion calculation step executes a step of calculating an emotion vector indicating intensity related to multidimensional emotions corresponding to each of the plurality of section audio data extracted in the audio extraction step. Specifically, the emotion analysis unit 1042 of the server 10 acquires the section audio data, section video data, and section read-out text stored in the audio section table 1016 in S103, and applies these as input data to the emotion evaluation model 1031. The emotion evaluation model 1031 outputs as output data the intensity of each of multiple emotional states (dimensions) according to the input data, and an emotion vector quantitatively expressed as a numerical value.

[0102] The emotion calculation step executes a step of calculating an emotion scalar indicating intensity related to one-dimensional emotion corresponding to each of the plurality of section audio data extracted in the audio extraction step, based on the calculated emotion vector. The emotion analysis unit 1042 of the server 10 calculates an emotion scalar, which indicates the intensity of one-dimensional emotions, by applying principal component analysis, a learning model such as a deep learning model, and calculations for each component of the emotion vector to the emotion vector. For example, the emotion scalar is an index that quantitatively expresses the positivity or negativity of the speaker's emotional state in the speech segment information, and may be numerical data normalized to a range of values ​​from +1 (positive) to -1 (negative).

[0103] The emotion analysis unit 1042 of the server 10 stores the calculated emotion features, that is, the emotion vector and emotion scalar, in the emotion data field of the record to be analyzed in the speech segment table 1016. The emotion data field may be configured to store either the emotion vector or the emotion scalar.

[0104] In step S104, the emotion analysis unit 1042 of the server 10 searches the user table 1012 for a user ID based on the speaker ID of the record to be analyzed in the speech segment table 1016, and acquires the user attributes.

[0105] In step S105, the emotion analysis unit 1042 of the server 10 executes a label specification step of specifying label information for the dialogue based on the plurality of emotion feature amounts calculated in the emotion calculation step. Specifically, the emotion analysis unit 1042 of the server 10 searches the speech segment table 1016 for the dialogue ID based on the dialogue ID, and acquires the emotion data item. The emotion analysis unit 1042 of the server 10 searches the emotion condition master 1021 based on the emotion data for the presence or absence of a record that matches the emotion condition, and acquires the label data item of the corresponding record. In the present disclosure, the emotion analysis unit 1042 of the server 10 may be configured to calculate a plurality of pieces of speech section information extracted from one piece of dialogue information, and identify and acquire label data using a plurality of emotion features corresponding to a plurality of pieces of stored emotion data as emotion conditions.

[0106] In step S105, a label specifying step executes a step of specifying label information for the dialogue based on the plurality of emotion scalars calculated in the emotion calculating step. Specifically, the emotion analysis unit 1042 of the server 10 may calculate the emotion scalar included in the stored emotion data for each of the plurality of pieces of speech section information extracted for one piece of dialogue information, and identify the label data using the emotion scalar as the emotion condition.

[0107] In step S105, a label specifying step executes a step of specifying label information for the dialogue based on the plurality of emotion vectors calculated in the emotion calculation step. Specifically, the emotion analysis unit 1042 of the server 10 may calculate and store emotion vectors included in the emotion data for each of the plurality of speech section information extracted from one piece of dialogue information, and identify label data using the emotion vectors as emotion conditions. For example, the emotion conditions may be configured to be specified by the ranges for each element component of the emotion vector.

[0108] In step S105, a label specifying step executes a step of specifying label information for the dialogue based on the number of emotion feature amounts that are equal to or greater than a predetermined threshold value among the plurality of emotion feature amounts calculated in the emotion calculation step. Specifically, it is assumed that a predetermined threshold and a predetermined number of pieces of information equal to or greater than the threshold are stored in the emotion condition item of the emotion condition master 1021. The emotion analysis unit 1042 of the server 10 compares the emotion scalar values ​​corresponding to each of the multiple pieces of voice section information extracted for one piece of dialogue information with the predetermined threshold, and counts the number of pieces of voice section information (emotion scalars) equal to or greater than the predetermined threshold. Note that it is also possible to count the number of pieces of voice section information equal to or less than the predetermined threshold. If the number of counted speech section information pieces is greater than a predetermined number, the emotion analysis unit 1042 of the server 10 determines that the emotional condition is met, and acquires and identifies the label data items associated with the emotional condition in the emotional condition master 1021. For example, if the number of speech section information (emotion scalars) above a predetermined threshold is greater than a predetermined number, label information indicating that the emotional state in the dialogue is positive is identified. Similarly, if the number of speech section information (emotion scalars) below a predetermined threshold is greater than a predetermined number, label information indicating that the emotional state in the dialogue is negative is identified.

[0109] In step S105, a label specifying step executes a step of specifying label information for the dialogue based on the proportion of emotion feature amounts above or below a predetermined threshold among the plurality of emotion feature amounts calculated in the emotion calculation step. Specifically, it is assumed that a predetermined threshold and a percentage (predetermined percentage) of values ​​above the threshold are stored in the emotion condition item of the emotion condition master 1021. The emotion analysis unit 1042 of the server 10 compares the emotion scalar values ​​corresponding to each of the multiple pieces of voice section information extracted for one piece of dialogue information with the predetermined threshold, and counts the number of pieces of voice section information (emotion scalars) above the predetermined threshold. Note that it is also possible to count the number of pieces below the predetermined threshold. If the ratio of the number of counted speech section information to the number of all speech section information extracted for one piece of dialogue information is greater than a predetermined ratio, the emotion analysis unit 1042 of the server 10 determines that the emotional condition is met, and acquires and identifies the label data item associated with the emotional condition in the emotion condition master 1021. For example, if the proportion of speech section information (emotion scalars) above a predetermined threshold is greater than a predetermined proportion, label information indicating that the emotional state in the dialogue is positive is identified. Similarly, if the proportion of speech section information (emotion scalars) below a predetermined threshold is greater than a predetermined proportion, label information indicating that the emotional state in the dialogue is negative is identified.

[0110] Note that instead of an emotion scalar, a single element component included in an emotion vector, an index calculated based on one or more element components included in an emotion vector, or the like may be regarded as an emotion feature, and similar processing may be performed.

[0111] In step S105, a label specifying step executes a step of specifying label information for the dialogue based on the statistical values ​​of the plurality of emotion feature amounts calculated in the emotion calculation step. Specifically, it is assumed that information on a predetermined threshold is stored in the emotion condition item of the emotion condition master 1021. The emotion analysis unit 1042 of the server 10 calculates statistical values ​​such as the average, median, mode, maximum, and minimum of the emotion scalar values ​​corresponding to each of the multiple pieces of voice section information extracted for one piece of dialogue information, compares them with a predetermined threshold, and if the calculated value is equal to or greater than the predetermined threshold, determines that the emotion condition applies, and acquires and identifies the label data item associated with the emotion condition in the emotion condition master 1021. Note that the condition may also be equal to or less than the predetermined threshold.

[0112] In step S105, a label specifying step executes a step of specifying label information for the dialogue based on time-series changes in the plurality of emotion feature amounts calculated in the emotion calculation step. The label identification step includes a step of performing a regression analysis on time-series changes in the plurality of emotion features calculated in the emotion calculation step, and a step of identifying label information for the dialogue based on regression coefficients obtained as a result of the regression analysis. Specifically, it is assumed that a range of regression coefficients is stored in the emotion condition item of the emotion condition master 1021. For the target dialogue data, a regression analysis of Y=f(X) is performed for each of the multiple pieces of speech section information associated with the dialogue data, where the X axis represents the start date / time, end date / time, and any date and time between the start date / time and end date / time of the speech section information, and the Y axis represents the emotion scalar value included in the emotion data of the speech section information. Any regression analysis, such as linear regression or quadratic regression, may be applied. The regression coefficient is calculated by performing the regression analysis and compared with the range of the regression coefficient. If the value is within the range of the regression coefficient, it is determined that the emotion condition applies, and the label data item associated with the emotion condition in the emotion condition master 1021 is obtained and identified. For example, in the case of linear regression (first-order regression), if the intercept is negative and the slope is positive, label information indicating that the emotional state in the dialogue is improving is identified. Note that instead of an emotion scalar, a single element component included in an emotion vector, an index calculated based on one or more element components included in an emotion vector, or the like may be regarded as an emotion feature, and similar processing may be performed.

[0113] In step S105, the emotion analysis unit 1042 of the server 10 executes a step of identifying a first emotion group which is a set of multiple emotion features corresponding to multiple chronologically consecutive section audio data extracted in the audio extraction step. The emotion analysis unit 1042 of the server 10 executes a step of identifying a second emotion group which is a set of multiple emotion features corresponding to multiple chronologically consecutive section audio data extracted in the audio extraction step. Specifically, the emotion analysis unit 1042 of the server 10 may divide the plurality of pieces of speech segment information extracted for one piece of dialogue information into section groups each consisting of a plurality of pieces of speech segment information, and execute the label identification step already described for each section group, thereby identifying label information corresponding to each of the plurality of section groups. For example, the emotion analysis unit 1042 of the server 10 calculates an emotion scalar for each of the extracted speech segment information included in the segment group and stores it in the emotion data. The emotion scalar included in the stored emotion data may be used as an emotion condition to identify label data. For example, the emotion analysis unit 1042 of the server 10 calculates an emotion vector for each of the extracted speech segment information included in the segment group and stores the emotion vector in the emotion data. The emotion vector included in the stored emotion data may be used as an emotion condition to identify label data.

[0114] In step S105, the label identification step includes a step of identifying first label information for the dialogue based on a plurality of emotion feature amounts included in a first emotion group, and a step of identifying second label information for the dialogue based on a plurality of emotion feature amounts included in a second emotion group. Specifically, the emotion analysis unit 1042 of the server 10 divides the multiple pieces of speech section information extracted for one piece of dialogue information into section groups each consisting of multiple pieces of speech section information, and performs the label identification step already described for each section group, thereby identifying label information corresponding to each of the multiple section groups.

[0115] In step S105, the emotion analysis unit 1042 of the server 10 executes a label presenting step of presenting the first label information and the second label information to the first user. Specifically, the emotion analysis unit 1042 of the server 10 transmits the identified first label information and second label information to the first user terminal 20. The control unit 204 of the first user terminal 20 displays the received first label information and second label information on the display 2081 of the first user terminal 20, and presents them to the first user. Note that the first label information and second label information may be presented to any user, such as the second user, another administrator, or another user.

[0116] In step S105, the emotion analysis unit 1042 of the server 10 executes a selection receiving step of receiving, from the first user, a selection instruction to select at least one of the first label information and the second label information presented in the label presenting step. Specifically, the first user selects either the first label information or the second label information presented on the display 2081 of the first user terminal 20 by operating the input device 206 of the first user terminal 20 or the like. Note that the first user may not select either of the information. The control unit 204 of the first user terminal 20 transmits the selected label information to the server 10. The emotion analysis unit 1042 of the server 10 identifies the received label information.

[0117] In step S105, a label identification step executes a step of identifying label information for the dialogue based on the plurality of emotion features calculated in the emotion calculation step and the user attributes of the first user or the second user who uttered the section audio data corresponding to the plurality of emotion features. Specifically, when identifying the label information, the emotion analysis unit 1042 of the server 10 may identify the label information by taking into consideration the user attributes of the first user and the second user identified in step S104. For example, the emotion conditions in the emotion condition master 1021 may include the user attributes of the first user and the second user as conditions.

[0118] In step S105, the label identification step executes a step of identifying label information for the dialogue based on the plurality of emotion features corresponding to the section audio data related to the speech of the second user calculated in the emotion calculation step, without taking into consideration the plurality of emotion features corresponding to the section audio data related to the speech of the first user. Specifically, the emotion analysis unit 1042 of the server 10 may exclude speech section information whose speaker ID is the first user ID 2011 from the multiple speech section information extracted for one piece of dialogue information, and may perform the label identification step already described based only on speech section information whose speaker ID is the second user ID 3011. This allows for identification of label information that takes into account only the emotional state of the customer. Typically, a first user, such as an operator, is interested in the emotional state of the customer, not his or her own. This configuration allows for identification of label information that takes into account the emotional state of the customer in particular.

[0119] The emotion analysis unit 1042 of the server 10 may exclude speech section information whose speaker ID is the second user ID 3011 from the multiple speech section information extracted for one piece of dialogue information, and may perform the label identification step already described based only on the speech section information whose speaker ID is the first user ID 2011.

[0120] The emotion analysis unit 1042 of the server 10 may perform the label identification step already described for each of the speech section information whose speaker ID is the first user ID 2011 and the speech section information whose speaker ID is the second user ID 3011, and identify multiple pieces of label information, namely, the first label information and the second label information, respectively.

[0121] Furthermore, the emotion analysis unit 1042 of the server 10 may exclude, from the multiple pieces of speech section information extracted for one piece of dialogue information, speech section information in which the user identified by the speaker ID is the host user who is the organizer of the dialogue, and perform the label identification step already described based only on speech section information in which the user identified by the speaker ID is not the host user. This allows label information to be identified without taking into account the emotional state of the conversation host. Typically, a conversation host is interested in the emotional state of the conversation partner, not their own emotional state. This configuration allows label information to be identified that takes into account the emotional state of the conversation partner.

[0122] In step S106, the emotion analysis unit 1042 of the server 10 executes a storage step of storing the label information identified in the label identification step in association with the dialogue. Specifically, the emotion analysis unit 1042 of the server 10 stores the label information identified in step S105 in the label data item of the label table 1015 in association with the dialogue ID numbered in step S101. In step S105, the identified label information may be presented to the first user, and the label information for which a selection instruction is received from the first user may be stored as label data in the label table 1015.

[0123] In step S106, the storage step executes a step of storing the first label information or the second label information identified in the label identification step in association with the dialogue. The storage step executes a step of storing at least one of the first label information and the second label information in association with the dialogue based on the selection instruction accepted from the first user in the selection accepting step. Specifically, the label information for which a selection instruction is received from the first user may be stored as label data in the label table 1015.

[0124] In addition, the first user can display the label information stored in the label table 1015 from the server 10 on the display 2081 of the first user terminal 20 by operating the input device 206 of the first user terminal 20.

[0125] <Regarding the timing of emotion analysis processing> Steps S103 to S106 of the emotion analysis process may be configured to be executed after an online dialogue between multiple users has ended. As a result, after the online dialogue has ended and the dialogue content has been determined, label information according to the emotional state of the users in the dialogue is identified and stored in association with the dialogue information.

[0126] Furthermore, the emotion analysis process may be configured to be executed after the start of an online dialogue between multiple users and before the dialogue ends. That is, the steps may be executed at any timing during an online dialogue between multiple users. Also, steps S103 to S106 may be executed periodically in real time during the online dialogue. As a result, even during the online dialogue, label information according to the emotional state of the user in the dialogue up to that point may be identified and stored in association with the dialogue information. This allows the user to check the emotional states of other users participating in the online dialogue in real time during the online dialogue, and also allows the user to organize and manage the dialogue information based on the latest emotional states.

[0127] <Impression analysis processing> The impression analysis process analyzes dialogue information such as audio and video of online dialogues conducted by multiple users, identifies the impression states of the users participating in the dialogue, and presents the impression states and speaker types to the users.

[0128] <Outline of impression analysis processing> The impression analysis process is a series of processes that, when an online conversation between users is detected, stores conversation information about the conversation, divides the audio data and video data contained in the conversation information into section data such as section audio data and section video data for each speech section, calculates impression features for each section data, identifies a speaker type based on the impression features, and presents the identified speaker type to the user.

[0129] <Details of impression analysis processing> The impression analysis process will be described in detail below.

[0130] In step S301, online conversation between the user and the customer is started via the outgoing call processing, incoming call processing, room, etc., which have already been described.

[0131] In step S302, the impression analysis unit 1043 of the server 10 executes a dialogue acquisition step of acquiring dialogue information relating to a dialogue between the second user and the first user from the second user. Step S302 is the same as step S102 in the emotion analysis process, and therefore a description thereof will be omitted.

[0132] In step S303, the impression analysis unit 1043 of the server 10 executes a voice extraction step of extracting a plurality of section voice data for each speech section from the voice data of the second user received in step S302. Step S303 is the same as step S103 in the emotion analysis process, and therefore a description thereof will be omitted.

[0133] In step S304, the impression analysis unit 1043 of the server 10 executes an impression calculation step of calculating impression features relating to the impression that the second user gives to other users in the dialogue, based on the dialogue information of the second user acquired in the dialogue acquisition step. The impression calculation step executes a step of calculating impression features indicating the intensity of at least one impression of like, dislike, noisy, difficult to listen to, polite, difficult to understand, timid, nervous, intimidating, violent, and sexual, based on the dialogue information acquired from the second user in the dialogue acquisition step. The impression calculation step executes a step of calculating, as output data, impression features relating to the impression that the second user gives to other users in the dialogue by applying the dialogue information acquired from the second user in the dialogue acquisition step as input data to a learning model. Specifically, the impression analysis unit 1043 of the server 10 acquires the section audio data, section video data, and section read-out text stored in the audio section table 1016 in S303, excludes audio section information whose speaker ID is the first user ID 2011 from the audio section information, and applies only the audio section information whose speaker ID is the second user ID 3011 as input data to the impression evaluation model 1032, and the impression evaluation model 1032 outputs impression features according to the input data as output data. This makes it possible to evaluate the impression given by the second user based on the impression features. The input data applied to the impression evaluation model 1032 may exclude speech section information whose speaker ID is the second user ID 3011 from the speech section information, and may instead be speech section information whose speaker ID is the first user ID 2011. In this case, the impression given by the first user can be evaluated by impression features.

[0134] In step S304, the impression calculation step includes a step of calculating dialogue features related to the speaking style of the second user in the dialogue based on the dialogue information of the second user acquired in the dialogue acquisition step, and a step of calculating impression features based on the calculated dialogue features. The impression calculation step includes a step of calculating dialogue features related to the speaking style of the second user in the dialogue as output data by applying the dialogue information of the second user acquired in the dialogue acquisition step as input data to a first learning model, and a step of calculating impression features by applying the calculated dialogue features as input data to a second learning model. The impression calculation step includes a step of calculating dialogue features related to at least one of the speaking style of the second user in the dialogue, including speech rate, intonation, number of polite expressions, number of fillers, and number of grammatical utterances, based on the dialogue information of the second user acquired in the dialogue acquisition step.

[0135] Specifically, the impression analysis unit 1043 of the server 10 acquires the section audio data, section video data, and section readout text stored in the audio section table 1016 in S303, excludes audio section information whose speaker ID is the first user ID 2011 from the audio section information, and applies only the audio section information whose speaker ID is the second user ID 3011 as input data to the first impression evaluation model 1033, and the first impression evaluation model 1033 outputs dialogue features corresponding to the input data as output data. The impression analysis unit 1043 of the server 10 applies the dialogue features as input data to the second impression evaluation model 1034, and the second impression evaluation model 1034 outputs the impression features according to the input data as output data. This makes it possible to evaluate the impression given by the second user based on the impression features. The input data applied to the impression evaluation model 1032 may exclude speech section information whose speaker ID is the second user ID 3011 from the speech section information, and may instead be speech section information whose speaker ID is the first user ID 2011. In this case, the impression given by the first user can be evaluated by impression features.

[0136] In step S304, the impression analysis unit 1043 of the server 10 executes a storage step of storing the impression feature amount calculated in the impression calculation step in association with the second user. Specifically, the impression analysis unit 1043 of the server 10 stores the calculated impression feature in the impression data item of the record to be analyzed in the speech segment table 1016. As a result, the impression feature is stored in association with the second user via the speaker ID (second user ID) of the speech segment table 1016. Note that the impression feature may be stored in association with the second user ID by providing a column for storing impression data (not shown) in the customer table 5012 of the CRM system 50. The impression feature may be stored in association with the second user ID by providing a column for storing impression data (not shown) in the user table 1012 of the server 10. By storing the impression features of the user identified in the target conversation in the customer table 5012 of the CRM system 50, it is possible to share the impression features of the user with members of other departments in the company. For example, it is possible to carry out work efficiently according to the impression of the conversation partner identified by the impression features.

[0137] In step S305, the impression analysis unit 1043 of the server 10 executes an identification step of identifying a speaker type that labels the impression that the second user gives to other users, based on the impression feature amount calculated in the impression calculation step. Specifically, the impression analysis unit 1043 of the server 10 searches the speech segment table 1016 for a dialogue ID based on the dialogue ID, and acquires the impression data item. The impression analysis unit 1043 of the server 10 searches the speaker type master 1022 for a record that matches the impression condition based on the impression data, and acquires the speaker type item of the corresponding record. In the present disclosure, the impression analysis unit 1043 of the server 10 may be configured to calculate impression features for each of a plurality of pieces of speech section information extracted for one piece of dialogue information, and to identify and acquire the speaker type using the impression features related to the plurality of pieces of stored impression data as impression conditions.

[0138] In step S305, the impression analysis unit 1043 of the server 10 executes a storing step of storing the speaker type identified in the identifying step in association with the second user. Specifically, the impression analysis unit 1043 of the server 10 transmits the identified speaker type and second user ID to the CRM system 50. The control unit 504 of the CRM system 50 stores the received speaker type and second user ID in the speaker type and user ID items of the customer table 5012, respectively. In other words, the identified speaker type is stored in association with the user ID of the user who spoke in the dialogue. By storing the information in the customer table 5012 of the CRM system 50, the speaker type of the user identified in the target conversation can be shared with members of other departments within the company. For example, it is possible to provide efficient customer service depending on the speaker type of the person with whom the conversation is being held. In the present disclosure, the speaker type of the user is stored in the customer table 5012 of the CRM system 50, but it may be stored in the user table 1012 of the server 10 in association with the second user.

[0139] In step S306, the impression analysis unit 1043 of the server 10 executes a presenting step of presenting to the first user the impression feature amount that has been stored in association with the second user in the storing step. Specifically, the impression analysis unit 1043 of the server 10 transmits the impression feature identified in step S305 to the first user terminal 20. The control unit 204 of the first user terminal 20 displays the received impression feature on the display 2081 of the first user terminal 20, and presents it to the first user. Note that the impression feature may be presented to any user, such as the second user, another administrator, or another user.

[0140] In step S306, prior to a dialogue between the first user and the second user, the impression analysis unit 1043 of the server 10 executes a presentation step of presenting to the first user the impression features stored in association with the second user in the storage step. For example, when the first user or another user starts an online dialogue with the second user via outgoing call processing, incoming call processing, a room, etc., the impression features of the second user stored in association with the second user in step S305 may be displayed on an outgoing call screen for making a call to the second user, an incoming call screen for receiving a call from the second user, a room screen before the dialogue starts, etc., which are displayed on the display 2081 of the first user terminal 20, and presented to the first user. This allows the first user to prepare a response that matches the impression the second user has of the first user before the start of the dialogue.

[0141] In addition, the impression analysis unit 1043 of the server 10 may perform a presentation step of presenting to the first user the speaker type that was stored in association with the second user in the storage step prior to a dialogue between the first user and the second user. For example, when a first user or another user starts an online dialogue with a second user via outgoing call processing, incoming call processing, a room, etc., the speaker type of the second user stored in association with the second user in step S305 may be displayed on the display 2081 of the first user terminal 20 on an outgoing call screen for making a call to the second user, an incoming call screen for receiving a call from the second user, a room screen before the dialogue starts, etc., and presented to the first user. This allows the first user to prepare a response that is appropriate for the speaker type of the second user before the start of the dialogue.

[0142] The impression analysis unit 1043 of the server 10 may execute a presentation step of presenting to the first user the impression features stored in association with the second user in the storage step before the dialogue between the first user and the second user ends. For example, while the first user or another user is having an online conversation with the second user, the impression features of the second user stored in association with the second user in step S305 may be displayed and presented to the first user on a conversation screen, a room screen, or the like displayed on the display 2081 of the first user terminal 20. Note that the impression features may be presented to any user, such as the second user, another administrator, or another user. This allows the first user to prepare a response that is appropriate for the impression of the second user during the conversation.

[0143] The impression analysis unit 1043 of the server 10 may perform a presentation step of presenting to the first user the speaker type stored in association with the second user in the storage step before the dialogue between the first user and the second user ends. For example, while the first user or another user is having an online conversation with the second user, the speaker type of the second user stored in association with the second user in step S305 may be displayed and presented to the first user on a conversation screen, room screen, or the like displayed on the display 2081 of the first user terminal 20. Note that the impression features may be presented to any user, such as the second user, another administrator, or another user. This allows the first user to prepare a response that suits the speaker type of the second user during the conversation.

[0144] In the impression calculation step, the impression analysis unit 1043 of the server 10 may execute a presentation step of presenting one or more dialogue features that have a large influence on the impression feature among the plurality of dialogue features. Specifically, the impression analysis unit 1043 of the server 10 applies a plurality of dialogue features as input data to the second impression evaluation model 1034, and when the second impression evaluation model 1034 outputs impression features according to the input data as output data, it may identify one or more dialogue features that have a significant influence on the output impression features, transmit them to the first user terminal 20, the second user terminal 30, or any other user terminal, and present them to the user. For example, the second impression evaluation model 1034 may output one or more dialogue features that have a large influence on the output impression features as output data, thereby enabling dialogue features that have a large influence on the impression features to be acquired at high speed.

[0145] <Modification of impression analysis process> The impression analysis process may be configured to identify the impression state of the first user who is an operator, rather than the second user who is a customer. Furthermore, the method may also include a step of accepting target impression features and target speaker type that the first user wants to give to other users, calculating dialogue features that the first user should improve, and presenting them to the first user. In other words, the method may also include a step of proposing a preferred speaking style to the first user. In this case, the processing content is the same except that the second user is replaced with the first user in steps S301 to S305 of the impression analysis processing, and therefore a description thereof will be omitted.

[0146] In step S306, the impression analysis unit 1043 of the server 10 executes a target receiving step of receiving a target speaker type that is to be given as a target by the first user to other users in a dialogue. Specifically, the first user operates the input device 206 of the first user terminal 20 or the like to access a predetermined web page provided by the server 10 and selects a target speaker type from a list of multiple speaker types. The control unit 204 of the first user terminal 20 identifies the selected target speaker type and transmits it to the server 10. The server 10 receives and accepts the target speaker type. The target speaker type is a speaker type related to a desired impression state that the first user wants to give to other users, and may be selected by the first user himself or herself, or may be selected by the first user's manager or the like depending on the first user's job or the like.

[0147] In step S306, the impression analysis unit 1043 of the server 10 executes a target receiving step of receiving a target impression feature quantity that is a target to be given by the first user to another user in a dialogue. Specifically, the impression analysis unit 1043 of the server 10 searches the speaker type items in the speaker type master 1022 based on the received target speaker type and acquires impression conditions. Based on the acquired impression conditions, the impression analysis unit 1043 of the server 10 identifies impression features included in the range of the impression conditions as target impression features and accepts them. The impression analysis unit 1043 of the server 10 may be configured to acquire and accept target impression features output by applying the target speaker type as input data to a learning model (not shown). Alternatively, the impression analysis unit 1043 may be configured to accept target impression features from the first user via the input device 206 of the first user terminal 20 or the like.

[0148] In step S306, the impression analysis unit 1043 of the server 10 executes an improvement step of calculating a dialogue feature to be improved by the first user based on the impression feature calculated in the impression calculation step and the target impression feature received in the target reception step. Specifically, the impression analysis unit 1043 of the server 10 identifies, based on the identified target impression feature, a dialogue feature for obtaining the target impression feature, and accepts the identified target dialogue feature. The impression analysis unit 1043 of the server 10 may be configured to acquire and accept a target dialogue feature by applying the target impression feature as input data to a learning model (not shown) or the like. Examples of the dialogue feature that the first user should improve include "faster speaking rate," "slower speaking rate," "higher intonation," "softer intonation," etc. The dialogue feature that the first user should improve may also be a target dialogue feature (target dialogue feature).

[0149] The impression analysis unit 1043 of the server 10 compares the dialogue feature calculated in step S304 with the target dialogue feature. The impression analysis unit 1043 of the server 10 calculates the difference between the dialogue feature and the target dialogue feature as the dialogue feature that the first user should improve. The impression analysis unit 1043 of the server 10 also compares the dialogue feature with the target dialogue feature, and identifies the dialogue feature with a large degree of deviation as the dialogue feature that the first user should improve. The impression analysis unit 1043 of the server 10 transmits the dialogue features to be improved by the first user to the first user terminal 20. The control unit 204 of the first user terminal 20 displays the received dialogue features to be improved on the display 2081 of the first user terminal 20, and presents them to the first user. For example, among dialogue features such as the speech rate, intonation, number of polite expressions, number of fillers, and number of grammatical utterances of the first user in a dialogue, the dialogue feature that the first user should improve is identified, and the first user is presented with the extent to which the speech rate, intonation, number of polite expressions, number of fillers, etc. should be improved. This allows the operator or the like to improve the impression they give to others by specifically improving their speaking style. The dialogue feature may be presented to the second user and other users.

[0150] This allows the impression analysis unit 1043 of the server 10 to execute an improvement step of calculating dialogue features that the first user should improve based on the speaker type calculated in the impression calculation step and the target speaker type accepted in the target acceptance step. In other words, the user can understand the dialogue features that should be improved according to the accepted target speaker type, and by improving their speaking style based on the dialogue features that should be improved, they can make the impression they give to others closer to the target speaker type.

[0151] <Topic definition processing> The topic definition process is a process in which a user registers and stores a topic related to a predetermined topic, which is associated with a plurality of keywords.

[0152] <Topic definition process overview> A user can define and memorize a new topic based on keywords such as multiple words, nouns, adjectives, etc. Furthermore, for an already memorized topic, keywords highly relevant to that topic are presented based on previously stored conversation information, and those keywords are added to the keywords associated with the topic and memorized, thereby expanding the keywords associated with the topic.

[0153] <Details of topic definition process> The topic definition process is explained in detail below.

[0154] The topic processing unit 1044 of the server 10 executes a keyword presentation step in which, based on the voice data stored in the voice storage step and the plurality of keywords received in the keyword reception step, one or more new keywords to be newly associated with the first topic are presented to the first user. Specifically, the first user executes the application program 2012 and runs a browser application by operating the input device 206 or the like of the first user terminal 20. In the browser application, the first user inputs a predetermined URL (Uniform Resource Locator) that specifies a predetermined web server provided by the server 10, thereby sending a request to the server 10 for a page for defining a topic.

[0155] The topic processing unit 1044 of the server 10 searches the speaker ID field of the speech segment table 1016 based on the first user ID 2011 included in the received request, and acquires the segment reading text. The topic processing unit 1044 of the server 10 extracts character strings such as nouns, adjectives, and keywords contained in the section reading text by performing processing such as morphological analysis on the section reading text. At this time, the importance of each character string may be calculated based on the frequency of appearance of the character string for each dialogue information and speech section information. TF-IDF is one example of a method for calculating the importance. The topic processing unit 1044 of the server 10 identifies a predetermined number of character strings with high importance as keyword candidates.

[0156] The topic processing unit 1044 of the server 10 may acquire topic IDs and keywords from the topic master 1023, and identify as keyword candidates character strings that are in a co-occurrence relationship with multiple keywords associated with each of multiple topic IDs in one or more pieces of dialogue information or speech section information, but are not associated with the topic ID. Note that the importance of each keyword and character string may be taken into consideration when calculating the co-occurrence relationship. When identifying keyword candidates, a predetermined number of character strings may be identified as keyword candidates, taking into consideration the importance calculated based on the frequency of appearance, etc.

[0157] The topic processing unit 1044 of the server 10 transmits the identified keyword candidates to the first user terminal 20. The control unit 204 of the first user terminal 20 displays the received keyword candidates on the display 2081 of the first user terminal 20, and presents them to the first user.

[0158] The topic processing unit 1044 of the server 10 executes a keyword receiving step of receiving one or more keywords from the first user. Specifically, the first user operates the input device 206 of the first user terminal 20 or the like to select a keyword to associate with a new topic from the keyword candidates displayed on the display 2081 of the first user terminal 20. The control unit 204 of the first user terminal 20 transmits to the server 10 one or more keyword candidates selected by the first user.

[0159] The keyword receiving step executes a step of receiving one or more keywords selected by the first user from among the plurality of new keywords presented to the first user in the keyword presenting step. Specifically, the topic processing unit 1044 of the server 10 receives and accepts one or more keyword candidates from the first user terminal 20.

[0160] The topic processing unit 1044 of the server 10 executes a topic storage step of storing one or more keywords received in the keyword receiving step in association with a first topic related to a predetermined topic. Specifically, the topic processing unit 1044 of the server 10 associates the received multiple keyword candidates with topic IDs and stores them in the topic master 1023. Note that the one or multiple keyword candidates selected by the first user may be associated with a topic ID already stored in the topic master 1023, or a new topic ID may be generated and associated with the newly generated topic ID. When storing a topic ID in association with a topic ID already stored in the topic master 1023, the first user performs a selection operation to select the topic ID to be associated by operating the input device 206 of the first user terminal 20 or the like.

[0161] <Topic analysis processing> Topic analysis processing is a process that analyzes dialogue information such as audio and video of online conversations conducted by multiple users, calculates the degree of relevance between the dialogue information and one or more topics, and associates and stores the topics with the dialogue information based on the degree of relevance.

[0162] <Outline of topic analysis process> The topic analysis process is a series of processes that, when an online conversation between users is detected, stores conversation information about the conversation, divides the audio data and video data contained in the conversation information into section data such as section audio data and section video data for each speech section, calculates the relevance of each section data with multiple topics, identifies the topic for each section data, and stores representative topics as label information for the conversation information.

[0163] <Details of topic analysis process> The topic analysis process will be described in detail below.

[0164] In step S511, online conversation between the user and the customer is started via the outgoing call processing, incoming call processing, room, etc., which have already been described.

[0165] In step S512, the topic processing unit 1044 of the server 10 executes a receiving step of receiving voice data related to the dialogue, and a voice storage step of storing the voice data received in the receiving step. Step S512 is the same as step S102 in the emotion analysis process, and therefore a description thereof will be omitted.

[0166] In step S513, the topic processing unit 1044 of the server 10 executes a voice extraction step of extracting a plurality of section voice data for each utterance section from the voice data received in the receiving step. Step S513 is the same as step S103 in the emotion analysis process, and therefore a description thereof will be omitted.

[0167] In step S513, the voice extracting step may include a step of extracting a plurality of section voice data for each utterance section from the voice data received in the receiving step before the dialogue ends. In other words, the voice extraction step may be configured to be executed at any timing during an online dialogue between a plurality of users.

[0168] In step S514, the topic processing unit 1044 of the server 10 executes a topic identification step of identifying a first topic related to a predetermined topic and associated with a plurality of keywords. Specifically, the topic processing unit 1044 of the server 10 refers to the topic master 1023, and acquires and specifies the topic ID and one or more keywords associated with the topic ID that have been registered in advance by the topic definition process.

[0169] The relevance calculation step executes a step of calculating the relevance of each of the plurality of topics identified in the topic identification step for each of the plurality of section voice data. In this disclosure, for simplicity's sake, we will mainly describe one first topic and one or more keywords associated with the first topic, but the topic is not limited to one and similar processing may be performed on multiple topics (second topic, third topic, etc.).

[0170] In step S514, the topic processing unit 1044 of the server 10 executes a relevance calculation step of calculating, for each of the plurality of section voice data, a first relevance indicating the relevance with the first topic identified in the topic identification step. Specifically, the topic processing unit 1044 of the server 10 calculates a first relevance indicating the degree of relevance with the first topic, based on the relevance between the speech section information acquired in S513 and the keywords associated with the first topic.

[0171] An example of a method for calculating the first relevance is described below. The topic processing unit 1044 of the server 10 creates a high-dimensional vector (topic vector) as a distributed representation (embedded representation) based on keywords associated with the first topic. The topic processing unit 1044 of the server 10 also performs processing such as morphological analysis on the section reading text included in the multiple pieces of speech segment information to extract character strings such as nouns, adjectives, and keywords included in the section reading text, and creates a high-dimensional vector (speech segment vector) as a distributed representation based on the extracted character strings. Note that a method called Word2vec is known as a method for creating distributed representations. The topic processing unit 1044 of the server 10 calculates the first relevance by calculating the cosine similarity between the topic vector and the speech segment vector. Note that the first relevance may be calculated using an algorithm for calculating the distance between any multidimensional vectors, such as Euclidean distance, Mahalanobis distance, Manhattan distance, Chebyshev distance, or Minkowski distance. The first relevance calculated in this way reflects the overall tendency of similarity between the multiple keywords associated with the first topic and the character strings included in the multiple pieces of speech segment information. As a result, the character strings included in the speech segment information are not judged as different words with the same meaning due to paraphrases or differences in spelling of the keywords included in the topic, and a higher relevance can be obtained for speech segment information that has a high semantic relevance to the keywords included in the first topic. In the present disclosure, the calculation of the first relevance indicating the relevance with the first topic has been described, but the calculation of the relevance between an arbitrary topic and speech segment information is similar.

[0172] The relevance calculation step may include a step of calculating a first relevance indicating the relevance with the first topic identified in the topic identification step for each section audio data included in the plurality of section audio data before the dialogue ends. That is, the configuration may be such that the calculation is performed at any timing during an online dialogue between multiple users, thereby making it possible to calculate the relevance of each topic with respect to the speech segment information in the dialogue up to that point, even during the online dialogue.

[0173] The relevance calculation step may be configured to assign a smaller weight to the relevance of a keyword that is included more frequently in the multiple section audio data extracted in the audio extraction step among the multiple keywords associated with the first topic, and calculate the degree of match taking into account the weighting of the multiple keywords associated with the first topic for each of the multiple section audio data as the first relevance indicating the relevance with the first topic. Specifically, when calculating the relevance, different weights may be assigned to the importance of each of the multiple keywords associated with the first topic. For example, for multiple pieces of speech segment information extracted for one piece of dialogue information, the importance and weight of a keyword that frequently appears in many pieces of speech segment information may be set to a smaller value than that of other keywords so that the degree of influence that the keyword has on the relevance is reduced. This makes it possible to prevent the relevance of a common keyword that frequently appears in many pieces of speech segment information from being overestimated. In the present disclosure, the calculation of the first relevance indicating the relevance with the first topic has been described, but the calculation of the relevance between an arbitrary topic and speech segment information may be performed in the same manner.

[0174] The relevance calculation step may be such that, among the multiple keywords associated with the first topic, the more frequently a keyword is included in multiple section audio data up to a predetermined number of sections chronologically before the target section audio data for which the first relevance is being calculated, the smaller the weight given to the relevance, and the degree of match taking into account the weighting with the multiple keywords associated with the first topic for each of the multiple section audio data may be calculated as the first relevance indicating the relevance with the first topic. For example, rather than for all of the plurality of pieces of speech section information extracted for one piece of dialogue information, for a plurality of pieces of speech section information up to a predetermined number of pieces in chronological order prior to the target section speech information to be calculated, the importance and weight of a keyword that frequently appears in many pieces of speech section information may be set to a smaller value than other keywords so that the degree of influence on the relevance is reduced. This makes it possible to more accurately calculate the relevance between the most recent speech section information and the topic even at any timing during the dialogue before the dialogue ends. In the present disclosure, the calculation of the first relevance indicating the relevance with the first topic has been described, but the calculation of the relevance between an arbitrary topic and speech segment information may be performed in the same manner.

[0175] The topic processing unit 1044 of the server 10 stores the relevance calculated for each of multiple topics for multiple speech section information extracted for one piece of dialogue information in the section ID, topic ID, and relevance items of a new record in the topic relevance table 1017, respectively, along with the section ID that identifies the speech section information, the topic ID that identifies the topic, and the calculated relevance.

[0176] In step S515, among one or more topics having a relevance level equal to or higher than a predetermined value in each piece of speech section information, the topic with the highest relevance level is identified as the topic related to the predetermined topic mentioned by the speech section information. Note that the topic does not necessarily have to be identified. The topic processing unit 1044 of the server 10 stores the topic ID of the identified topic in the speech section table 1016 in the topic ID field of the record identified by the section ID of the speech section information to be used for calculating the relevance level. As a result, the speech section information is stored in association with a topic with a high relevance level.

[0177] In step S516, the topic processing unit 1044 of the server 10 executes a label specifying step of specifying label information for the dialogue based on the relevance of each of the multiple topics calculated in the relevance calculation step. The topic processing unit 1044 of the server 10 executes a storage step of storing the label information specified in the label specifying step in association with the dialogue. Specifically, in step S515, the topic processing unit 1044 of the server 10 counts up the topic IDs stored for each of the multiple pieces of speech section information extracted for one piece of dialogue information, and identifies one or more topic IDs in descending order of the number of counted topic IDs as topics that characterize the piece of dialogue information. Note that one or more topic IDs for which the number of counted topic IDs is equal to or greater than a predetermined number may be identified as topics that characterize the piece of dialogue information. The topic processing unit 1044 of the server 10 identifies the topic name, label, etc. of the identified topic ID as label information. Note that a configuration may be adopted in which arbitrary label information is identified based on the identified topic ID by referring to a table (not shown) or the like. The identified label information and the dialogue ID of the dialogue information are stored in the label data and dialogue ID fields of a new record in the label table 1015. This allows the dialogue information and the topic that characterizes the dialogue information to be associated and stored as label information, which can be conveniently used when searching for dialogue information, etc.

[0178] <About the timing of topic analysis processing> Steps S513 to S516 of the topic analysis process may be configured to be executed after an online conversation between multiple users has ended. As a result, after the online conversation has ended and the content of the conversation has been determined, topics related to the conversation are identified and stored in association with the conversation information.

[0179] The topic analysis process may be configured to be executed after the start of an online conversation between multiple users and before the conversation ends. That is, the steps may be executed at any timing during an online dialogue between multiple users. Also, steps S513 to S516 may be executed periodically in real time during the online dialogue. As a result, even during the online dialogue, a topic according to the dialogue up to that point may be identified and stored in association with the dialogue information. This allows users to check topics being mentioned by other users participating in the online dialogue in real time during the online dialogue, and also allows users to organize and manage dialogue information based on the latest topics.

[0180] <Topic presentation process> The topic presentation process visually visualizes and presents to users dialogue information, such as audio and video, of online dialogues between multiple users, and also presents topics associated with the dialogue information to the users. Users can check the dialogue information and the topics associated with the dialogue information at a glance, and intuitively grasp the outline of the dialogue content.

[0181] <Outline of topic presentation process> This is a series of processes that accepts from the user the dialogue information to be presented, acquires the dialogue information, acquires section data and topics for each section data, analyzes the dialogue information, presents the user with a speech graph that allows visual confirmation of the speech situation for each speaker, and overlays the topics for each speech section on the speech graph and presents them to the user.

[0182] <Details of topic presentation process> The topic presentation process will be described in detail below.

[0183] In step S521, the first user selects the dialogue information whose topic he or she wishes to check. Specifically, the first user executes the application program 2012 and runs a browser application by operating the input device 206 or the like of the first user terminal 20. In the browser application, the first user inputs a predetermined URL (Uniform Resource Locator) that specifies a predetermined web server provided by the server 10, thereby sending a request to the server 10 for a page to present a topic. The topic processing unit 1044 of the server 10 searches the user ID item in the dialogue table 1014 based on the first user ID 2011 included in the received request, and acquires a dialogue ID. The topic processing unit 1044 of the server 10 transmits the acquired one or more dialogue IDs to the first user terminal 20. The control unit 204 of the first user terminal 20 displays the received one or more dialogue IDs on the display 2081 of the first user terminal 20, thereby presenting them to the first user. The first user selects a predetermined dialogue ID from the presented dialogue IDs by operating the input device 206 of the first user terminal 20. The control unit 204 of the first user terminal 20 transmits the selected predetermined dialogue ID to the server 10. The server 10 receives and accepts the dialogue ID.

[0184] Note that, when the first user is engaged in a dialogue using the online dialogue service according to the present disclosure, the dialogue information during the dialogue may be selected. In other words, the topic presentation process may be executed on the dialogue screen displayed on the display 2081 of the first user terminal 20 during the dialogue.

[0185] In step S522, the topic processing unit 1044 of the server 10 searches the dialogue ID item in the dialogue table 1014 based on the received dialogue ID, and obtains dialogue information such as the user ID, customer ID, dialogue category, sending / receiving type, audio data, and video data.

[0186] In step S523, the topic processing unit 1044 of the server 10 searches the dialogue ID item in the speech segment table 1016 based on the received dialogue ID, and acquires the items of segment ID, start date and time, end date and time, and topic ID. Based on the acquired segment ID, the topic processing unit 1044 of the server 10 searches the segment ID item in the topic relevance table 1017, and acquires the topic ID and relevance. That is, the topic processing unit 1044 of the server 10 acquires a plurality of pieces of speech section information associated with a dialogue ID, and a topic ID and a degree of association for each piece of speech section information.

[0187] In step S524, the topic processing unit 1044 of the server 10 outputs a speech graph indicating the time series transition of the speaker's speech situation based on the dialogue information acquired in step S522, and transmits the speech graph to the first user terminal 20. The control unit 204 of the first user terminal 20 displays the received speech graph on the display 2081 of the first user terminal 20, and presents it to the first user. An example screen 70 including the speech graph presented to the first user is shown in FIG. 20. The voice graph may be presented to any user, such as the second user, other administrators, or other users.

[0188] The voice graph is a graph with the horizontal axis representing the conversation time, the vertical axis (upper) representing the voice output volume of the first user, and the vertical axis (lower) representing the voice output volume of the second user, where the solid line L1 represents the voice of the first user and the dashed line L2 represents the voice of the second user. Looking at the solid line L1 and the dashed line L2, we can see that basically, while the first user is making a sound (talking), the second user is not making a sound (listening silently), and while the second user is making a sound (talking), the first user is not making a sound (listening silently). Here, the points indicated by Z3 are when both users are making sounds at the same time (overlapping), and it is possible that the first user started speaking before the second user finished speaking. The points indicated by Z1 and Z2 are times when neither user is making a sound (time of silence). The points indicated by P1 and P2 are times when a specific keyword appeared.

[0189] In step S525, the topic processing unit 1044 of the server 10 executes a section group identification step to identify a first section group that includes, from among the multiple section audio data, one or more section audio data whose first relevance calculated in the relevance calculation step is equal to or greater than a predetermined value. Specifically, when the topic processing unit 1044 of the server 10 determines in the topic analysis process that one or more pieces of speech section information, each of which has a first relevance degree equal to or greater than a predetermined value, mentions a topic related to a first topic, the topic processing unit 1044 identifies one or more pieces of speech section information including the one or more pieces of speech section information as a first section group. For example, if the association of multiple pieces of chronologically consecutive speech section information with topics is as follows: Section 1: Topic A, Section 2: Topic A, Section 3: No Topic, Section 4: Topic A, Section 5: No Topic, Section 6: Topic B, Section 7: Topic B, and Section 8: Topic B, the topic processing unit 1044 identifies Sections 1 to 4 as a section group related to Topic A, and Sections 6 to 8 as a section group related to Topic B. Even if a section of Topic A includes a speech section associated with another topic, such as Section 3, if Sections 1 to 4 as a whole are considered to mention the topic of Topic A, the topic processing unit 1044 may identify Sections 1 to 4 collectively as a section group related to Topic A.

[0190] In the present disclosure, the first section group is identified, but one or more section audio data related to a first topic on a predetermined topic may be identified from among a plurality of section audio data. Also, the first section group may be identified by selecting one or more section audio data, or the first section group, through an input operation by the first user or the second user.

[0191] In step S525, the topic processing unit 1044 of the server 10 executes a presentation step of associating the first segment group identified in the segment group identification step with the first topic and presenting the first segment group to the first user or the second user. The presentation step executes a step of presenting the first segment group identified in the segment group identification step on the same time series axis as the audio graph in an audio graph showing the time series progression of the speaker's speech situation, which is obtained by analyzing the audio data received in the reception step, and associating the first topic with the first segment group and presenting the first topic to the first user or the second user. Specifically, in the speech graph of FIG. 20, the topic processing unit 1044 of the server 10 presents a first segment group T1 associated with a first topic, a second segment group T2 associated with a second topic, and a third segment group T3 associated with a third topic as drawing objects superimposed on the speech graph. For example, the first segment group T1, the second segment group T2, and the third segment group T3 may be configured to be drawn as drawing objects in different colors assigned to each topic. This allows the first user to associate the segment groups with related topics and visually recognize them superimposed on the speech graph. This allows the first user to visually see at a glance which parts of the speech graph are being discussed about which topics. In addition, the topic processing unit 1044 of the server 10 may be configured to present the first section group identified in the section group identification step to any user, such as an administrator other than the first user and the second user, or another user.

[0192] In step S525, the section group identification step may include a step of calculating a moving average based on the first relevance calculated for each of multiple section audio data arranged in chronological order, and a step of identifying section audio data whose calculated moving average is equal to or greater than a predetermined value as a first section group. Specifically, when identifying a segment group, the topic processing unit 1044 of the server 10 arranges the speech segment information acquired from the topic relevance table in chronological order based on the start date and time of the speech segment information, etc. The topic processing unit 1044 of the server 10 calculates a moving average of the most recent N relevance degrees for a given speech segment information, where N is an arbitrary integer. The calculated moving average is regarded as a new relevance degree for the given speech segment information, and speech segment information whose relevance degree is equal to or greater than a predetermined value is identified as a first segment group associated with the first topic. In this disclosure, for simplicity's sake, the moving average for the relevance of one first topic has been described, but the topic is not limited to one, and similar processing may be performed for multiple topics. This makes it possible to identify groups of sections that mention a topic by smoothing the topic relevance, even when topics with high relevance for each utterance section change frequently. This makes it easier for users to check what topics a speaker is talking about in online dialogue services.

[0193] In step S525, the section group identification step may include a step of identifying, as a first section group, a plurality of consecutive section audio data among a plurality of section audio data arranged in chronological order, the plurality of consecutive section audio data having a calculated first relevance greater than or equal to a predetermined value. Specifically, when identifying the section group, the topic processing unit 1044 of the server 10 arranges the speech section information acquired from the topic relevance table in chronological order based on the start date and time of the speech section information, etc. The topic processing unit 1044 of the server 10 identifies multiple consecutive pieces of speech section information whose relevance is equal to or greater than a predetermined value as a first section group associated with the first topic. In this disclosure, for simplicity's sake, the moving average for the relevance of one first topic has been described, but the topic is not limited to one, and similar processing may be performed for multiple topics. This allows for the identification of consecutive sections of speech data that are highly relevant to a specific topic as a group of sections that mention that topic, making it easier for users to confirm what topics speakers are talking about in online dialogue services.

[0194] In step S525, the topic processing unit 1044 of the server 10 executes a summarization step of generating a summary text summarizing text information included in one or more section audio data based on one or more section audio data among the plurality of section audio data and the first topic identified in the topic identification step. The summarization step executes a step of generating a summary text summarizing text information included in one or more section audio data by extracting only parts of the text information included in the one or more section audio data that are highly relevant to the first topic identified in the topic identification step.

[0195] In step S525, the summarization step executes a step of generating a summary text by applying the text information contained in one or more sections of audio data and a plurality of keywords associated with the first topic as input data to a learning model. Specifically, section data including at least one of section audio data, section video data, and section reading text, and a plurality of keywords associated with the topic of the section data, are applied to summary model 1035 as input data, and summary text, which is text information that summarizes the text information included in the section data, is obtained as output data. This makes it possible to extract only parts of the text information included in the section data that are particularly highly relevant to the topic, and to obtain summary text that summarizes the text information included in the section data.

[0196] In step S525, the summarization step executes a step of generating a summary text that summarizes the text information contained in one or more section audio data based on one or more section audio data included in the first section group identified in the section group identification step and the first topic identified in the topic identification step. Specifically, one or more interval data included in the interval group and multiple keywords associated with the topic of the interval group are used as input data and applied to summary model 1035, and summary text, which is text information that summarizes the text information included in the interval group, is obtained as output data. This makes it possible to extract parts of the text information included in the interval data that are particularly highly relevant to the topic, and to obtain summary text that summarizes the text information included in the interval data.

[0197] In step S525, the topic processing unit 1044 of the server 10 executes a presentation step of presenting the summary text generated in the summarization step in association with one or more pieces of section speech data. In step S525, the topic processing unit 1044 of the server 10 executes a presenting step of presenting the summary text generated in the summarizing step in association with the first segment group identified in the segment group identifying step. 20, the topic processing unit 1044 of the server 10 presents summary text 701 related to the first topic of the first segment group T1 in association with the first segment group T1. Note that the topic processing unit 1044 of the server 10 may present summary text 701 in association with any one or more speech segments instead of a segment group. In addition, the topic processing unit 1044 of the server 10 may be configured to present the first section group identified in the section group identification step to any user, such as the first user, the second user, other administrators, or other users.

[0198] <Learning process> The learning processes of the emotion evaluation model 1031, impression evaluation model 1032, first impression evaluation model 1033, and second impression evaluation model 1034 will be described below.

[0199] <Learning process of emotion evaluation model 1031> The learning process of the emotion evaluation model 1031 is a process of learning the learning parameters of the deep neural network included in the emotion evaluation model 1031 by deep learning.

[0200] <Overview of the learning process of emotion evaluation model 1031> The learning process of the emotion evaluation model 1031 is a process in which the learning parameters of the deep neural network included in the emotion evaluation model 1031 are learned by deep learning using the interval audio data, interval video data, and interval read-aloud text as input data (input vectors) so that the emotion feature amount, that is, the emotion vector or emotion scalar, becomes the output data (teaching data). Any of the section audio data, section video data, and section read-aloud text may be omitted from the input data of the emotion evaluation model 1031 .

[0201] <Details of the learning process of emotion evaluation model 1031> The learning unit 1051 of the server 10 creates learning data using section audio data, section video data, section read-aloud text, etc. as input data (input vectors) so that predetermined emotion features become output data (teaching data). The learning unit 1051 of the server 10 creates data sets such as training data, test data, and verification data for training the deep neural network of the emotion evaluation model 1031 based on the learning data. The learning unit 1051 of the server 10 learns the learning parameters of the deep neural network included in the emotion evaluation model 1031 by deep learning based on the created data set.

[0202] <Learning process of impression evaluation model 1032> The learning process of the impression evaluation model 1032 is a process of learning the learning parameters of the deep neural network included in the impression evaluation model 1032 by deep learning.

[0203] <Overview of the learning process for impression evaluation model 1032> The learning process of the impression evaluation model 1032 is a process in which the learning parameters of the deep neural network included in the impression evaluation model 1032 are learned by deep learning using the section audio data, section video data, and section reading text as input data (input vectors) so that the impression features become output data (teaching data). Any of the section audio data, section video data, and section read-aloud text may be omitted from the input data of the impression evaluation model 1032 .

[0204] <Details of the learning process for impression evaluation model 1032> The learning unit 1051 of the server 10 creates learning data using section audio data, section video data, section read-aloud text, etc. as input data (input vectors) so that predetermined impression features become output data (teaching data). The learning unit 1051 of the server 10 creates data sets such as training data, test data, and verification data for training the deep neural network of the impression evaluation model 1032 based on the learning data. The learning unit 1051 of the server 10 learns the learning parameters of the deep neural network included in the impression evaluation model 1032 by deep learning based on the created data set.

[0205] <Learning process of first impression evaluation model 1033> The learning process of the first impression evaluation model 1033 is a process of learning the learning parameters of the deep neural network included in the first impression evaluation model 1033 by deep learning.

[0206] <Outline of the learning process of the first impression evaluation model 1033> The learning process of the first impression evaluation model 1033 is a process in which the learning parameters of the deep neural network included in the first impression evaluation model 1033 are learned by deep learning using the section audio data, section video data, and section reading text as input data (input vectors) and the dialogue features as output data (teaching data). Any of the section audio data, section video data, and section read-aloud text may be omitted from the input data of the first impression evaluation model 1033 .

[0207] <Details of the learning process of the first impression evaluation model 1033> The learning unit 1051 of the server 10 creates learning data using section voice data, section video data, section read-aloud text, and the like as input data (input vectors) so that predetermined dialogue features become output data (teaching data). The learning unit 1051 of the server 10 creates data sets such as training data, test data, and verification data for training the deep neural network of the first impression evaluation model 1033 based on the learning data. The learning unit 1051 of the server 10 learns the learning parameters of the deep neural network included in the first impression evaluation model 1033 by deep learning based on the created data set.

[0208] <Learning process of second impression evaluation model 1034> The learning process of the second impression evaluation model 1034 is a process of learning the learning parameters of the deep neural network included in the second impression evaluation model 1034 by deep learning.

[0209] <Outline of the learning process of the second impression evaluation model 1034> The learning process of the second impression evaluation model 1034 is a process of learning the learning parameters of the deep neural network included in the second impression evaluation model 1034 by deep learning so that the dialogue features are input data (input vectors) and the impression features are output data (teacher data).

[0210] <Details of the learning process of the second impression evaluation model 1034> The learning unit 1051 of the server 10 creates learning data using dialogue features and the like as input data (input vectors) so that predetermined impression features become output data (teaching data). The learning unit 1051 of the server 10 creates data sets such as training data, test data, and verification data for training the deep neural network of the second impression evaluation model 1034 based on the learning data. The learning unit 1051 of the server 10 learns the learning parameters of the deep neural network included in the second impression evaluation model 1034 by deep learning based on the created data set.

[0211] <Details of the training process for summary model 1035> The learning unit 1051 of the server 10 uses section data including at least one of section audio data, section video data, and section reading text, and multiple keywords associated with topics related to a predetermined topic, as input data (input vectors), and creates learning data such that summary text, which is text information that summarizes the text information included in the section data, becomes output data (teaching data). The learning unit 1051 of the server 10 creates data sets such as training data, test data, and validation data for training the deep neural network of the summary model 1035 based on the learning data. The learning unit 1051 of the server 10 uses deep learning to learn the learning parameters of the deep neural network included in the summary model 1035 based on the created dataset.

[0212] <Basic computer hardware configuration> 21 is a block diagram showing the basic hardware configuration of a computer 90. The computer 90 includes at least a processor 901, a main memory device 902, an auxiliary memory device 903, and a communication IF 991 (interface), which are electrically connected to one another by a communication bus 921.

[0213] The processor 901 is hardware for executing an instruction set written in a program, and is composed of an arithmetic unit, registers, peripheral circuits, and the like.

[0214] The main storage device 902 is for temporarily storing programs, data to be processed by the programs, etc. For example, it is a volatile memory such as a DRAM (Dynamic Random Access Memory).

[0215] The auxiliary storage device 903 is a storage device for saving data and programs, such as a flash memory, a hard disk drive (HDD), a magneto-optical disk, a CD-ROM, a DVD-ROM, or a semiconductor memory.

[0216] The communication IF 991 is an interface for inputting and outputting signals for communicating with other computers via a network using wired or wireless communication standards. The network is composed of the Internet, a LAN, various mobile communication systems constructed by wireless base stations, etc. For example, the network includes 3G, 4G, and 5G mobile communication systems, LTE (Long Term Evolution), and wireless networks (e.g., Wi-Fi (registered trademark)) that can connect to the Internet via a predetermined access point. In the case of a wireless connection, communication protocols include, for example, Z-Wave (registered trademark), ZigBee (registered trademark), and Bluetooth (registered trademark). In the case of a wired connection, the network also includes a direct connection using a USB (Universal Serial Bus) cable, etc.

[0217] It should be noted that the computer 90 can be virtually realized by distributing all or part of each hardware configuration across multiple computers 90 and interconnecting them via a network. In this way, the computer 90 is a concept that includes not only a computer 90 housed in a single housing or case, but also a virtualized computer system.

[0218] <Basic functional configuration of computer 90> The following describes the functional configuration of a computer realized by the basic hardware configuration (FIG. 21) of the computer 90. The computer includes at least the functional units of a control unit, a storage unit, and a communication unit.

[0219] The functional units of the computer 90 can also be realized by distributing all or part of the functional units among multiple computers 90 interconnected via a network. The computer 90 is a concept that includes not only a single computer 90 but also a virtualized computer system.

[0220] The control unit is realized by the processor 901 reading out various programs stored in the auxiliary storage device 903, expanding them in the main storage device 902, and executing processing in accordance with the programs. The control unit can realize functional units that perform various types of information processing depending on the type of program. In this way, the computer is realized as an information processing device that performs information processing.

[0221] The storage unit is realized by a main storage device 902 and an auxiliary storage device 903. The storage unit stores data, various programs, and various databases. Furthermore, the processor 901 can allocate a storage area corresponding to the storage unit in the main storage device 902 or the auxiliary storage device 903 in accordance with the programs. Furthermore, the control unit can cause the processor 901 to execute processes for adding, updating, and deleting data stored in the storage unit in accordance with the various programs.

[0222] A database refers to a relational database, which manages data sets called masters and tables in a tabular format structurally defined by rows and columns, by relating them to each other. In a database, a table is called a table, a master, a column in a table is called a column, and a row in a table is called a record. In a relational database, relationships between tables and masters can be set and associated. Typically, each table and each master has a column set as a primary key to uniquely identify a record, but setting a primary key to a column is not essential. The control unit can cause the processor 901 to add, delete, or update records in specific tables and masters stored in the storage unit according to various programs.

[0223] Note that the databases and masters in this disclosure may include any data structure in which information is structurally defined (such as a list, dictionary, associative array, or object). The data structure also includes data that can be considered as a data structure by combining data with functions, classes, methods, etc. written in any programming language.

[0224] The communication unit is realized by the communication IF 991. The communication unit realizes a function of communicating with other computers 90 via a network. The communication unit can receive information transmitted from other computers 90 and input the information to the control unit. The control unit can cause the processor 901 to execute information processing on the received information in accordance with various programs. In addition, the communication unit can transmit information output from the control unit to other computers 90.

[0225] <Additional Notes> The matters described in the above embodiments will be supplemented below.

[0226] (Appendix 1) a processor and a storage unit, and a program for causing a computer to process information relating to a dialogue between a first user and a second user, the program causing the processor to execute: a reception step (S102) of receiving voice data relating to the dialogue; a voice extraction step (S103) of extracting multiple section voice data for each speech section from the voice data received in the reception step; an emotion calculation step (S104) of calculating multiple emotion features relating to the emotional state of a speaker in the section voice data, corresponding to each of the multiple section voice data extracted in the voice extraction step; a label identification step (S105) of identifying label information for the dialogue based on the multiple emotion features calculated in the emotion calculation step; and a storage step (S106) of storing the label information identified in the label identification step in association with the dialogue. This allows for managing dialogue information between speakers in a dialogue based on the emotional states of the speakers.

[0227] (Appendix 2) The program described in Appendix 1, wherein the emotion calculation step (S104) includes a step of calculating an emotion vector indicating the intensity of multidimensional emotions corresponding to each of the multiple section audio data extracted in the audio extraction step, and a step of calculating an emotion scalar indicating the intensity of one-dimensional emotions corresponding to each of the multiple section audio data extracted in the audio extraction step based on the calculated emotion vector, and the label identification step (S105) is a step of identifying label information for the dialogue based on the multiple emotion scalars calculated in the emotion calculation step. This allows label information to be identified based on a one-dimensional emotion scalar that integrates elements of emotion vectors, such as anger, disgust, fear, happiness, sadness, and surprise, and allows dialogue information between speakers to be managed.

[0228] (Appendix 3) The program described in Appendix 1, wherein the emotion calculation step (S104) is a step of calculating an emotion vector indicating the intensity of multidimensional emotions corresponding to each of the multiple section audio data extracted in the audio extraction step, and the label identification step (S105) is a step of identifying label information for the dialogue based on the multiple emotion vectors calculated in the emotion calculation step. This allows label information to be identified based on multidimensional emotion vectors such as anger, disgust, fear, happiness, sadness, and surprise, which are elements of the emotion vector, and allows conversation information between speakers to be managed.

[0229] (Appendix 4) The program according to claim 1, wherein the label identification step (S105) is a step of identifying label information for the dialogue based on the number of emotion features that are above or below a predetermined threshold, among the plurality of emotion features calculated in the emotion calculation step. This makes it possible to estimate the emotional state of the speaker, and to manage dialogue information between speakers in a dialogue based on the emotional state of the speaker.

[0230] (Appendix 5) The program according to claim 1, wherein the label identification step (S105) is a step of identifying label information for the dialogue based on the proportion of emotion features that are above or below a predetermined threshold among the plurality of emotion features calculated in the emotion calculation step. This makes it possible to estimate the emotional state of the speaker, and to manage dialogue information between speakers in a dialogue based on the emotional state of the speaker.

[0231] (Appendix 6) The program according to claim 1, wherein the label specifying step (S105) is a step of specifying label information for the dialogue based on statistical values ​​of the plurality of emotion features calculated in the emotion calculation step. This makes it possible to estimate the emotional state of the speaker, and to manage dialogue information between speakers in a dialogue based on the emotional state of the speaker.

[0232] (Appendix 7) The program according to claim 1, wherein the label specifying step (S105) is a step of specifying label information for the dialogue based on time-series changes in the plurality of emotion features calculated in the emotion calculation step. This makes it possible to manage dialogue information between speakers based on time-series changes in the emotional states of the speakers.

[0233] (Appendix 8) The program according to Appendix 7, wherein the label identification step (S105) includes a step of performing regression analysis on time-series changes in the plurality of emotion features calculated in the emotion calculation step, and a step of identifying label information for the dialogue based on regression coefficients obtained as a result of the regression analysis. This makes it possible to manage dialogue information between speakers based on time-series changes in the emotional states of the speakers.

[0234] (Appendix 9) The program causes a processor to execute a step (S105) of identifying a first emotion group which is a set of a plurality of emotion features corresponding to a plurality of chronologically consecutive section audio data extracted in the audio extraction step, and a step (S105) of identifying a second emotion group which is a set of a plurality of emotion features corresponding to a plurality of chronologically consecutive section audio data extracted in the audio extraction step, wherein the label identification step (S105) includes a step of identifying first label information for the dialogue based on the plurality of emotion features included in the first emotion group, and a step of identifying second label information for the dialogue based on the plurality of emotion features included in the second emotion group, and the storage step (S106) is a step of storing the first label information or the second label information identified in the label identification step in association with the dialogue. This allows multiple pieces of label information to be identified based on the emotional states of multiple speakers involved in one conversation, making it possible to manage conversation information between speakers more accurately.

[0235] (Appendix 10) The program causes a processor to execute a label presenting step (S105) of presenting the first label information and the second label information to a first user, and a selection receiving step (S105) of receiving, from the first user, a selection instruction to select at least one of the first label information and the second label information presented in the label presenting step, and a storage step (S106) of storing at least one of the first label information and the second label information in association with the dialogue based on the selection instruction received from the first user in the selection receiving step. This allows multiple label information to be identified based on the emotional states of multiple speakers involved in a single dialogue, and presented to the user, allowing for more accurate management of dialogue information based on the label information selected by the user.

[0236] (Appendix 11) The label identification step (S105) is a step of identifying label information for the dialogue based on the plurality of emotion features calculated in the emotion calculation step and the user attributes of the first user or the second user who spoke the section audio data corresponding to the plurality of emotion features. This makes it possible to identify more appropriate label information that takes into account the user attributes of each user, and to more appropriately manage dialogue information between speakers in a dialogue based on the emotional states of the speakers.

[0237] (Appendix 12) The program according to Appendix 1, wherein the label identification step (S105) is a step of identifying label information for the dialogue based on the plurality of emotion features corresponding to the section audio data related to the speech of the second user calculated in the emotion calculation step, without taking into consideration the plurality of emotion features corresponding to the section audio data related to the speech of the first user. This allows the dialogue information between speakers in a dialogue to be managed based only on the emotional state of the speaker related to the second user. For example, the dialogue information can be managed without considering the emotional state of the speaker of the first user.

[0238] (Appendix 13) 13. The program of claim 12, wherein the first user is a host user who is an organizer of the conversation, and the second user is not a host user. This allows the dialogue information between speakers in a dialogue to be managed based on the emotional state of the second user who is the dialogue partner, without taking into consideration the emotional state of the host user who is the organizer of the dialogue.

[0239] (Appendix 14) 13. The program of claim 12, wherein the second user is a host user who is an organizer of the conversation, and the first user is not a host user. This allows the dialogue information between speakers in a dialogue to be managed based on the emotional state of the host user who is the organizer of the dialogue, without taking into consideration the emotional state of the second user who is the dialogue partner.

[0240] (Appendix 15) The program according to claim 1, wherein the emotion calculation step (S104) calculates emotion features as output data by applying the section audio data extracted in the audio extraction step as input data to a learning model. This allows for managing dialogue information between speakers in a dialogue based on the emotional states of the speakers.

[0241] (Appendix 16) An information processing device comprising a processor and a storage unit, wherein the processor executes a program according to any one of Supplementary Notes 1 to 15. This allows for managing dialogue information between speakers in a dialogue based on the emotional states of the speakers.

[0242] (Appendix 17) An information processing system including an information processing device having a processor and a storage unit, wherein the processor executes a program according to any one of Supplementary Notes 1 to 15. This allows for managing dialogue information between speakers in a dialogue based on the emotional states of the speakers.

[0243] (Appendix 18) An information processing method executed by a computer having a processor and a memory unit, the information processing method causing the computer to execute a program according to any one of Supplementary Notes 1 to 15. This allows for managing dialogue information between speakers in a dialogue based on the emotional states of the speakers.

[0244] (Appendix 19) An information processing terminal comprising a processor and a display device, wherein the processor is capable of displaying label information identified by the label identification step executed in the information processing device according to Supplementary Note 16 on the display device. This allows the user to check the dialogue information between speakers in a dialogue as label information based on the emotional states of the speakers. [Explanation of symbols]

[0245] 1 System, 10 Server, 101 Memory unit, 104 Control unit, 106 Input device, 108 Output device, 20 First user terminal, 201 Memory unit, 204 Control unit, 206 Input device, 208 Output device, 30 Second user terminal, 301 Memory unit, 304 Control unit, 306 Input device, 308 Output device, 50 CRM system, 501 Memory unit, 504 Control unit, 506 Input device, 508 Output device, 60 Voice server (PBX), 601 Memory unit, 604 Control unit, 606 Input device, 608 Output device

Claims

1. A program for causing a computer to process information regarding a dialogue between a first user and a second user, the program comprising: a processor; and a storage unit, The program causes the processor to: a receiving step of receiving voice data relating to the dialogue; a data input step of inputting data corresponding to at least a part of the voice data received in the receiving step into a learning model as input data; a data output step of outputting an evaluation result regarding a speaker's emotion or impression according to the input data from the learning model; A program that executes the following.

2. The input data includes voice data or text data regarding the content of a user's utterance in video data. The program according to claim 1.

3. The data output step is a step of outputting the evaluation result including label information indicating a speaker's emotion or impression according to the input data from the learning model. The program according to claim 1.

4. A program comprising a processor and a memory unit, the program causing a computer to process information relating to a dialogue between a first user and a second user, The program causes the processor to: a receiving step of receiving voice data relating to the dialogue; a data input step of inputting data corresponding to at least a part of the voice data received in the receiving step into a learning model as input data; a label acquisition step of acquiring label information indicating a speaker's emotion or impression according to the input data based on an output result from the learning model; A program that executes the following.

5. The learning model is trained to be able to output a predetermined feature quantity in response to inputting data corresponding to at least a portion of the voice data. The program according to any one of claims 1 to 4.

6. The program causes the processor to: executing a presentation step of presenting the evaluation result to a user; The program according to any one of claims 1 to 3.

7. The program causes the processor to: executing a presentation step of presenting the label information to a user; The program according to claim 4.

8. A method executed on an information processing device having a processor and a memory unit, wherein the processor executes all of the steps executed in an invention relating to any one of claims 1 to 4.

9. An information processing device comprising a processor and a memory unit, wherein the processor executes all of the steps executed in an invention relating to any one of claims 1 to 4.

10. A system comprising means for executing all steps performed in an invention according to any one of claims 1 to 4.