Information processing system, information processing device, information processing method, and program

The information processing system analyzes customer emotions from voice data by extracting emotional values and patterns, enhancing operator performance and reducing stress through accurate emotional analysis and context understanding.

JP7805427B2Active Publication Date: 2026-01-23MITSUBISHI ELECTRIC DIGITAL INNOVATION CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024192951
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-11-01
Publication Date
2026-01-23
Estimated Expiration
2043-06-02

AI Technical Summary

Technical Problem

Existing call center technologies fail to accurately analyze customer emotions from voice data by understanding the content and context of conversations, necessitating operators to adjust their speaking style based on incomplete emotional evaluations.

Method used

An information processing system that extracts emotional values and frequently occurring words from voice data, using a learning model to predict and estimate patterns in conversations, enabling accurate emotional analysis and context understanding.

Benefits of technology

Enables effective emotional analysis and context understanding from voice data, improving operator skills and reducing stress by providing feedback on speaking style adjustments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007805427000001
    Figure 0007805427000001
  • Figure 0007805427000002
    Figure 0007805427000002
  • Figure 0007805427000003
    Figure 0007805427000003
Patent Text Reader

Abstract

To provide an information processing system capable of performing sentiment analysis after comprehending the content and context of conversation from the text data based on voice data.SOLUTION: The information processing system includes a prediction unit that predicts an evaluation value by inputting an emotion value into a generated learning model based on the correspondence between multiple emotion values corresponding to text data for each speaker that represents the utterance contents of the speaker and representing the emotions of a speaker and the evaluation value for the conversation corresponding to the text data.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing system, an information processing device, an information processing method, and a program. [Background technology]

[0002] In call center operations, in addition to improving customer satisfaction, it is desirable to improve the skills of operators and reduce operator stress. For example, Patent Document 1 discloses an information provision system that includes an analysis unit that uses voice data of a call between a customer and an operator that has been quantified based on emotions to classify the call into one of the evaluation indicators of the customer's emotions, and a provision unit that provides the results of the classification. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-154230 Summary of the Invention [Problem to be solved by the invention]

[0004] Typically, call center operations require operators to appropriately change intonation, words, and other aspects of their speech depending on the content of the conversation and the customer. For this reason, it is necessary to analyze emotions by understanding the content and context of the conversation from text data based on the audio data of the conversation, for example. However, the technology described in Patent Document 1 is based on the premise that text data based on voice data is not used to accurately evaluate customer emotions. Therefore, the operator needs to appropriately change the speaking style, such as intonation and words, depending on the content of the conversation and the customer. This has led to the problem that it is not possible to analyze emotions by understanding the content and context of the conversation from text data based on the voice data of the conversation.

[0005] One aspect of the present invention has been made in consideration of the above points, and aims to provide an information processing system, an information processing device, an information processing method, and a program that can analyze emotions after understanding the content and context of a conversation from text data based on voice data. [Means for solving the problem]

[0006] The present invention has been made to solve the above-mentioned problems, and one aspect of the present invention is an information processing system including a prediction unit that corresponds to text data for each speaker that represents the content of an utterance by the speaker, and that predicts the evaluation value by inputting the emotional values ​​that represent the emotions of the speaker into a learning model that is generated based on a correspondence between the emotional values ​​and an evaluation value for a conversation that corresponds to the text data.

[0007] Another aspect of the present invention is an information processing system comprising: a frequent word extraction unit that extracts one or more frequently occurring words included in text data for each speaker that represents the content of an utterance by the speaker; and an estimation unit that estimates which of a plurality of types of pattern the speaker belongs to based on the frequently occurring words and a plurality of emotion values ​​that represent the emotions of the speaker, wherein the emotion values ​​are either statistics of evaluation values ​​that represent each emotion over a predetermined period from the start of a conversation, statistics of evaluation values ​​that represent each emotion over a predetermined period going back a certain period from the end of the conversation, or statistics of evaluation values ​​that represent each emotion over a predetermined period that has elapsed since a specific keyword was uttered in the conversation, and the estimation unit estimates which of a plurality of types of pattern the speaker belongs to using the emotion values, numerical values ​​that represent the frequently occurring words, and weights for each of the plurality of types of patterns.

[0008] Another aspect of the present invention is an information processing device that includes a prediction unit that predicts an evaluation value by inputting a plurality of emotional values ​​that correspond to text data for each speaker that represents the content of an utterance by the speaker and that represent the emotions of the speaker into a learning model that is generated based on a correspondence between the emotional values ​​and an evaluation value for a conversation that corresponds to the text data.

[0009] Another aspect of the present invention is an information processing device comprising: a frequent word extraction unit that extracts one or more frequently occurring words included in text data for each speaker that represents the content of an utterance by the speaker; and an estimation unit that estimates which of a plurality of types of pattern the speaker belongs to based on the frequently occurring words and a plurality of emotion values ​​that represent the emotions of the speaker, wherein the emotion values ​​are either statistics of evaluation values ​​that represent each emotion over a predetermined period from the start of a conversation, statistics of evaluation values ​​that represent each emotion over a predetermined period going back a certain period from the end of the conversation, or statistics of evaluation values ​​that represent each emotion over a predetermined period that has elapsed since a specific keyword was uttered in the conversation, and the estimation unit estimates which of a plurality of types of pattern the speaker belongs to using the emotion values, numerical values ​​that represent the frequently occurring words, and weights for each of the plurality of types of patterns. is.

[0010] Another aspect of the present invention is an information processing method to be executed by a computer of an information processing device, the information method including a prediction step of predicting the evaluation value by inputting a plurality of emotional values ​​that correspond to text data for each speaker that represents the content of the utterance by the speaker and that represent the emotions of the speaker into a learning model that is generated based on a correspondence between the emotional values ​​and an evaluation value for the conversation that corresponds to the text data.

[0011] Another aspect of the present invention is an information processing method to be executed by a computer of an information processing device, the information processing method comprising: a frequent word extraction step of extracting one or more frequently occurring words included in text data for each speaker that represents utterance content by the speaker; and an estimation step of estimating which of a plurality of types of pattern the speaker belongs to based on the frequently occurring words and a plurality of emotion values ​​that represent the emotions of the speaker, wherein the emotion values ​​are either statistics of evaluation values ​​that represent each emotion over a predetermined period from the start of a conversation, statistics of evaluation values ​​that represent each emotion over a predetermined period going back a predetermined period from the end of the conversation, or statistics of evaluation values ​​that represent each emotion over a predetermined period that has elapsed since a specific keyword was uttered in the conversation, and the estimation step estimates which of a plurality of types of pattern the speaker belongs to using the emotion values, numerical values ​​that represent the frequently occurring words, and weights for each of the plurality of types of patterns.

[0012] Another aspect of the present invention is a program for causing a computer of an information processing device to execute a prediction step of predicting the evaluation value by inputting the emotional values ​​into a learning model generated based on a correspondence between a plurality of emotional values ​​that correspond to text data for each speaker representing the content of the speaker's utterance and represent the emotions of the speaker, and an evaluation value for the conversation corresponding to the text data.

[0013] Another aspect of the present invention is a program that causes a computer of an information processing device to execute: a frequent word extraction step of extracting one or more frequently occurring words included in text data for each speaker that represents the content of utterances by the speaker; and an estimation step of estimating which of a plurality of types of pattern the speaker belongs to, based on the frequently occurring words and a plurality of emotion values ​​that represent the emotions of the speaker, wherein the emotion values ​​are either statistics of evaluation values ​​that represent each emotion over a predetermined period from the start of a conversation, statistics of evaluation values ​​that represent each emotion over a predetermined period going back a predetermined period from the end of the conversation, or statistics of evaluation values ​​that represent each emotion over a predetermined period that has elapsed since a specific keyword was uttered in the conversation, and wherein the estimation step estimates which of a plurality of types of pattern the speaker belongs to, using the emotion values, numerical values ​​that represent the frequently occurring words, and weights for each of the plurality of types of patterns. [Effects of the Invention]

[0014] According to the present invention, it is possible to analyze emotions by understanding the content and context of a conversation from text data based on voice data. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a system configuration diagram showing an example of the configuration of an information processing system according to a first embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing an example of a hardware configuration of the information processing device according to the present embodiment. [Figure 3] FIG. 2 is a block diagram illustrating an example of a functional configuration of the information processing device according to the present embodiment. [Figure 4] 1A to 1C are explanatory diagrams showing examples of an emotion-word matrix, a weight matrix, and a pattern matrix according to the present embodiment. [Figure 5] FIG. 2 is a diagram showing an example of text data stored in a storage unit of the information processing device according to the present embodiment. [Figure 6] FIG. 2 is a diagram illustrating an example of integrated data stored in a storage unit of the information processing device according to the embodiment. [Figure 7] FIG. 10 is a diagram showing an example of an output display in the information processing device according to the embodiment. [Figure 8] 10 is a flowchart illustrating an example of information processing in the information processing device according to the present embodiment. [Figure 9] FIG. 10 is a block diagram showing an example of a functional configuration of an information processing device according to a second embodiment of the present invention. [Figure 10] 5 is a diagram showing an example of learning data used by a pattern learning unit in the information processing device according to the present embodiment. FIG. [Figure 11] FIG. 10 is a diagram showing an example of an output display in the information processing device according to the embodiment. [Figure 12] 10 is a flowchart illustrating an example of information processing in the information processing device according to the present embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0016] [First embodiment] Hereinafter, each embodiment of the present invention will be described with reference to the drawings. <Configuration of information processing system SYS> First, the configuration of the information processing system SYS will be described.

[0017] FIG. 1 is a system configuration diagram showing an example of the configuration of an information processing system SYS according to the first embodiment of the present invention. The information processing system SYS includes an information processing device 100 and a call terminal device 200. The information processing device 100 and the call terminal device 200 are connected via a network NW so as to be able to communicate with each other.

[0018] The information processing system SYS will now be described in more detail.

[0019] The information processing system SYS is a system that extracts multiple emotional values ​​representing multiple emotions of multiple people (e.g., customers and operators) from voice data of calls between customers and operators at a call center, analyzes the content of the conversation (content of the call) and context by extracting frequently occurring words and the like from text data obtained by speech recognition of the voice data, and infers which of multiple patterns the conversation and speaker (e.g., operator) falls into based on the emotional values ​​and frequently occurring words, and outputs the inferred result. Here, the frequently occurring words include single words and phrases consisting of multiple words, but may also be either single words or phrases consisting of multiple words.

[0020] More specifically, the information processing system SYS acquires voice data and extracts, for each speaker, multiple emotional values ​​representing the emotions of the speaker in the voice data and text data representing the content of the speaker's utterances through voice recognition. The information processing system SYS also extracts one or more frequently occurring words contained in the text data, estimates which of multiple patterns the speaker belongs to based on the frequently occurring words and emotional values, and outputs the estimated pattern.

[0021] With this configuration, the information processing system SYS can grasp the context of a conversation from frequently occurring words extracted from text data based on voice data and estimate patterns according to the frequently occurring words and emotional values, thereby performing emotional analysis according to the conversation.

[0022] The information processing device 100 acquires voice data of a call and extracts from the voice data multiple emotion values ​​that represent the emotions of each speaker (e.g., the operator and the customer) over a predetermined period of time, for example, one minute from the start of the conversation. The emotions are, for example, joy, anger, sadness, happiness, empathy, kindness, sincerity, and other emotions felt by the customer or the operator during the call. The emotion values ​​are numerical values ​​that indicate the degree of the above-mentioned emotions.

[0023] The emotion is not limited to these, and any emotion that a customer or an operator may have may be used. The predetermined period for extracting the emotion value of a conversation from the voice data may be a period going back a certain period from the end of the conversation, or a certain period from the timing when a specific keyword is uttered in the conversation.

[0024] Furthermore, the information processing device 100 extracts text data of the conversation from the voice data by voice recognition. The text data corresponds to the voice data from which the emotion value was extracted. The information processing device 100 grasps the content and context of a conversation by extracting frequently occurring words from text data. The information processing device 100 estimates which of multiple patterns the conversation falls into using the frequently occurring words and emotion values. The information processing device 100 outputs the estimation result.

[0025] The call terminal device 200 is a terminal device used for calls. The call terminal device 200 may be a telephone, a smartphone, a call application, a chat tool, or may be a recorder or camera capable of recording audio data. In this embodiment, the call terminal device 200 is a terminal device used by an operator for calls with customers. The call terminal device 200 has a function for making calls with customers and a function for recording the calls.

[0026] In this embodiment, a call at a call center is described as an example, but the present invention can be applied to any conversation between multiple people, such as face-to-face or online.

[0027] Next, the hardware configuration of the information processing device 100 will be described.

[0028] <Hardware configuration> FIG. 2 is a block diagram showing an example of the hardware configuration of the information processing device according to this embodiment. The information processing device 100 includes a CPU 101, a storage medium interface unit 102, a storage medium 103, an input device 104, an output device 105, a ROM 106 (Read Only Memory), a RAM 107 (Random Access Memory), an auxiliary storage unit 108, and a network interface unit 109. The CPU 101, the storage medium interface unit 102, the input device 104, the output device 105, the ROM 106, the RAM 107, the auxiliary storage unit 108, and the network interface unit 109 are connected to each other via a bus.

[0029] The CPU 101 referred to here refers to a processor in general, and includes not only a device called a CPU in the narrow sense, but also, for example, a GPU, a DSP, etc. Furthermore, the CPU 101 referred to here is not limited to being realized by a single processor, but may be realized by combining multiple processors of the same or different types.

[0030] <cpu101> The CPU 101 controls the information processing device 100 by reading and executing programs stored in the auxiliary storage unit 108, the ROM 106, and the RAM 107, and by reading various data stored in the auxiliary storage unit 108, the ROM 106, and the RAM 107 and writing the various data to the auxiliary storage unit 108 and the RAM 107. The CPU 101 also reads various data stored in the storage medium 103 via the storage medium interface unit 102 and writes the various data to the storage medium 103.

[0031] <Storage medium 103> The storage medium 103 is a portable storage medium such as a magneto-optical disk, a flexible disk, or a flash memory, and stores various data.

[0032] <Storage medium interface unit 102> The storage medium interface unit 102 is an interface for reading and writing data from and to the storage medium 103 .

[0033] <Input device 104> The input device 104 is an input device such as a mouse, a keyboard, a touch panel, a microphone, a volume control button, a power button, a setting button, and an infrared receiver.

[0034] <Output Device 105> The output device 105 is an output device such as a display unit and a speaker.

[0035] <ROM106、RAM107> The ROM 106 and RAM 107 store programs for operating the various functional units of the information processing device 100 and various data.

[0036] <Auxiliary storage section 108> The auxiliary storage unit 108 is a hard disk drive, a flash memory, or the like, and stores programs for operating each functional unit of the information processing device 100 and various data.

[0037] <Network Interface Unit 109> The network interface unit 109 has a communication interface and is connected to the network NW via wireless communication.

[0038] For example, the CPU 101 of the information processing device 100 corresponds to the control unit 15 in the functional configuration shown in Fig. 3. The ROM 106, RAM 107, auxiliary storage unit 108, or any combination thereof of the information processing device 100 corresponds to the storage unit 12 in the functional configuration shown in Fig. 3. The input device 104 and the output device 105 of the information processing device 100 correspond to the input unit 13 and the output unit 14 in the functional configuration shown in Fig. 3.

[0039] Although illustration and description of the hardware configuration of the call terminal device 200 will be omitted, the call terminal device 200 has the same hardware configuration as the information processing device 100 shown in FIG.

[0040] Next, the functional configuration of the information processing device 100 will be described.

[0041] <Functional configuration of information processing device 100> FIG. 3 is a block diagram showing an example of the functional configuration of the information processing device 100 according to this embodiment. The information processing device 100 includes a communication unit 11, a storage unit 12, an input unit 13, an output unit 14, and a control unit 15. The communication unit 11, the storage unit 12, the input unit 13, the output unit 14, and the control unit 15 are connected to each other via a bus.

[0042] <Communications Department 11> The communication unit 11 has a function of communicating with the call terminal device 200. The communication unit 11 outputs various information received from the call terminal device 200 to the control unit 15. In addition, the communication unit 11 transmits information input from the control unit 15 to the call terminal device 200.

[0043] <Storage section 12> The storage unit 12 is configured by a storage medium, such as a hard disk drive (HDD), flash memory, electrically erasable programmable read-only memory (EEPROM), random access read / write memory (RAM), read-only memory (ROM), or any combination of these storage media. The storage unit 12 can be, for example, a nonvolatile memory.

[0044] The storage unit 12 stores text data 121, integrated data 122, and pattern data 123. The text data 121 is generated by speech recognition of the voice data. Here, the text data 121 is data associated with an emotion value extracted from the voice data. The integrated data 122 is data in which frequently occurring words extracted from text data, emotional values ​​of conversations in the voice data corresponding to the text data, and call types are associated with each other. The call type is attribute information that identifies the type of call, such as an inbound call (received to a call center), an outbound call (made from a call center), or the content of an inquiry. The pattern data 123 is data of a plurality of model patterns in which frequently occurring words are associated with emotional values.

[0045] <Input section 13> The input unit 13 is an input device such as a mouse, keyboard, or microphone connected to the information processing device 100. The input unit 13 accepts an operation input input from outside. The input unit 13 outputs an operation signal corresponding to the operation input to the control unit 15.

[0046] <Output section 14> The output unit 14 is an output device such as a display device, etc. The output unit 14 outputs the presentation information output from the control unit 15 to the output device or another device such as the call terminal device 200.

[0047] <Control unit 15> The control unit 15 has a function of controlling the information processing device 100. The control unit 15 reads out various data, applications, programs, etc. stored in the storage unit 12 and controls the information processing device 100.

[0048] The processing of the control unit 15 will be described in more detail. The control unit 15 includes a voice data acquisition unit 151 , a text data extraction unit 152 , an emotion / text integration unit 153 , a pattern estimation unit 154 , and a pattern output unit 155 .

[0049] <Audio data acquisition unit 151> The voice data acquisition unit 151 acquires voice data from the communication terminal device 200 via the network NW and the communication unit 11. The voice data acquisition unit 151 outputs the voice data to the text data extraction unit 152.

[0050] <Text data extraction unit 152> When voice data is input from voice data acquisition unit 151, text data extraction unit 152 extracts text data by voice recognition of the voice data. The text data is text data corresponding to voice data for a predetermined period of time, for example, one minute, from the start of the call. The text data is also text data obtained by converting speakers and the content of each speaker's speech into text through speaker recognition. Text data extraction unit 152 stores the extracted text data in storage unit 12 in association with identification information that identifies the call. Furthermore, the text data extraction unit 152 extracts, for each utterance, multiple emotional values ​​of the speaker, for example, the customer and the operator, for a predetermined period from the start of the call from the voice data, using, for example, known technology. The text data extraction unit 152 associates the extracted emotional values ​​with the text data and stores them in the storage unit 12. Here, speaker recognition is performed by recognizing the voices of multiple speakers using existing voice separation technology or the like.

[0051] The emotion value may be a statistical quantity such as a mean value, median value, or variance value obtained by averaging each emotion value over a predetermined period. The predetermined period may be a period until the end of the conversation, a period going back a certain period from the end of the conversation, or a period of a certain amount of time elapsed from the time a specific keyword was uttered during the conversation.

[0052] <Emotion and Text Integration Section 153> The emotion / text integrator 153 reads the text data 121 from the storage unit 12 and integrates the emotion value with the text data for each utterance. Specifically, the emotion / text integrator 153 extracts words (frequently occurring words) and phrases for each utterance in the text data and quantifies the importance of each extracted word and phrase using, for example, TF-IDF (Term Frequency - Inverse Document Frequency). For each utterance, the emotion / text integrator 153 associates the emotion value, frequently occurring words, and scores such as the importance of each word and phrase, and stores them in the storage unit 12 as integrated data 122.

[0053] <Pattern estimation unit 154> The pattern estimation unit 154 reads the integrated data 122 from the storage unit 12 and extracts patterns of emotions and words. Specifically, the pattern estimation unit 154 extracts emotion values ​​and word scores from the integrated data 122 and defines them as an emotion-word matrix. The pattern estimation unit 154 decomposes the emotion-word matrix into a weight matrix and a pattern matrix. For example, singular value decomposition, non-negative matrix factorization, neural network, etc. are used for the decomposition into the weight matrix and the pattern matrix. The pattern estimation unit 154 associates the emotion-word matrix, weight matrix, and pattern matrix and stores them in the storage unit 12 as pattern estimation results.

[0054] In this embodiment, for example, non-negative matrix factorization is used for decomposition into a weight matrix and a pattern matrix. In this way, all of the values ​​of the decomposed matrixes are non-negative, making it easier to intuitively understand the pattern matrix and weight matrix obtained as a result of the decomposition.

[0055] Furthermore, the pattern estimation unit 154 is not limited to using a linear algorithm such as matrix decomposition, and may extract patterns of emotions and words using a nonlinear algorithm such as a k-nearest neighbor method, UMAP (Uniform Manifold Approximation and Projection), or t-SNE (t-distributed stochastic neighbor embedding), or a combination thereof.

[0056] This will be explained with reference to FIG.

[0057] FIG. 4 is an explanatory diagram of the emotion-word matrix, weight matrix, and pattern matrix according to this embodiment. The illustrated example is an example in which there is one type of operator emotion value (OP emotion value), one type of customer emotion value (CU emotion value), and one type of word, and the number of patterns is three. If the number of calls is N, the (N × 3) emotion-word matrix is ​​decomposed into an (N × 3) weight matrix and a (3 × 3) pattern matrix. The weight matrix indicates the weight of each pattern in each call, and also indicates which emotions and words have a strong relationship in each conversation. Each column of the pattern matrix indicates the relative strength of the relationship between emotion and word by its weight, and each row value of the pattern matrix represents a common pattern of words and emotions in N conversations.

[0058] The size of the matrix is ​​not limited to this, and can be changed depending on the number of emotion values ​​that can be obtained, the type of words, and the number of words. Also, numerical values ​​other than word scores and emotion values, such as call type, may be used.

[0059] <Pattern output unit 155> The pattern output unit 155 reads out the pattern estimation results from the storage unit 12, and generates a display image including emotions corresponding to the emotion values ​​using the emotion-word matrix, words corresponding to the word scores, weights in the weight matrix, patterns in the pattern matrix, and pattern data based on past cases. The pattern output unit 155 outputs the generated display image to an output device or the call terminal device 200.

[0060] The pattern output unit 155 may output both each pattern of the pattern matrix and the word corresponding to the word score, that is, the score of the frequently occurring word, or may output either one of them.

[0061] Next, various data stored in the storage unit 12 of the information processing device 100 will be described.

[0062] FIG. 5 is a diagram showing an example of text data 121 stored in the storage unit 12 of the information processing device 100 according to this embodiment. The illustrated text data is data in which a call identification number, a start time, an end time, a speaker, the content of the speech, an emotion value (A), an emotion value (B), and an emotion value (C) are associated with each other. The call identification number is identification information (number) that identifies a call. The start time and end time are time information indicating the start time and end time of an utterance, and are acquired from time information included in the voice data, for example. The speaker represents the speaker whose voice has been separated by speaker recognition. Note that as long as the speaker can be distinguished, such as Speaker A and Speaker B, it is not necessary to distinguish between different types of people, such as "operator" and "customer." The speech content is text data representing the content of the speech by the speaker, and is generated by speech recognition. Emotion value (A), emotion value (B), and emotion value (C) are numerical values ​​that represent each of the preset emotions. For example, emotion value (A) is a numerical value that represents joy, emotion value (B) is a numerical value that represents anger, and so on. Emotion values ​​corresponding to multiple emotions are output from the voice data for each utterance using known technology. As shown in the figure, the text data associates the content of an utterance for each speaker with a plurality of emotion values ​​for each utterance content.

[0063] FIG. 6 is a diagram showing an example of the integrated data 122 stored in the storage unit 12 of the information processing device 100 according to this embodiment. The integrated data shown in the figure is data in which call identification numbers, operator word importance levels, customer word importance levels, operator emotion values, and customer emotion values ​​are associated with each other. The integrated data may be associated with a call type. The call identification number is identification information that identifies a call. The operator word importance indicates the score of a word extracted from the speech content of the operator. For example, the operator word importance indicates a predetermined number of words in descending order of word importance, as a numerical value associated with the importance of each word. The customer word importance indicates the score of a word extracted from the customer's speech. For example, the customer word importance indicates a predetermined number of words (frequently occurring words) in descending order of word importance, and is represented as a numerical value associated with the importance of each word. The operator emotion value and the customer emotion value are emotion values ​​that represent emotions extracted from the speech content of the operator and the customer.

[0064] Although an example of generating integrated data by converting the emotional values ​​into numerical values ​​has been described, the integrated data may also be generated by converting the emotional values ​​into text.

[0065] Next, an example of output from the output unit 14 of the information processing device 100 will be described.

[0066] FIG. 7 is a diagram showing an example of an output display in the information processing device 100 according to this embodiment. The output display shown in the figure is an example of a case where there are three types of patterns and the display image generated by the pattern output unit 155 is output to the output device. For example, let us assume that emotion values ​​(A) of emotion A and emotion values ​​(B) of emotion B of the operator and the customer are extracted from the voice data, that the words UUU, VVV, and WWW are extracted as frequently occurring words based on the voice data, and that the word importance (word scores) of the frequently occurring words are extracted. The information processing device 100 generates an emotion-word matrix according to the emotion values ​​and word scores, and decomposes the emotion-word matrix into a weight matrix and a pattern matrix. The displayed image includes an area for displaying the weight of emotion / word patterns, an area for listing emotion / word patterns, an area for pattern examples, and an area for overall evaluation.

[0067] The emotion / word pattern weight display area reads each column of the weight matrix and displays the weight of the pattern for each call. The strength (large or small) of the weight of each pattern expresses the characteristics of the call content. By using the pattern weights in this way, it is possible to distinguish and compare the trends of each call. In the example shown, it is displayed as "Pattern 1 weight 1.5," "Pattern 2 weight 0.1," and "Pattern 3 weight 0.5."

[0068] The emotion / word pattern list area is an area that reads the values ​​of each row of the pattern matrix and displays the patterns in various graphs. The emotion / word pattern list area is also an area that displays, for example, a predetermined number of highly important words (words with large values ​​in the weight matrix) as "frequent words." Instead of words, important phrases or prohibited phrases may be displayed to the call center.

[0069] The pattern example area is a surface area that displays pattern data stored in the memory unit 12 as pattern data. For example, the pattern data may hold emotion and word patterns of highly skilled operators and emotion and word patterns of less experienced and less skilled operators. The pattern output unit 155 searches the pattern data stored in the memory unit 12 and displays good patterns to follow and bad patterns to serve as counterexamples. This allows operators to, for example, refer to customer emotion values ​​from past cases to check how emotionally they should speak, what words they should use, etc. It is also possible to identify the content of a call from words, for example, and compare it with the emotion of a highly skilled operator at that time.

[0070] Pattern data may also contain emotion and word patterns linked to other indicators related to call center and sales conversations. Indicators include, for example, operator turnover rate and customer satisfaction surveyed separately via social media, the web, or questionnaires. Linking these indicators to emotion and word patterns makes it possible to understand trends, such as which emotion and word patterns improve indicators.

[0071] The comprehensive evaluation area is an area that displays advice that compares the emotion and word patterns in a certain call with patterns stored as pattern data. In the example shown, the conversation content is estimated based on frequently occurring words and emotional values, and the display shows an evaluation of whether the conversation is being conducted with appropriate emotions and whether appropriate words are being used, advice on how to improve emotions, and advice on words that should be used to improve customer impressions such as satisfaction, such as "Evaluation: In the current conversation content, pattern 1 is closest. In the content of this call, it is close to an undesirable pattern.", "Emotional advice: Try to lower the tone of your voice," and "Word advice: Talking to XXX will improve customer satisfaction."

[0072] Next, the flow of information processing according to this embodiment will be described.

[0073] FIG. 8 is a flowchart showing an example of information processing in the information processing device 100 according to this embodiment. In step S101, the information processing device 100 acquires audio data. Then, the information processing device 100 executes the process of step S102. In step S102, the information processing device 100 extracts text data from the voice data by voice recognition. Then, the information processing device 100 executes the process of step S103.

[0074] In step S103, information processing device 100 extracts a plurality of emotion values ​​for each conversation from the voice data. After that, information processing device 100 executes the process of step S104. In step S104, information processing device 100 generates integrated data that integrates the text data and emotional values. Specifically, information processing device 100 calculates word importance for each conversation and generates integrated data by associating a predetermined number of words with high word importance with the word importance and the emotional values ​​extracted from the audio data of the conversation. Information processing device 100 then performs the process of step S105.

[0075] In step S105, the information processing device 100 reads the integrated data and generates an emotion-word matrix. The information processing device 100 also decomposes the emotion-word matrix into a weight matrix and a pattern matrix to estimate a conversation pattern. The information processing device 100 then executes the process of step S106.

[0076] Here, estimating a conversation pattern includes indicating weights for multiple patterns, such as the above-mentioned "weight of 1.5 for pattern 1," "weight of 0.1 for pattern 2," and "weight of 0.5 for pattern 3." Estimating a conversation pattern also includes selecting the weight with the largest weight value from the weights of multiple patterns and estimating, for example, that it is pattern 1 or that pattern 1 is strong based on the selected weight. Furthermore, estimating a conversation pattern also includes setting a threshold for the weights of multiple patterns and estimating a tendency from patterns with weights above the threshold, such as when the threshold is 0.3 in the above-mentioned example, pattern 1 is strong, followed by pattern 3.

[0077] In step S106, the information processing device 100 outputs the estimation result, and then the information processing device 100 ends the information processing according to FIG.

[0078] As described above, the information processing system SYS according to this embodiment includes a voice data acquisition unit 151 that acquires voice data, an extraction unit (text data extraction unit 152) that uses voice recognition to extract, for each speaker, from the voice data, a plurality of emotion values ​​that represent the emotions of the speaker in the voice data and text data that represents the content of the utterance by the speaker, a frequent word extraction unit (emotion / text integrating unit 153) that extracts one or more frequently occurring words included in the text data and further integrates the text data and the emotion values ​​into a numerical value or text format to generate integrated data, an estimation unit (pattern estimation unit 154) that estimates which of a plurality of types of pattern the utterance belongs to based on the frequently occurring words and emotion values, and an output unit (pattern output unit 155) that outputs the estimated pattern.

[0079] This allows for sentiment analysis after grasping the conversation content and context from text data based on voice data. Furthermore, by extracting and displaying patterns simultaneously with words and phrases in addition to emotional values, it is possible to visualize the speaker's skills according to the content of the utterance. Furthermore, by providing feedback on model emotions and speaking tendencies, it is possible to improve the speaker's skills according to the content of the utterance.

[0080] Furthermore, in the information processing system SYS, the emotion values ​​are statistics such as the average values ​​of the emotion values ​​over a predetermined period of time from the start of the conversation, and the estimation unit (pattern estimation unit 154) estimates which of the multiple patterns the speaker belongs to using the emotion values, numerical values ​​representing frequently occurring words, and weights for each of the multiple types of patterns.

[0081] By doing this, it is possible to use statistical quantities such as emotion values ​​and average emotion values ​​over a predetermined period from the start of a conversation, which tend to reveal the characteristics of the conversation content, thereby improving the accuracy of estimating patterns suited to the conversation content.

[0082] In the information processing system SYS, the output unit (pattern output unit 155) outputs patterns and frequently occurring words.

[0083] By doing this, you can check what words and emotions to use in the conversation.

[0084] In the information processing system SYS, the voice data is call voice data, and the output unit (pattern output unit 155) matches a speaker with a call partner of the speaker according to emotions and word patterns. Specifically, for an operator, if the pattern is similar (approximate) to, for example, "Pattern 1 weight 1.5," "Pattern 2 weight 0.1," or "Pattern 3 weight 0.5," it is considered a friendly pattern, and if the pattern is similar to, for example, "Pattern 1 weight 1.0," "Pattern 2 weight 1.0," or "Pattern 3 weight 1.0," it is considered a calm pattern. Furthermore, with regard to customers, if the pattern is similar to, for example, "Pattern 1 weight 1.0," "Pattern 2 weight 1.0," or "Pattern 3 weight 1.0," it is considered a sociable pattern, and if the pattern is similar to, for example, "Pattern 1 weight 1.5," "Pattern 2 weight 1.0," or "Pattern 3 weight 1.0," it is considered a critical pattern.

[0085] In this case, matching patterns are previously defined in the information processing system SYS so that a customer with a sociable pattern is matched with an operator with a friendly pattern, and a customer with a critical pattern is matched with an operator with a calm pattern. As a result, the output unit (pattern output unit 155) determines patterns for the characteristics of the customer and the agent, and matches the customer with the agent according to the matching pattern.

[0086] The information processing device 100 also includes a voice data acquisition unit 151 that acquires voice data, an extraction unit (text data extraction unit 152) that extracts, for each speaker, a plurality of emotion values ​​that represent the emotions of the speaker in the voice data and text data that represents the content of the utterance by the speaker from the voice data by voice recognition, a frequent word extraction unit (emotion / text integrating unit 153) that extracts one or more frequently occurring words included in the text data, an estimation unit (pattern estimation unit 154) that estimates which of a plurality of patterns the speaker belongs to based on the frequently occurring words and the emotion values, and an output unit (pattern output unit 155) that outputs the estimated pattern.

[0087] This allows for sentiment analysis after grasping the conversation content and context from text data based on voice data. Furthermore, by extracting and displaying patterns simultaneously with words and phrases in addition to emotional values, it is possible to visualize the speaker's skills according to the content of the utterance. Furthermore, by providing feedback on model emotions and speaking tendencies, it is possible to improve the speaker's skills according to the content of the utterance.

[0088] [Second embodiment] In the second embodiment, an example of learning conversation-related indices by supervised learning will be described. Here, in the second embodiment, the differences from the first embodiment will be mainly described, and the first embodiment will be used for the other parts, and the description thereof will be omitted.

[0089] FIG. 9 is a block diagram showing an example of the functional configuration of an information processing apparatus 100 according to the second embodiment of the present invention. The information processing device 100 includes a communication unit 11, a storage unit 12, an input unit 13, an output unit 14, and a control unit 15. The communication unit 11, the storage unit 12, the input unit 13, the output unit 14, and the control unit 15 are connected to each other via a bus.

[0090] <Storage section 12> The storage unit 12 stores text data 121, integrated data 122, pattern data 123, index data 124, and a learning model 125. The index data 124 is, for example, data on indexes to be checked by a QA (Quality Administrator) of a call center, index data related to management, index data to be improved by analyzing emotions and words, and related data thereof.

[0091] Examples of indicator data that QA should check include the following: Response Time: The time it takes for an operator to respond to a customer inquiry. · Average Handling Time: The average length of time a customer spends on a call. First Call Resolution Rate: The percentage of customers who reach a resolution on their first call. Customer Satisfaction: An indicator that measures how satisfied a customer is with the operator's response. Operator Compliance: An indicator that evaluates whether operators comply with corporate and industry rules, legal compliance, and privacy protection. Operator Skills and Knowledge: An indicator that evaluates whether the operator has the appropriate knowledge of products and services and the appropriate skills to handle the calls. Call Quality: Indicators used to evaluate the quality of a call, such as the ease with which the operator's voice can be heard, clarity, appropriate phrases and expressions, phrases and expressions that should not be spoken, and the timing of speech. Escalation Rate: An indicator of the rate at which operators escalate inquiries to a senior or supervisor.

[0092] Furthermore, the management indicators include, for example, the following indicator data: Cost Performance: An indicator of the balance between operational costs and customer satisfaction. Operator Efficiency: An indicator that evaluates whether operators are performing their duties efficiently, including call time, waiting time, and time to resolution. Operator Retention Rate: An indicator of the percentage of operators who have been working at a call center for a long period of time.

[0093] Furthermore, for example, it may be an index related to digital marketing collected from the web, social media, etc. Conversion Rate: An indicator that shows the percentage of visitors who actually make a purchase or make an inquiry as a result of marketing activities. Bounce Rate: An indicator that shows the percentage of visitors who leave the site shortly after visiting it. Click-Through Rate: An indicator that shows the percentage of people who clicked on content such as an advertisement or email newsletter. · Social Share Rate: An indicator showing the percentage of people who shared content on social media etc. · Repeat Rate: An indicator that shows the percentage of people who have previously visited a site who will return. Net Promoter Score (NPS): A metric that measures how likely a customer is to recommend a brand or service to others.

[0094] The learning model 125 is a learning model generated by the pattern learning unit 156 .

[0095] <Control unit 15> The control unit 15 includes a voice data acquisition unit 151 , a text data extraction unit 152 , an emotion / text integration unit 153 , a pattern estimation unit 154 , a pattern output unit 155 , and a pattern learning unit 156 .

[0096] <Pattern Learning Unit 156> Pattern learning unit 156 sets the index data stored in storage unit 12 as a target variable, sets the text data including emotional values ​​stored in storage unit 12 as an explanatory variable, and generates a learning model by learning the correspondence between them using machine learning including deep learning, etc. As a result, by inputting text data including emotional values ​​into the learning model, various indices are output.

[0097] The pattern learning unit 156 may generate a learning model according to the format of the learning data. For example, in the case of quantified learning data, a machine learning model that is strong in table data, such as Lightgbm, may be used, and in the case of text learning data, a language model, such as BERT, may be used.

[0098] <Pattern estimation unit 154> The pattern estimation unit 154 utilizes the learning model stored in the memory unit 12, inputs text data including emotional values ​​as explanatory variables into the learning model, and obtains the index set as the objective variable as output, thereby estimating (predicting) the index.

[0099] <Pattern output unit 155> The pattern output unit 155 generates a display image that shows the relationship between the predicted index, the content of the conversation, and the emotion value.

[0100] Fig. 10 is a diagram showing an example of learning data used by the pattern learning unit 156 in the information processing device 100 according to this embodiment. Fig. 11 is a diagram showing an example of an output display in the information processing device 100 according to this embodiment. The learning data in this embodiment is, for example, text data including emotion values ​​stored in the storage unit 12. The learning data is data in which call identification numbers, text data, operator emotion values, and customer emotion values ​​are associated with each other. Furthermore, the learning data used is text data associated with emotion values ​​shown in FIG. 10 and index data corresponding to the text data. Note that instead of the text data including emotion values ​​shown in Fig. 10, the text data including emotion values ​​shown in Fig. 5 may be used as training data. For example, the emotion values ​​are converted into text, such as "Emotion A=10, Emotion B=20, Emotion C=0", and combined with the text data of the speech recognition results to be used as training data.

[0101] In the learning data shown in FIG. 10, learning may be performed using learning data in which index data stored in the storage unit 12 is associated with the operator emotion value and the customer emotion value, instead of or in addition to the operator emotion value and the customer emotion value.

[0102] For example, when text data including the emotion values ​​shown in FIG. 10 are input as explanatory variables to a learning model trained in this way, indicators A and B corresponding to the call identification number, text data, operator emotion value, and customer emotion value in FIG. 10 are output as objective variables, as shown in FIG. 11. The learning model may also output which words or emotion values ​​were focused on when predicting an index. For example, a deep learning model with an attention mechanism can be used as the learning model, and the output values ​​of the intermediate layer of the learning model can be used to output which words or emotion values ​​were focused on when predicting an index. Here, the output display shown in FIG. 11 in the above description is an example of a display image generated by the pattern output unit 155 based on the index obtained as an output from the learning model. Note that which words or emotion values ​​were focused on when estimating the index may be displayed in a way that is easy for the user to understand by changing the display mode, such as the color of the text or the background color. In this way, it is possible to visualize in an easy-to-understand manner which words or emotion values ​​contributed to the index prediction.

[0103] Next, information processing according to this embodiment will be described.

[0104] FIG. 12 is a flowchart showing an example of information processing in the information processing device 100 according to this embodiment. In step S201, the information processing device 100 acquires audio data. Then, the information processing device 100 executes the process of step S202. In step S202, the information processing device 100 extracts text data from the voice data by voice recognition. Then, the information processing device 100 executes the process of step S203.

[0105] In step S203, information processing device 100 extracts a plurality of emotion values ​​for each conversation from the audio data. After that, information processing device 100 executes the process of step S204. In step S204, information processing device 100 generates integrated data by integrating the text data and emotion values. Specifically, information processing device 100 calculates word importance for each conversation and generates integrated data by associating a predetermined number of words with high word importance with the word importance and the emotion values ​​extracted from the speech data of the conversation. Alternatively, information processing device 100 may generate integrated data by converting the emotion values ​​into text and combining the emotion values ​​converted into text with text data of the speech recognition result. Thereafter, information processing device 100 executes the process of step S205.

[0106] In step S205, when text data including emotion values ​​is input, information processing device 100 generates a learning model that outputs an index (index value). After that, information processing device 100 executes the process of step S206. In step S206, the information processing device 100 reads out the learning model 125 stored in the storage unit 12, and inputs text data including emotion values ​​into the learning model 125 to obtain an index as an output. The information processing device 100 estimates the output index as a pattern. Thereafter, the information processing device 100 executes the process of step S207.

[0107] In step S207, the information processing device 100 outputs the estimation result, and then the information processing device 100 ends the information processing according to FIG.

[0108] As described above, the information processing system SYS according to this embodiment includes a voice data acquisition unit 151 that acquires voice data, an extraction unit (text data extraction unit 152) that uses voice recognition to extract, for each speaker, a plurality of emotion values ​​that represent the emotions of the speaker in the voice data and text data that represents the content of the utterance by the speaker, from the voice data, a frequent word extraction unit (emotion / text integration unit 153) that extracts one or more frequently occurring words included in the text data and further integrates the text data and emotion values ​​into a numerical or text format to generate integrated data, a pattern learning unit 156 that machine-learns a learning model based on the training data, an estimation unit (pattern estimation unit 154) that receives text and emotion values ​​as input and predicts an index using the learning model, and an output unit (pattern output unit 155) that outputs the estimated index.

[0109] In this way, sentiment analysis can be performed by grasping the content and context of the conversation from text data based on the voice data. Furthermore, by simultaneously performing machine learning on not only emotional values ​​but also words and phrases, and predicting and displaying various indicators, it is possible to visualize the speaker's skill according to the content of the utterance. Furthermore, by displaying the emotional values ​​and words that the learning model focused on when predicting the indicators, it is possible to provide feedback on the emotions and speaking tendencies that should be modeled, thereby improving the speaker's skill according to the content of the utterance.

[0110] In the information processing system SYS, the output unit (pattern output unit 155) displays the predicted index. Also, by displaying the emotional values ​​and words that were focused on when predicting the index, it is possible to provide feedback on the emotions and speech tendencies that should be modeled, thereby improving the speaker's skills according to the content of the speech.

[0111] In the information processing system SYS, the voice data is call voice data, and the output unit (pattern output unit 155) matches the speaker with the call partner of the speaker according to the pattern of the predicted index. Specifically, for example, an operator's pattern is friendly if it resembles "index 1 value 1.5" and "index 2 value 0.1," while a calm pattern is similar to "index 1 value 1.0" and "index 2 value 1.0." For customers, a pattern is sociable if it resembles "index 3 value 1.0" and "index 4 value 1.0," while a pattern is critical if it resembles "index 3 value 1.5" and "index 4 value 1.0." The information processing system SYS predefines matching patterns such that a customer with a sociable pattern is matched with an operator with a friendly pattern, and a customer with a critical pattern is matched with an operator with a calm pattern. The output unit (pattern output unit 155) determines the personality patterns of the customer and the operator, and matches the customer and the operator according to the matching pattern.

[0112] In this way, it is possible to match a speaker suitable for the other party according to the pattern.

[0113] The information processing system SYS further includes a prediction unit (not shown) that predicts an evaluation value by inputting the emotional value and text data into a learning model generated by learning the correspondence between text data, an emotional value corresponding to the text data, and an evaluation value for a conversation corresponding to the text data.

[0114] Since the evaluation value can be output along with the pattern, it is possible to visualize the speaker's skill according to the content of the utterance. In addition, it is possible to provide feedback on the emotions and speaking tendencies that should be modeled, and it is possible to improve the speaker's skill according to the content of the utterance.

[0115] The information processing device 100 also includes a voice data acquisition unit 151 that acquires voice data, an extraction unit (text data extraction unit 152) that extracts, for each speaker, multiple emotional values ​​that represent the emotions of the speaker in the voice data and text data that represent the content of the speaker's utterance from the voice data by voice recognition, an integration unit that integrates the text data and the emotional values ​​into a numerical value or text format, a frequent word extraction unit (emotion / text integration unit 153) that integrates the text data and the emotional values ​​into a numerical value or text format to generate integrated data in the process of extracting one or more frequently occurring words included in the text data, a pattern learning unit 156 that machine-learns a learning model based on the learning data, an estimation unit (pattern estimation unit 154) that inputs text and emotional values ​​and predicts an index using the learning model, and an output unit (pattern output unit 155) that outputs the estimated index.

[0116] In this way, sentiment analysis can be performed after understanding the content and context of the conversation from text data based on the voice data. Furthermore, by simultaneously learning not only emotional values ​​but also words and phrases through machine learning and predicting and displaying various indicators, it is possible to visualize the speaker's skill according to the content of the utterance. Furthermore, by displaying the emotional values ​​and words that the learning model focused on when predicting the indicators, it is possible to provide feedback on the emotions and speaking tendencies that should be modeled, thereby improving the speaker's skill according to the content of the utterance.

[0117] <Additional Notes> One aspect of the present invention is an information processing system comprising: a voice data acquisition unit that acquires voice data; an extraction unit that extracts, for each speaker, from the voice data using voice recognition, a plurality of emotional values ​​that represent the emotions of the speaker in the voice data and text data that represents the content of the utterances made by the speaker; a frequent word extraction unit that extracts one or more frequent words included in the text data; an estimation unit that estimates which of a plurality of patterns the speaker belongs to based on the frequent words and the emotional values; and an output unit that outputs a score of the estimated pattern or the frequent words.

[0118] Another aspect of the present invention is an information processing device that includes: a voice data acquisition unit that acquires voice data; an extraction unit that extracts, for each speaker, from the voice data using voice recognition, a plurality of emotional values ​​that represent the emotions of the speaker in the voice data and text data that represents the content of the utterance by the speaker; a frequent word extraction unit that extracts one or more frequent words included in the text data; an estimation unit that estimates which of a plurality of patterns the speaker belongs to based on the frequent words and the emotional values; and an output unit that outputs a score of the estimated pattern or the frequent words.

[0119] Another aspect of the present invention is an information processing method to be executed by a computer of an information processing device, the information processing method having: a voice data acquisition step of acquiring voice data; an extraction step of extracting, for each speaker, from the voice data using voice recognition, a plurality of emotional values ​​that represent the emotions of the speakers in the voice data and text data that represent the content of the utterances made by the speakers; a frequent word extraction step of extracting one or more frequent words included in the text data; an estimation step of estimating which of a plurality of patterns the speaker belongs to based on the frequent words and the emotional values; and an output step of outputting a score of the estimated pattern or the frequent words.

[0120] Another aspect of the present invention is a program that causes a computer of an information processing device to execute the following steps: an audio data acquisition step of acquiring audio data; an extraction step of extracting, for each speaker, from the audio data using speech recognition, a plurality of emotional values ​​that represent the emotions of the speaker in the audio data and text data that represents the content of the utterances made by the speaker; a frequent word extraction step of extracting one or more frequent words included in the text data; an estimation step of estimating which of a plurality of patterns the speaker belongs to based on the frequent words and the emotional values; and an output step of outputting a score of the estimated pattern or the frequent words.

[0121] <Appendix 1> an extraction unit that extracts, for each speaker, from the voice data using voice recognition, a plurality of emotional values ​​that represent the emotions of the speaker in the voice data and text data that represents the content of the speech by the speaker; a frequent word extraction unit that extracts one or more frequently occurring words included in the text data; an estimation unit that estimates which of a plurality of patterns the speaker belongs to based on the frequent words and the emotional values; and an output unit that outputs a score of the estimated pattern or the frequent words.

[0122] <Appendix 2> The information processing system according to claim 1, wherein the emotion values ​​are either statistics of emotion values ​​for a predetermined period from the start of a conversation, statistics of emotion values ​​for a predetermined period going back a certain period from the end of the conversation, or statistics of emotion values ​​for a predetermined period a certain period after a specific keyword was uttered in the conversation, and the estimation unit estimates which of a plurality of patterns the speaker belongs to using the emotion values, numerical values ​​representing the frequently occurring words, and weights for each of a plurality of types of patterns.

[0123] <Appendix 3> and a prediction unit that predicts the evaluation value by inputting the emotional value and the text data into a learning model that is generated by learning a correspondence relationship between the text data, the emotional value corresponding to the text data, and an evaluation value for a conversation that corresponds to the text data.

[0124] <Appendix 4> 2. The information processing system according to claim 1, wherein the frequent word extraction unit integrates the text data and the emotion values ​​corresponding to the text data as all numerical values ​​or all text.

[0125] <Appendix 5> The information processing system described in Appendix 3, wherein the voice data is call voice data, and the output unit matches the speaker with a call partner of the speaker according to the pattern or the evaluation value.

[0126] <Appendix 6> an extraction unit that extracts, for each speaker, from the voice data using voice recognition, a plurality of emotional values ​​that represent the emotions of the speaker in the voice data and text data that represents the content of the utterance by the speaker; a frequent word extraction unit that extracts one or more frequently occurring words included in the text data; an estimation unit that estimates which of a plurality of patterns the speaker belongs to based on the frequent words and the emotional values; and an output unit that outputs a score of the estimated pattern or the frequent words.

[0127] <Appendix 7> An information processing method to be executed by a computer of an information processing device, the information processing method comprising: a voice data acquisition step of acquiring voice data; an extraction step of extracting, for each speaker, from the voice data using voice recognition, a plurality of emotional values ​​that represent the emotions of the speakers in the voice data and text data that represent the content of the utterances made by the speakers; a frequent word extraction step of extracting one or more frequently occurring words included in the text data; an estimation step of estimating which of a plurality of patterns the speaker belongs to based on the frequent words and the emotional values; and an output step of outputting a score of the estimated pattern or the frequent words.

[0128] <Appendix 8> A program for causing a computer of an information processing device to execute the following steps: an audio data acquisition step for acquiring audio data; an extraction step for extracting, for each speaker, from the audio data using voice recognition, a plurality of emotional values ​​that represent the emotions of the speaker in the audio data and text data that represents the content of the speech by the speaker; a frequent word extraction step for extracting one or more frequently occurring words included in the text data; an estimation step for estimating which of a plurality of patterns the speaker belongs to based on the frequent words and the emotional values; and an output step for outputting a score of the estimated pattern or the frequent words.

[0129] Each embodiment of the present invention has been described in detail above with reference to the drawings, but the specific configuration is not limited to that described above, and various design changes can be made within the scope that does not deviate from the gist of the present invention.

[0130] For example, in each of the above-described embodiments, an example has been described in which the information processing device 100 and the call terminal device 200 are configured as individual devices, but one aspect of the present invention may also be realized by a device that combines some or all of these devices, or a device that rearranges some of these devices.

[0131] The program running on the information processing device 100 and the call terminal device 200 according to one aspect of the present invention may be a program that controls one or more processors, such as a central processing unit (CPU), to realize the functions described in the above-described embodiments and modifications related to one aspect of the present invention (a program that causes a computer to function). The term "computer" as used herein also includes quantum computers. Information handled by each of these devices may be temporarily stored in random access memory (RAM) during processing, and then stored in various storage devices such as flash memory and hard disk drives (HDDs), and may be read, modified, or written by the CPU or the like as needed.

[0132] Note that part or all of the information processing device 100 and the call terminal device 200 in each of the above-described embodiments and modifications may be realized by a computer having one or more processors. In this case, a program for realizing this control function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read into a computer system and executed to realize the control function.

[0133] The term "computer system" used here refers to a computer system built into the information processing device 100 and the communication terminal device 200, and includes hardware such as an OS and peripheral devices. Also, the term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into the computer system.

[0134] Furthermore, the term "computer-readable recording medium" may include a medium that dynamically stores a program for a short period of time, such as a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, or a medium that stores a program for a fixed period of time, such as volatile memory within a computer system that serves as a server or client in such a case. The program may also be one that realizes part of the above-mentioned functions, or one that can realize the above-mentioned functions in combination with a program already stored in the computer system.

[0135] Furthermore, part or all of the information processing device 100 and the call terminal device 200 in each of the above-described embodiments and modifications may be realized as an LSI, which is typically an integrated circuit, or as a chipset. Furthermore, each functional block of the information processing device 100 and the call terminal device 200 in each of the above-described embodiments and modifications may be individually formed into a chip, or part or all of them may be integrated into a chip. Furthermore, the integrated circuit method is not limited to LSI, and may be realized using a dedicated circuit and / or a general-purpose processor. Furthermore, if an integrated circuit technology that replaces LSI emerges due to advances in semiconductor technology, it is also possible to use an integrated circuit based on that technology.

[0136] While the embodiments and modifications have been described above in detail with reference to the drawings as one aspect of the present invention, the specific configuration is not limited to the embodiments and modifications, and design changes within the scope of the present invention are also included. Furthermore, various modifications of one aspect of the present invention are possible within the scope of the claims, and embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention. Furthermore, configurations in which elements described in the above embodiments and modifications are substituted with elements that achieve the same effect are also included. [Explanation of symbols]

[0137] SYS Information Processing System 100 Information processing device 101 CPU 102 Storage medium interface unit 103 Storage medium 104 Input Device 105 Output Device 106 ROM 107 RAM 108 Auxiliary storage 109 Network Interface Unit 11 Communications Department 12 Storage section 121 Text Data 122 Integrated Data 123 Pattern Data 124 Index Data 125 Learning Model 13 Input section 14 Output section 15 Control Unit 151 Audio data acquisition unit 152 Text Data Extraction Unit 153 Emotion and Text Integration Department 154 Pattern Estimation Unit 155 Pattern output section 156 Pattern Learning Unit 200 Telephone terminal device

Claims

1. a prediction unit that predicts the evaluation value by inputting text data for each speaker that represents the content of an utterance by the speaker, a plurality of emotion values ​​that correspond to the text data and represent the emotions of the speaker, and an evaluation value for a conversation that corresponds to the text data; and An information processing system comprising:

2. The emotion values ​​include an operator emotion value representing an emotion extracted from the content of the operator's utterance, and a customer emotion value representing an emotion extracted from the content of the customer's utterance, the prediction unit predicts the evaluation value by inputting the text data, the operator emotion value, and the customer emotion value into the learning model generated based on correspondence relationships among the text data, the operator emotion value, the customer emotion value, and the evaluation value; The evaluation value includes at least one of the following: the customer satisfaction level of the customer with respect to the response of the operator; an index value evaluating the response skills of the operator; an index value evaluating the knowledge of the operator regarding the product or service; an index value evaluating the call quality of the operator; and a net promoter score indicating the degree to which the customer recommends the product or service. The information processing system according to claim 1 .

3. Further, an output unit outputs the evaluation value corresponding to the text data, the operator emotion value, and the customer emotion value. The information processing system according to claim 2 .

4. The learning model is a deep learning model having an attention mechanism, the prediction unit predicts, based on an output value of an intermediate layer of the learning model, a word in the text data that contributed to the output of the evaluation value or the emotion value; the output unit outputs the words in the text data that contributed to the output of the evaluation value or the emotion value in a display format different from other displays. The information processing system according to claim 3 .

5. a prediction unit that predicts the evaluation value by inputting text data for each speaker that represents the content of an utterance by the speaker, a plurality of emotion values ​​that correspond to the text data and represent the emotions of the speaker, and an evaluation value for a conversation that corresponds to the text data; and An information processing device comprising:

6. The emotion values ​​include an operator emotion value representing an emotion extracted from the content of the operator's utterance, and a customer emotion value representing an emotion extracted from the content of the customer's utterance, the prediction unit predicts the evaluation value by inputting the text data, the operator emotion value, and the customer emotion value into the learning model generated based on correspondence relationships among the text data, the operator emotion value, the customer emotion value, and the evaluation value; The evaluation value includes at least one of the following: the customer satisfaction level of the customer with respect to the response of the operator; an index value evaluating the response skills of the operator; an index value evaluating the knowledge of the operator regarding the product or service; an index value evaluating the call quality of the operator; and a net promoter score indicating the degree to which the customer recommends the product or service. The information processing device according to claim 5 .

7. Further, an output unit outputs the evaluation value corresponding to the text data, the operator emotion value, and the customer emotion value. The information processing device according to claim 6 .

8. The learning model is a deep learning model having an attention mechanism, the prediction unit predicts, based on an output value of an intermediate layer of the learning model, a word in the text data that contributed to the output of the evaluation value or the emotion value; the output unit outputs the words in the text data that contributed to the output of the evaluation value or the emotion value in a display format different from other displays. The information processing device according to claim 7 .

9. An information processing method to be executed by a computer of an information processing device, comprising: a prediction step of predicting the evaluation value by inputting text data for each speaker that represents the content of an utterance by the speaker, a plurality of emotion values ​​that correspond to the text data and represent the emotions of the speaker, and an evaluation value for the conversation that corresponds to the text data; An information processing method comprising:

10. A computer of an information processing device, a prediction step of predicting the evaluation value by inputting text data for each speaker that represents the content of an utterance by the speaker, a plurality of emotion values ​​that correspond to the text data and represent the emotions of the speaker, and an evaluation value for the conversation that corresponds to the text data; A program to execute.

Citation Information

Patent Citations

  • Business assessment method, business assessment device and business assessment program

    JP2018041120A

  • Speaker estimation device

    JP2019124835A

  • Information providing system, information providing method, and computer program

    JP2022154230A