Meeting minutes generation method, real-time customer service support method, voice interaction method, meeting minutes generation device, real-time customer service support device, voice interaction device, and recording medium
Patent Information
- Application Number
- PCT/JP2025/011369
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2026-10-01
Smart Images

Figure JP2025011369_01102026_PF_FP_ABST
Abstract
Description
Minutes generation method, real-time customer service support method, voice interaction method, minutes generation device, real-time customer service support device, voice interaction device, and recording medium
[0001] The present disclosure relates to a minutes generation method, a real-time customer service support method, a voice interaction method, a minutes generation device, a real-time customer service support device, a voice interaction device, and a recording medium.
[0002] It has been widely known in behavioral economics that user emotions have a significant impact on purchasing behavior and decision-making. Patent Document 1 discloses storing utterance contents and emotion analysis results of an operator and a customer in a storage device. Further, Patent Document 1 discloses setting a predetermined threshold for the emotion levels of the customer and the operator respectively, and proposing an appropriate response to the operator or issuing a warning to a supervisor when the emotion level exceeds the threshold.
[0003] Japanese Unexamined Patent Application Publication No. 2020-115244
[0004] An example of an object of the present disclosure is to develop technology related to customer response.
[0005] According to one aspect of the present disclosure, there is provided a minutes generation method characterized in that one or more computers: acquire video data and audio data; analyze emotions from the video data and the audio data; generate text data from the audio data; integrate the emotion analysis result and the text data in time series; and generate minutes reflecting a flow of emotions based on the integrated data.
[0006] Further, according to one aspect of the present disclosure, there is provided a minutes generation device comprising: means for acquiring video data and audio data; means for analyzing emotions from the video data and the audio data; means for generating text data from the audio data; means for integrating the emotion analysis result and the text data in time series; and means for generating minutes reflecting a flow of emotions based on the integrated data.
[0007] Furthermore, according to one aspect of this disclosure, a recording medium is provided which records a program that causes a computer to function as: means for acquiring video data and audio data; means for analyzing emotions from the video data and audio data; means for generating text data from the audio data; means for integrating the results of the emotion analysis and the text data in chronological order; and means for generating meeting minutes that reflect the flow of emotions based on the integrated data.
[0008] Furthermore, according to one aspect of this disclosure, a real-time customer service support method is provided, characterized in that one or more computers acquire voice data and video data of a customer during a call in real time, analyze the emotions from the voice data and video data, generate a response policy or question based on the results of the emotion analysis, and present the response policy or question to the operator.
[0009] Furthermore, according to one aspect of this disclosure, a real-time customer service support device is provided, comprising: means for acquiring voice and video data of a customer during a call in real time; means for analyzing emotions from the voice and video data; means for generating a response policy or question based on the results of the emotion analysis; and means for presenting the response policy or question to the operator.
[0010] Furthermore, according to one aspect of this disclosure, a recording medium is provided which records a program that causes a computer to function as: a means for acquiring voice and video data of a customer during a call in real time; a means for analyzing emotions from the voice and video data; a means for generating a response policy or question based on the results of the emotion analysis; and a means for presenting the response policy or question to an operator.
[0011] Furthermore, according to one aspect of this disclosure, a voice dialogue method is provided characterized by one or more computers acquiring customer attribute data and product feature data, optimizing voice parameters based on the customer attribute data and product feature data, generating suggested content based on the customer attribute data and product feature data, and synthesizing the suggested content into speech based on the optimized voice parameters.
[0012] Furthermore, according to one aspect of this disclosure, a voice dialogue device is provided that includes means for acquiring customer attribute data and product feature data; means for optimizing voice parameters based on the customer attribute data and product feature data; means for generating proposal content based on the customer attribute data and product feature data; and means for synthesizing the proposal content into speech based on the optimized voice parameters.
[0013] Furthermore, according to one aspect of this disclosure, a recording medium is provided which records a program that causes a computer to function as: means for acquiring customer attribute data and product characteristic data; means for optimizing voice parameters based on the customer attribute data and product characteristic data; means for generating proposal content based on the customer attribute data and product characteristic data; and means for synthesizing the proposal content into speech based on the optimized voice parameters.
[0014] This example of disclosure allows for the development of technologies related to customer service.
[0015] Figure 1 shows an example of a system overview. Figure 2 shows an example of a system function configuration diagram. Figure 3 shows an example of an emotion analysis processing flow. Figure 4 shows an example of a meeting minutes generation process. Figure 5 shows an example of an emotion data model ER (entity relationship diagram). Figure 6 shows an example of a customer profile data structure. Figure 7 shows an example of a voice parameter control model. Figure 8 shows an example of an operator support screen. Figure 9 shows an example of a real-time emotion analysis result display screen. Figure 10 shows an example of a dialogue support processing sequence. Figure 11 shows an example of a voice synthesis parameter control screen. Figure 12 shows an example of an error handling flowchart. Figure 13 shows an example of an operation monitoring screen. Figure 14 shows an example of a system hardware configuration.
[0016] The embodiments of this disclosure will be described below with reference to the drawings. In this disclosure, the drawings are associated with one or more embodiments.
[0017] (Premise) First, we will describe the prerequisites for the system in this embodiment. Emotions play an extremely important role in the human decision-making process. Research in psychology and behavioral economics has shown that approximately 80% of human decisions are based on emotions, with rational judgments coming afterward. In particular, in high-involvement decision-making situations such as purchasing financial products or concluding important contracts, the emotional state of the customer, such as anxiety, expectations, and trust, greatly influences the final decision.
[0018] In traditional face-to-face communication, skilled operators and sales representatives built trust by reading customers' emotional states from their facial expressions and tone of voice and responding accordingly. However, with the spread of online communication and the expansion of remote work, non-face-to-face customer interactions are increasing, making it difficult to respond in a way that captures such subtle emotional nuances.
[0019] Recent research has also highlighted the importance of "emotional transitions" in purchasing decisions. It has become clear that not only the emotional state at a single point in time, but also the process of emotional changes during the conversation—such as shifting from anxiety to reassurance, or from interest to enthusiasm—significantly influences the customer's final decision.
[0020] Against this backdrop, there is a strong demand for systems that can detect customers' emotional states in real time, track their changes, and support optimal responses. In particular, in situations where customer risk perception is crucial, such as when proposing financial products or signing insurance contracts, optimizing explanation methods that take emotional states into consideration is essential. In today's increasingly digital society, complementing the advantages of face-to-face communication with technology and achieving a new understanding of customers through the accumulation and analysis of emotional data is becoming an extremely important element in improving a company's competitiveness.
[0021] However, the technology to capture emotional changes in real time and directly utilize them in purchasing behavior and customer service operations is still underdeveloped, and there were challenges such as the following: Challenge 1: There was no efficient method for creating meeting minutes that reflected the flow of customer emotions when documenting the content of customer interactions. Challenge 2: There was a problem of increased burden on operators during customer interactions and inconsistency in the quality of customer service. Challenge 3: In customer service utilizing generative AI (artificial intelligence), it was difficult to flexibly optimize voice responses to reflect customer attributes and product characteristics, resulting in uniform responses.
[0022] (Overview) Next, an overview of this embodiment will be described. This embodiment relates to a technology that supports communication between customers and operators using an emotion analysis and dialogue control system. Figure 1 is a diagram showing the system overview of this embodiment.
[0023] In Figure 1, the emotion analysis and dialogue control system shown in the center consists of three main engines: an emotion analysis engine, a dialogue control engine, and a speech synthesis engine. The emotion analysis engine analyzes the customer's emotional state in real time from the voice and video data input from the customer. The dialogue control engine determines the optimal dialogue strategy for the interaction with the customer based on the analyzed emotional state. The speech synthesis engine generates an appropriate voice response for the customer based on the determined dialogue strategy.
[0024] Audio and video data are input into the system from the customer shown on the left side of the diagram. This data is processed by an emotion analysis engine, which analyzes the customer's emotional state in real time. The system returns an audio response to the customer based on the analysis of their emotional state.
[0025] The operator, shown on the right side of the diagram, can review the customer's emotional state analysis provided by the system and send control instructions to the system as needed. For example, if the operator determines that a more detailed explanation is needed based on the customer's emotional state, they can adjust the system's response strategy.
[0026] In this way, this disclosure system enables optimal dialogue by constantly monitoring the customer's emotional state while also taking into account the operator's judgment. This allows for detailed responses tailored to the customer's situation, simultaneously achieving improved customer satisfaction and increased operator efficiency.
[0027] (Device Configuration) The functional configuration of this embodiment will be explained with reference to Figure 2. As shown in Figure 2, the system disclosed hereby includes input means, processing means (emotion analysis means, dialogue control means), output means, and storage means.
[0028] The input means include video input means, audio input means, and customer attribute acquisition means.
[0029] The video input means accepts video input of a customer's facial expression (face). For example, the customer communicates with the system (by making a call, chatting, etc.) by operating the system's input device or a terminal device connected to the system in a communicative manner. The video input means may be equipped with a camera that captures the customer's face and acquires the image generated by the camera. Alternatively, the video input means may receive an image of the customer's face generated by the camera of a terminal device operated by the customer. The video input means can acquire the video in real time.
[0030] The voice input means collects customer speech. The voice input means may be equipped with a microphone, which collects the customer's speech and generates audio data. Alternatively, the voice input means may receive audio data from a terminal device operated by the customer, which contains audio data of the customer's speech collected by the microphone of the terminal device. The voice input means can acquire this audio data in real time.
[0031] The customer attribute acquisition means acquires attribute information such as customer profile information and risk preference. In one example, attribute information including at least one of the following—customer profile information, risk preference, and hobbies and preferences—is pre-stored in the storage means. The customer attribute acquisition means can read the customer profile information and attribute information such as risk preference and hobbies and preferences from the storage means.
[0032] The core of the processing system, the emotion analysis system, consists of a facial expression analysis unit, a voice emotion analysis unit, and an emotion integration processing unit.
[0033] The facial expression analysis unit analyzes the customer's facial expressions from the video. The unit analyzes the customer's facial image (video) to recognize their facial expressions and can estimate the customer's emotion (primary emotion) based on the recognized facial expressions. The unit can also estimate the intensity of the emotion (emotion score). This estimation of the customer's emotion can be achieved using widely known image analysis techniques. For example, the use of deep learning models is one example, but it is not limited to this.
[0034] The voice emotion analysis unit extracts emotional features (voice features) from the voice. Based on the emotional features extracted from the voice, the voice emotion analysis unit can estimate the customer's emotions (secondary emotions). This estimation of the customer's emotions can be achieved using widely known voice emotion analysis techniques. Furthermore, the voice emotion analysis unit can also estimate the intensity of the emotions (emotion score). This estimation can also be achieved using widely known voice analysis techniques. For example, the use of deep learning models is one example, but it is not limited to this.
[0035] Examples of emotions estimated by the facial expression analysis unit and the voice emotion analysis unit include anxiety, satisfaction, interest, anger, fear, surprise, sadness, joy, acceptance, disgust, expectation, gratitude, peace, hope, pride, amusement, encouragement, awe, and normal.
[0036] The emotion integration processing unit integrates the analysis results (first emotion and second emotion) from the facial expression analysis unit and the voice emotion analysis unit to estimate the overall emotional state. There are various methods of integration. For example, the analysis results from the facial expression analysis unit and the voice emotion analysis unit show the confidence level (numerical value) for each of the multiple emotions. In one example, the emotion integration processing unit calculates statistical values of the confidence level calculated by the facial expression analysis unit and the confidence level calculated by the voice emotion analysis unit for each type of emotion. The statistical values are, but are not limited to, the mean, weighted mean, maximum, and minimum values. When calculating the weighted mean, weight values for the confidence levels of the emotions calculated by the facial expression analysis unit and the confidence levels of the emotions calculated by the voice emotion analysis unit are defined in advance. The emotion integration processing unit can calculate the weighted mean using these weight values. For example, after calculating statistical values for each type of emotion, the emotion with the largest statistical value can be estimated as the customer's emotion at that time.
[0037] The dialogue control means includes a response policy generation unit, an additional question generation unit, and a voice parameter control unit.
[0038] The response strategy generation unit determines an appropriate dialogue strategy (response strategy) for customer interaction based on the analysis results of the customer's emotions. The process of generating the response strategy includes detecting when a change in the customer's emotional state meets predetermined attention conditions and generating a response strategy in response to such detection. The attention conditions include when the customer's emotional state, expressed numerically, changes at a speed exceeding a predetermined threshold. The customer's emotional state, expressed numerically, is, for example, the emotional value of each emotion (the confidence level or intensity of each emotion).
[0039] Here, we will explain an example of a dialogue strategy (response strategy) generated by the response strategy generation unit. For example, if the response strategy generation unit estimates that the customer is in a state of anxiety, it will generate a dialogue strategy that recommends providing detailed explanations and specific examples to reassure the customer. If the customer shows interest during the meeting, the response strategy generation unit may generate a dialogue strategy that provides relevant information to further deepen that interest or encourages questions. Also, if the customer shows feelings of confusion or anger, it may generate a step-by-step dialogue strategy that first shows empathy and then makes suggestions for solving the problem.
[0040] This section describes an example of the specific process by which the response strategy generation unit generates a dialogue strategy. The response strategy generation unit may select the optimal response pattern from a predefined response pattern database, taking the emotion category (e.g., anxiety, interest, anger, etc.) and emotion intensity obtained as a result of emotion analysis as input.
[0041] The database contains combinations of detected emotion categories, emotion intensity, and product category and phase information (product description stage, price presentation stage, contract stage, etc.) as keys, and stores specific response templates, recommended speech patterns, example utterances, recommended question lists, expressions to avoid lists, and supplementary materials to refer to as values. For example, for the key "anxiety, intensity, investment product, risk explanation stage," the database might store response templates such as "emphasize long-term stability while showing specific past market fluctuation examples" and "explain in the format of 'even if XX happens, there is a countermeasure of YY'."
[0042] Furthermore, when generating an operator's response policy in accordance with emotions in product sales, a machine learning model may be used if past dialogue data, contract results, and other information have been accumulated. For example, by learning the response pattern with the highest contract closure rate for each combination of emotional state and customer attribute, a more accurate response policy may be automatically generated. As the machine learning model, various models such as decision trees, random forests, support vector machines, and neural networks can be applied; time series models that can consider dialogue context and graph-based models are examples, but the model is not limited to these.
[0043] The additional question generation unit generates questions to be presented to the operator based on the analysis result of the customer's emotion. The additional question generation unit may determine the necessity of a question based on the customer's emotional state. Then, the additional question generation unit may generate a question when it is determined that a question is necessary.
[0044] The additional question generation unit may determine the necessity of a question based on the recognition result of the customer's emotional state. A plurality of factors may be considered in the specific determination of question necessity. For example, when the rate of change of the customer's emotional state exceeds a threshold (e.g., when the emotion abruptly changes from "interested" to "anxious"), the additional question generation unit may determine that a question is necessary to accurately grasp the situation. Furthermore, when ambiguity or contradiction is detected in the customer's utterance content, or when there is little emotional response after an explanation regarding a specific product, the additional question generation unit may determine that it is necessary to clarify the customer's intention through an additional question. As another example, in order to determine question timing that follows the natural flow of dialogue, a certain period of silence after the customer's utterance, breaks in product explanation, and the like may also be used as determination criteria for question necessity. Such predetermined comprehensive condition evaluation achieves optimal question generation that does not result in excessively frequent questions and does not miss important timing for information collection.
[0045] This section describes an example of the specific actions taken by the additional question generation unit when it determines that a question is necessary. The additional question generation unit may refer to a question template database that uses emotion category and emotion intensity as keys and select an appropriate question pattern. This database contains question templates corresponding to each emotional state and is further subdivided by product category and dialogue phase (initial contact, during product description, after price presentation, etc.). For example, for the keys "satisfied, weak, insurance product, after explanation," the database may contain questions that explore related needs, such as "Are there any other family members who are considering purchasing insurance?"
[0046] Furthermore, question generation may utilize not only the template-based method described above, but also question generation techniques using machine learning models. For example, the additional question generation unit may employ language generation methods that leverage pre-trained language models such as the Transformer architecture or BERT (Bidirectional Encoder Representations from Transformers). By using such natural language processing models, it becomes possible to dynamically generate more natural and situation-appropriate questions by learning the context of the dialogue, the customer's emotional state, and successful patterns from similar past cases. For example, by building a model that predicts the next optimal question using the dialogue history and sentiment analysis results as input, it becomes possible to flexibly handle a variety of dialogue situations that cannot be adequately addressed by standard templates.
[0047] Further, the additional question generation unit may generate a question based on the response policy generated by the response policy generation unit. For example, when the response policy generation unit generates a policy of "identifying potential anxiety factors of the customer and presenting specific safety measures" as the response policy, the additional question generation unit generates a sequence of stepwise questions. Specifically, the additional question generation unit may first start with a broad question such as "Is there any point that you are particularly concerned about regarding XX?". Then, the additional question generation unit may output a question including a specific solution such as "There is a countermeasure such as ΔΔ for that point, what do you think about that?" in accordance with the answer. Alternatively, when the response policy generated by the response policy generation unit includes an element of "making a proposal in accordance with the customer's future plan", the additional question generation unit may propose a question including a medium- to long-term perspective such as "Do you plan to have any change in your lifestyle in the next five years?".
[0048] The voice parameter control unit optimizes characteristics (voice parameters) of response voice output to the customer in accordance with the customer's emotional state or the like. For example, the voice parameter control unit can adjust the tone and speed of the response voice in accordance with the customer's emotional state or the like. Details of voice parameter control will be described later.
[0049] The output means includes a minutes generation unit, an operator support display unit, and a voice synthesis output unit.
[0050] The minutes generation unit creates minutes that integrate dialogue content and emotion analysis results. Specifically, the minutes generation unit performs voice analysis on voice data to generate character data. Transcription of voice data can be realized by using widely known techniques. Then, the minutes generation unit integrates the analysis result of the customer's emotion and the character data indicating the customer's utterance content in time series. Said integration includes a process of associating, for each utterance text included in time-series character string data, the analysis result of the customer's emotion at the utterance timing of each utterance text, to generate time-series integrated data. The minutes generation unit can generate minutes reflecting the flow of the customer's emotion based on such time-series integrated data. Details of minutes generation will be described later.
[0051] The operator support display unit presents operators with various information generated based on customer emotions and other factors. For example, the operator support display unit presents operators with the results of customer emotion analysis, text indicating the content of customer speech, and recommended actions determined by the system. An example of a screen displayed to the operator will be described later.
[0052] The speech synthesis output unit generates response speech based on parameters optimized by the speech parameter control unit. Details of speech synthesis will be described later.
[0053] The memory system includes an emotion database and a customer profile database, and permanently stores analysis results and customer information. This data will be used as reference information in future interactions. Details of the data stored in the memory system will be described later.
[0054] (Processing Flow) Next, we will explain the sentiment analysis processing sequence diagram in Figure 3. Figure 3 shows the series of processing flows in chronological order, from processing the video and audio data input from the customer to outputting the sentiment analysis results.
[0055] First, the customer sends video and audio data to the input processing unit. The input processing unit then forwards the received data to the video processing unit and the audio processing unit, respectively.
[0056] The video processing unit and the audio processing unit perform processing in parallel. The video processing unit first extracts feature points of facial expressions from the customer's facial image, and then performs facial expression classification. This result is sent to the emotion analysis unit as a facial expression score. The facial expression score indicates, for example, the confidence that the customer's facial expression corresponds to each of several types of emotions. Meanwhile, the audio processing unit extracts features from the audio and performs emotion pattern analysis. This result is sent to the emotion analysis unit as an audio emotion score. The audio emotion score indicates, for example, the confidence that the customer's voice corresponds to each of several types of emotions.
[0057] The emotion analysis unit integrates the received facial expression score and voice emotion score to estimate the overall emotional state. The emotion analysis unit also evaluates the reliability of the estimation results. The emotion analysis unit can save the analysis results to a database. Furthermore, the emotion analysis unit can transmit the analysis results to the output processing unit.
[0058] Although not shown in Figure 3, the output processing unit can display the received analysis results to the operator in real time. Furthermore, the output processing unit can record the data as time-series data in a database. This allows the operator to understand the customer's emotional state in real time. The time-series data stored in the database can be used for subsequent analysis.
[0059] Each processing step handles specific data such as sentiment scores, confidence levels, and time information, and this information is systematically managed.
[0060] (Meeting Minutes Generation Process) The meeting minutes generation process performed by the meeting minutes generation unit will be explained using Figure 4. As shown in Figure 4, the meeting minutes generation process consists of five main processing stages. First, audio data, sentiment analysis results, and time information are acquired as input data.
[0061] In the next text processing stage, the minutes generation unit converts the audio data into text data using speech recognition technology. At this stage, in addition to accurately capturing the spoken content as text, speaker identification information is also added. Speaker identification is achieved using widely known technology. For example, the speaker of each utterance can be identified based on the voice characteristics of each speaker that have been registered in advance.
[0062] In the subsequent emotional information integration stage, the minutes generation unit correlates text data with emotional analysis results in chronological order. Each utterance is associated with the customer's emotional state at that time (e.g., level of interest, level of anxiety, etc.), and the data is organized into a format that allows for understanding the changes in emotions within the context of the dialogue.
[0063] During the document structuring stage, the minutes generation unit structures the document as minutes based on the integrated text and sentiment information. Specifically, the minutes generation unit divides the dialogue into sections and assigns appropriate headings. In this process, the minutes generation unit may also consider points of change in sentiment when structuring the document. The assignment of headings and section division are achieved using widely known technologies. For example, generative AI, such as the Transformer architecture or pre-trained language models like BERT, may be used. The desired processing is achieved by including instructions to consider points of change in sentiment in the prompts input to the generative AI.
[0064] During the summary generation process, the minutes generation unit extracts key points from the structured data and generates a summary that includes emotional tendencies. Particular emphasis is placed on extracting sections where emotional values change significantly or where important business decisions were made. Summary generation is achieved using widely known technologies. For example, generation AI may be used. By including instructions in the prompts input to the generation AI to focus on extracting sections where emotional values change significantly or where important business decisions were made, the desired extraction can be achieved.
[0065] In the final meeting minutes output, the minutes generation unit integrates these processing results to generate structured minutes that include a textual record of the dialogue and the summary based on the customer sentiment analysis results. By creating the summary based on data that integrates the customer sentiment analysis results and the textual data showing the content of the utterances in chronological order, it is possible to generate meeting minutes that reflect the flow of the customer's emotions. These meeting minutes can be used as reference information for subsequent business negotiations and customer interactions.
[0066] Each processing stage is executed in the order indicated by the arrows, and the quality of the meeting minutes is gradually improved by referring to the results of the previous stage as needed. In addition, the results of each stage are stored in a storage device and managed in a form that can be reused as needed.
[0067] (Data Structure) Figure 5 is an ER diagram showing the sentiment data model in this embodiment. This data model consists of four main entities: session, sentiment analysis results, customer, and sentiment statistics.
[0068] A session entity manages the basic unit of interaction. It uses a session ID (identifier) as its primary key and includes attributes such as a customer ID (foreign key), start time, end time, and session type. This allows for the unique identification and management of each individual interaction session.
[0069] The sentiment analysis results entity records real-time sentiment analysis results. It uses a sentiment analysis ID as its primary key and includes attributes such as a session ID (foreign key), timestamp, sentiment category, sentiment intensity, and confidence level. This allows for detailed tracking of the time-series changes in emotional state during a session.
[0070] The customer entity manages the basic information of each customer. It uses a customer ID as the primary key and includes attributes such as customer attribute IDs (foreign keys), basic information, risk preference, hobbies and preferences, transaction history reference, and last updated date. This allows for the appropriate management of individual customer characteristics and historical information. Basic information includes the customer's ID, name, age, address, etc. Risk preference indicates a tendency to prefer high-risk transactions. Transaction history shows the history of each customer's past transactions (buys and sells).
[0071] The emotion statistics entity statistically aggregates the results of emotion analysis for the entire session. It uses a statistical ID as its primary key and has attributes such as a session ID (foreign key), major emotion tendency, emotion change pattern, important event records, and aggregation period. This allows for analysis of long-term emotion trends and comparisons between sessions. The major emotion tendency indicates the customer's dominant emotion during each session. For example, the major emotion tendency may be the emotion that appeared for the longest period of time during each session. Alternatively, the major emotion tendency may be at least one emotion that appeared for a predetermined period of time or longer during each session. The emotion change pattern indicates the pattern of emotion changes that occurred during each session. The emotion change pattern is shown as a pair of emotions before and after the change, such as "anger → calm." The important event records indicate important events that occurred during each session and the timing of their occurrence. Important events include, for example, the start of product description and the start of risk explanation. For example, such important events can be detected by detecting predetermined keywords from the dialogue history.
[0072] The relationships between entities are indicated by arrows, representing one-to-many relationships. For example, multiple sentiment analysis results can be linked to a single session, and similarly, multiple sessions can be associated with a single customer.
[0073] This data model allows for the persistent storage of sentiment analysis results in a structured format, enabling multifaceted analysis and reference as needed. Furthermore, by linking it with customer information, it supports the realization of detailed responses tailored to individual customer characteristics.
[0074] Furthermore, the data structure of the customer profile will be explained using Figure 6. The customer profile in this system consists of four tables: basic customer information, investment profile, communication history, and product preferences.
[0075] The customer basic information table manages basic attribute information such as the customer ID, which serves as the primary key for uniquely identifying customers throughout the system, as well as name, age, and contract type. This information is used as basic reference information during interactions.
[0076] The investment profile table manages the characteristics of a customer's investments. Risk tolerance is expressed as an integer value from 1 to 5, for example, and along with years of investment experience and desired investment period, it is an important factor in making decisions when proposing products. This table holds the customer ID as a foreign key and is linked to the customer basic information table.
[0077] The communication history table manages the sentiment analysis results of past interactions in chronological order. Each interaction is identified by a unique history ID, and the sentiment score is recorded along with the date and time of the interaction. This history data serves as important reference information when determining the strategy for interacting with customers.
[0078] The product preference table manages the customer's level of interest in each product category using integer values from 1 to 5. This information, along with the last update date, is used to optimize recommendations. While product categories vary, examples of investment products include stocks, bonds, mutual funds, foreign currency deposits, life insurance, government bonds, real estate, and commodities.
[0079] These tables are interconnected via customer IDs, enabling consistent customer profile management across the entire system. The data types of each table are appropriately configured according to the characteristics of the information they store.
[0080] (Voice Parameter Control) The voice parameter control model will be explained using Figure 7. This model consists of three main parts: input elements, control logic, and output parameters.
[0081] The voice parameter control unit receives customer emotional state (anxiety, satisfaction, interest, etc.) and risk tolerance (high, medium, low) as input elements. Emotional state is mainly obtained from the results of facial expression analysis and voice emotion analysis, while risk tolerance is obtained from the customer profile.
[0082] The control logic consists of three logics: tone adjustment, speed adjustment, and emotion expression adjustment. In the tone adjustment logic, the voice parameter control unit selects a calm tone if the customer is showing anxiety, a bright tone if they are showing interest, and a gentle tone if they are satisfied. In the speed adjustment logic, the voice parameter control unit adjusts the speaking speed according to the risk tolerance level, providing slower explanations to customers with low risk tolerance and faster explanations to customers with high risk tolerance. In the emotion expression adjustment logic, the voice parameter control unit also controls the intensity of emotion expression based on risk tolerance. Figure 7 shows an example of the control content of each logic.
[0083] The speech parameter control unit can generate three numerical parameters as output parameters: speech tone (range 0.5-2.0), speech rate (range 0.8-1.5), and emotion intensity (range 0.0-1.0). These parameters are input to the speech synthesis engine to determine the characteristics of the final speech output.
[0084] Furthermore, this processing may be performed depending on the type of product mentioned in the conversation with the customer. The mentioned product can be identified, for example, by analyzing the content of the conversation (keyword detection, etc.). The voice parameter control unit may, for example, adjust the speech speed according to the customer's risk preference if a risk product is mentioned in the conversation with the customer. In one example, product feature data indicating the characteristics of each product is registered in the system in advance. The product feature data includes at least one of the product's risk level and selling points. The voice parameter control unit can read (acquire) the product feature data. The voice parameter control unit can then identify products with a risk level above a threshold as risk products.
[0085] Furthermore, the voice parameter control unit may also adjust voice parameters based on customer preferences and product selling points. For example, if the customer's preferences data identifies them as a "classical music lover," a more refined and calm speaking style (lower pitch, moderate intonation, relaxed tempo) may be selected and reflected during speech synthesis. Similarly, for an attribute such as "sports spectator," a more lively speaking style (brighter tone, slightly faster tempo, increased emphasis) may be selected.
[0086] Regarding the selling points of a product, for example, if "safety" is the main selling point, a calm voice tone that fosters trust can be combined with emphasis on keywords (a slight increase in power). For products where "innovation" is the selling point, a slightly higher pitch and energetic speaking style can be chosen to emphasize novelty and potential through vocal expression. For products emphasizing "luxury," adjustments can be made to apply a relaxed speaking pace and refined intonation patterns.
[0087] By tailoring the parameters of the voice output directed at customers to their personality traits, customer satisfaction and purchase intent can be increased. For example, an active and sociable approach may be effective for extroverted customers, while a more reserved approach emphasizing detailed information may be effective for introverted customers. Furthermore, the use of emotional language in product descriptions, for instance, can increase purchase intent for entertainment products.
[0088] This model makes it possible to generate optimized voice responses based on the customer's status and profile. In particular, when explaining risk products, it can adapt the explanation style to the customer's risk tolerance.
[0089] Furthermore, the system can generate suggestions based on customer attribute data and product characteristic data. Specifically, the system can match customer attribute data (age, occupation, family structure, asset status, risk tolerance, investment experience, hobbies and preferences, etc.) with product characteristic data (risk characteristics, investment period, liquidity, institutional characteristics, required funds, etc.) and select product candidates that are suitable for the customer's situation through a matching process.
[0090] For example, for a customer with the attributes of "low risk tolerance and short-to-medium-term funding needs," the system can generate proposals that include information centered on "Financial Product A, which is highly stable and suitable for medium-term asset management." On the other hand, for a customer with the attributes of "high risk tolerance and experience in long-term investment," the system can generate proposals that include information on "Financial Product B, which is suitable for long-term asset building."
[0091] In the proposal generation process, the system quantifies the degree of fit between customer attributes and product characteristics. From the product group whose degree of fit exceeds a threshold, it further considers the customer's past transaction history and the preferences of other customers with similar attributes to narrow down the final proposal candidates. The system can then assemble the selected product's features, anticipated benefits, risk factors to consider, and appropriate usage methods into a structured proposal.
[0092] Furthermore, any technology can be used to generate the proposed content, including methods based on predefined templates and rules, methods utilizing natural language processing technology for text generation, or methods using machine learning models learned from past successful cases.
[0093] The speech synthesis output unit can then synthesize the proposed content into speech based on optimized speech parameters. In one example, the speech synthesis output unit can perform speech synthesis using a speech synthesis engine, as shown in Figure 7. The speech synthesis output unit can then output the generated speech to the customer (voice response). Speech synthesis of proposed content means generating speech data that speaks the proposed content based on optimized speech parameters. The speech tone, speaking speed, emotional intensity, etc., of the generated speech data are determined according to the optimized speech parameters.
[0094] (UI (user interface) screen) The operator support screen of this embodiment will be explained using Figure 8. This screen is designed to provide operators with a highly visible overview of the information they need to efficiently handle customer interactions. The screen is composed of six main areas.
[0095] The top header area displays the system name and the name of the assigned operator. This allows the login status to the system to be constantly monitored.
[0096] In the customer basic information area in the upper left, in addition to basic attributes such as the ID, name, and age of the customer being handled, important information for product recommendations, such as risk tolerance and transaction history, is displayed.
[0097] In the sentiment analysis results area at the top center, the customer's emotional state is displayed in real time as a graph. By visualizing the changes in the emotional values of various emotions (trust level and intensity of each emotion) over time, it is possible to understand emotional changes in the flow of conversation. Furthermore, by indicating the timing of event occurrences (product description, risk explanation, etc.) within the graph, it is possible to understand the relationship between events and emotions. The timing of event occurrences can be detected, for example, by analyzing the content of the conversation (keyword detection, etc.). In addition, the figure shows the frequency of occurrence of various emotions within a short period of time as a bar graph.
[0098] In the recommended action area in the upper right, suggested response strategies and questions generated based on the sentiment analysis results are presented. This allows the operator to select the appropriate response for the situation.
[0099] In the dialogue history area at the bottom left, the system and customer interactions are displayed chronologically. Each utterance is accompanied by time information, allowing for an accurate understanding of the flow of the conversation.
[0100] In the product information area in the lower right, detailed information about the products mentioned during the conversation is displayed for easy reference. Important product characteristics such as risk level and recommended investment period can be instantly viewed.
[0101] By appropriately arranging this information, operators can accurately understand the customer's situation and respond efficiently. Furthermore, a consistent design is adopted across the entire screen, ensuring visibility even during long work sessions.
[0102] Next, we will explain the real-time sentiment analysis results display screen shown in Figure 9. This screen is a support interface that allows operators to check the sentiment analysis results in real time during conversations with customers and take appropriate action. The screen is composed of three main areas.
[0103] In the upper left corner, the current emotional state area displays the most recent analysis results: the primary emotion, its confidence level, and the measurement time. If the primary emotion is a predefined attention-grabbing state such as "anxiety," the information is highlighted in red or another color to draw the operator's attention.
[0104] In the upper right emotional value trend area, the temporal changes in three emotional values—positive, negative, and neutral—are visualized in a line chart. Various emotions are pre-classified into positive, negative, and neutral. The temporal changes in the emotional values (trust level and intensity of each emotion) for each category are then displayed graphically. If there are multiple types of emotions within a category, the statistical values of the emotional values of all those types can be used as the emotional value for each category. For example, the lines for each category are displayed in different colors, making it easy to grasp the trend of change over time. Grid lines are displayed on the graph to assist in reading the values. Furthermore, by indicating the timing of event occurrences (product description, risk explanation, etc.) within the graph, the relationship between events and emotions can be understood.
[0105] Important notification information is located at the bottom of the screen. If a situation requiring attention is detected, a red alert will be displayed. For example, if a rapid change in emotional state (emotional value changing faster than a threshold) is detected, the situation and recommended course of action will be presented. In addition, the important events history area records important events during the conversation in chronological order, helping to understand the context of emotional changes.
[0106] In this way, this screen is designed to present the sentiment analysis results clearly from multiple perspectives, supporting operators in making quick and appropriate decisions.
[0107] (Use Case) Figure 10 shows the time-series processing sequence of dialogue support in this system. The system achieves a series of dialogue controls through the interaction of six main actors: customer, input processing unit, sentiment analysis engine, dialogue control engine, operator, and speech synthesis engine.
[0108] One cycle of dialogue begins with voice and video input from the customer. The input processing unit receives this data and transfers it to the emotion analysis engine. The emotion analysis engine sequentially performs three stages of processing: facial expression analysis, voice emotion analysis, and emotion integration processing, to generate a comprehensive emotion analysis result.
[0109] The analysis results are transferred to the dialogue control engine, which first displays the customer's emotional state to the operator. The dialogue control engine generates a response plan based on the analysis results, but if the customer's emotional state exceeds a predetermined threshold, it notifies the operator of an alert. In this case, the operator can modify the response plan as needed.
[0110] Once the response strategy is determined, the dialogue control engine generates specific response content, sets appropriate voice parameters, and transmits them to the speech synthesis engine. The speech synthesis engine performs speech synthesis processing based on these parameters and outputs the final voice response to the customer.
[0111] Thus, this system achieves flexible and effective dialogue support by combining automated dialogue control based on sentiment analysis with operator intervention as needed. Furthermore, by clearly indicating the activation status at each processing stage, the system's operating status can be clearly understood.
[0112] (Voice Parameter Control Screen) Figure 11 shows the control screen for adjusting speech synthesis parameters. This screen is presented to the operator. The operator can make various settings on this screen. This screen consists of two panels, left and right. The left panel is used for detailed parameter settings, and the right panel is used for preview and control.
[0113] In the basic parameter setting area on the left panel, key parameters related to speech synthesis can be adjusted using sliders. Specifically, operators can intuitively control four parameters: voice pitch (70%), speech rate (60%), emotional expression (50%), and clarity (65%). In addition, operators can easily recall frequently used settings by operating the preset selection area located at the bottom of the screen.
[0114] The preview and control area on the right panel allows for immediate confirmation of the effects of the set parameters. The audio waveform is visually displayed at the top, with basic control buttons such as play and stop located below. Furthermore, a preview text input field is provided. By manipulating this input field, the operator can check the speech synthesis results with any text they choose.
[0115] The operator can apply the adjusted settings to the actual operating environment by pressing the "Apply Parameters" button at the bottom of the screen. This screen layout allows operators to intuitively optimize audio parameters and immediately see the effects.
[0116] (Error Handling) Next, the error handling flow of this system will be explained using Figure 12. Figure 12 is a flowchart showing the main error cases that may occur in the system and how to deal with them. When an error is detected, the type of error is determined first, and based on the result, the appropriate response flow is executed.
[0117] In this example system, errors are broadly classified into three types. First, there are data loss errors, which occur when some video or audio data is missing. In this case, the system performs processing using substitute data. Specifically, the system performs imputation using the most recent valid data or estimates the missing value using statistical methods.
[0118] Secondly, there are system errors, which are detected when problems occur in core system functions such as the sentiment analysis engine or the dialogue control engine. In this case, the system performs a system restart process. During the restart, the system takes measures to securely save data being processed and maintain system integrity.
[0119] Thirdly, there are communication errors, which occur when data cannot be sent or received properly due to network connection problems. In this case, the system performs a communication reconnection process. During the reconnection process, the system re-establishes the communication protocol and controls the retransmission of any unsent data.
[0120] After each error is processed, the system evaluates the results, confirms that the system has recovered successfully, and terminates the process. At each stage of error processing, the operator is notified in parallel, and the system is designed to allow for manual intervention as needed.
[0121] Thus, this system is equipped with appropriate countermeasures for various error cases, ensuring stable operation and reliability.
[0122] (Operation Monitoring Screen) The configuration of the operation monitoring screen will be explained using Figure 13. This system's operation monitoring screen is designed to allow system administrators to grasp the system status in real time and take necessary actions quickly. The screen is composed of four main sections.
[0123] At the top of the screen are three panels that provide an overview of the system status. The left panel visually displays the system's operational status, the center panel shows resource usage such as CPU and memory with a progress bar, and the right panel displays the current number of active sessions and the status of pending processes.
[0124] A graph showing the system performance over time is located in the center of the screen. This graph allows you to understand the changes in system performance over time, which can be useful for the early detection of abnormal fluctuation patterns.
[0125] At the bottom of the screen, a chronological history of alerts generated by the system is displayed. Each alert entry records the time of occurrence and its specific details, and high-priority issues are visually highlighted by changing the background color according to their importance.
[0126] This information is automatically updated, ensuring that the latest status is always displayed. Furthermore, the information in each section is interconnected, allowing for analysis such as linking performance degradation with the occurrence of alerts.
[0127] In this way, this monitoring screen efficiently aggregates the information necessary for operational management, enabling continuous monitoring of the system's health.
[0128] (Effects of this embodiment) The effects of this embodiment will now be explained. The following effects can be expected from the introduction of this system.
[0129] Firstly, it improves the quality of customer service. Real-time sentiment analysis by the system and optimal dialogue control based on that analysis enable meticulous responses tailored to the customer's emotional state. This leads to increased customer satisfaction and a reduction in the rate of complaints.
[0130] Secondly, the system improves the efficiency of operators' work. By visualizing the results of sentiment analysis and having the system suggest response strategies, operators can accurately understand the customer's state and respond efficiently. This makes it possible to increase the number of cases handled per person and reduce the response time.
[0131] Thirdly, standardization of dialogue quality is achieved. Consistent sentiment analysis and response strategy generation by the system enable uniform customer service that is independent of the operator's experience and skills. This contributes particularly to the rapid development of new operators.
[0132] Fourthly, it promotes a deeper understanding of customers. By automatically generating meeting minutes that integrate dialogue content and sentiment analysis results, it becomes possible to accurately record and analyze the history of customer interactions, including emotional aspects. This leads to a deeper long-term understanding of customers and enables the formulation of more appropriate product proposals and response strategies.
[0133] Fifth, system operation becomes more efficient. Integrated operational monitoring functions allow for constant monitoring of the system's status, enabling early detection and rapid response to problems. This results in stable system operation and reduced operating costs.
[0134] Based on the above effects, this system is expected to function as an effective customer service support system that simultaneously improves customer satisfaction and operational efficiency. Furthermore, the accumulated dialogue history and sentiment analysis data can be used to further improve customer service and develop new products and services.
[0135] Although this embodiment uses the example of selling financial products for explanation, the scope of application of this embodiment is not limited to this. This embodiment can be used in various dialogue scenarios where the emotional state of the customer or consultant influences decision-making and satisfaction, such as the sale of daily necessities and home appliances, proposals for travel and entertainment services, contracts for insurance and real estate, interview support for job hunting and career changes, medical and health consultations, and educational and guidance settings.
[0136] (Minimum configuration of meeting minutes generation technology) The minimum configuration for realizing the meeting minutes generation technology disclosed here includes the following means: • Means for acquiring video and audio data • Means for analyzing emotions from video and audio data • Means for generating text data from audio data • Means for integrating the emotion analysis results and text data in chronological order • Means for generating meeting minutes that reflect the flow of emotions based on the integrated data
[0137] A device equipped with such means can generate time-series data showing the content of speech and the speaker's emotions at the time each speech was uttered. Based on this time-series data, the device can generate meeting minutes that reflect the flow of the speaker's emotions. By utilizing such meeting minutes, technologies related to customer service can be advanced.
[0138] (Minimum configuration of real-time customer service support technology) The minimum configuration for realizing the real-time customer service support technology disclosed here includes the following means: • Means for acquiring customer voice and video data in real time during a call • Means for analyzing emotions from the voice and video data • Means for generating response policies or questions based on the results of the emotion analysis • Means for presenting the response policies or questions to the operator
[0139] A device equipped with such means can identify the speaker's emotions in real time and generate response strategies and questions based on the results. This device, capable of performing real-time processing in response to customer emotions, can advance customer service technologies.
[0140] (Minimum configuration of voice dialogue technology) The minimum configuration for realizing the voice dialogue technology of this disclosure includes the following means: • Means for acquiring customer attribute data and product feature data • Means for optimizing voice parameters based on customer attribute data and product feature data • Means for generating proposal content based on customer attribute data and product feature data • Means for synthesizing the proposal content into speech based on the optimized voice parameters
[0141] A device equipped with such means can optimize the parameters of the voice output to the customer based on the customer's attributes and the characteristics of the product being mentioned. This device can advance customer service technology.
[0142] <Hardware Configuration> Next, an example of the hardware configuration of the system disclosed will be described. Each functional part of the system disclosed is realized by any combination of hardware and software. It will be understood by those skilled in the art that there are various variations in the implementation method and apparatus. The software includes programs that are pre-installed at the time of shipment of the device, as well as programs downloaded from recording media such as CDs (Compact Discs) or from servers on the Internet.
[0143] Figure 14 is a block diagram illustrating the hardware configuration of the system of this disclosure. As shown in Figure 14, the system of this disclosure includes a processor 1A, memory 2A, input / output interface 3A, peripheral circuitry 4A, and bus 5A. Peripheral circuitry 4A includes various modules. The system of this disclosure does not necessarily have peripheral circuitry 4A. The system of this disclosure may consist of multiple physically and / or logically separated devices. In this case, each of the multiple devices may have the above hardware configuration.
[0144] Bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuit 4A, and input / output interface 3A to send and receive data to and from each other. The processor 1A is a processing unit such as a CPU (Central Processing Unit) or GPU (Graphics Processing Unit). Memory 2A is a memory such as RAM (Random Access Memory) or ROM (Read Only Memory). The input / output interface 3A includes interfaces for acquiring information from input devices, external devices, external servers, external sensors, cameras, etc., and interfaces for outputting information to output devices, external devices, external servers, etc. The input / output interface 3A also includes interfaces for connecting to a communication network such as the Internet. Input devices include, for example, a keyboard, mouse, microphone, physical buttons, touch panel, etc. Output devices include, for example, a display, projection device, speaker, printer, mailer, etc. The processor 1A can issue commands to each module and perform calculations based on their calculation results.
[0145] Although this disclosure has been described above with reference to embodiments, this disclosure is not limited to the embodiments described above. Various modifications to the structure and details of this disclosure are possible, which can be understood by those skilled in the art within the scope of this disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0146] Furthermore, the sequence diagrams used in the above explanation show multiple processes (processes) in order. However, the execution order of the processes performed in each embodiment is not limited to the order in which they are described. In each embodiment, the order of the illustrated processes can be changed to the extent that it does not impede the content.
[0147] Some or all of the above embodiments may also be described as follows, but are not limited to the following: 1. A method for generating meeting minutes, characterized in that one or more computers acquire video data and audio data, analyze emotions from the video data and audio data, generate text data from the audio data, integrate the results of the emotion analysis and the text data in chronological order, and generate meeting minutes that reflect the flow of emotions based on the integrated data. 2. A method for generating meeting minutes according to 1, characterized in that the analysis of emotions involves recognizing facial expressions from the video data to estimate a first emotion, analyzing audio features from the audio data to estimate a second emotion, and integrating the first and second emotions to determine the emotional state. 3. 1. A method for generating meeting minutes according to 1 or 2, wherein integrating the emotion analysis results and the character data in chronological order includes linking the emotion analysis results at the timing of each utterance of the utterance of each utterance to each utterance text included in the chronological string data, thereby generating the chronologically integrated data. 4. A method for generating meeting minutes according to any one of 1 to 3, wherein the video data and the audio data are acquired in real time. 5. A real-time customer service support method characterized in that one or more computers acquire audio data and video data of a customer during a call in real time, analyze emotions from the audio data and the video data, generate a response policy or question based on the emotion analysis results, and present the response policy or question to the operator. 6. A real-time customer service support method according to 5, wherein generating the response policy includes detecting that a change in the customer's emotional state meets predetermined attention conditions, and generating the response policy in response to the detection. 7. In the real-time customer service support method described in 6, the cautionary condition includes the real-time customer service support method in which the emotional state of the customer, expressed numerically, changes at a speed exceeding a predetermined threshold.8. A real-time customer service support method according to any one of 5 to 7, wherein generating the question includes determining the necessity of the question based on the emotional state, and generating the question when it is determined that the question is necessary. 9. A voice dialogue method characterized in that one or more computers acquire customer attribute data and product feature data, optimize voice parameters based on the customer attribute data and product feature data, generate proposal content based on the customer attribute data and product feature data, and synthesize the proposal content into voice based on the optimized voice parameters. 10. A voice dialogue method according to 9, wherein the customer attribute data includes at least one of the customer's risk preference and hobbies and preferences, and the product feature data includes at least one of the product's risk level and selling points. 11. A voice dialogue method according to 9 or 10, wherein optimizing the voice parameters includes adjusting the tone and speed of the voice. 12. A voice dialogue method according to any one of 9 to 11, wherein the voice synthesized with the voice parameters is output to the customer. 13. 14. A voice dialogue method according to any one of 9 to 12, wherein information indicating the emotional state of a customer is acquired, and the voice parameters are optimized based on the emotional state. 15. A voice dialogue method according to 13, wherein the tone of voice is adjusted based on the emotional state. 16. A meeting minutes generation device having means for acquiring video data and audio data, means for analyzing emotions from the video data and audio data, means for generating text data from the audio data, means for integrating the results of the emotion analysis and the text data in chronological order, and means for generating meeting minutes that reflect the flow of emotions based on the integrated data.16. A recording medium on which a program is recorded that causes a computer to function as: means for acquiring video data and audio data; means for analyzing emotions from the video data and audio data; means for generating text data from the audio data; means for integrating the results of the emotion analysis and the text data in chronological order; and means for generating meeting minutes that reflect the flow of emotions based on the integrated data. 17. A real-time customer service support device having means for acquiring audio data and video data of a customer during a call in real time; means for analyzing emotions from the audio data and video data; means for generating a response policy or question based on the results of the emotion analysis; and means for presenting the response policy or question to the operator. 18. A recording medium on which a program is recorded that causes a computer to function as: means for acquiring audio data and video data of a customer during a call in real time; means for analyzing emotions from the audio data and video data; means for generating a response policy or question based on the results of the emotion analysis; and means for presenting the response policy or question to the operator. 19. A voice dialogue device comprising: means for acquiring customer attribute data and product characteristic data; means for optimizing voice parameters based on the customer attribute data and product characteristic data; means for generating proposal content based on the customer attribute data and product characteristic data; and means for synthesizing the proposal content into speech based on the optimized voice parameters. 20. A recording medium storing a program that causes a computer to function as: means for acquiring customer attribute data and product characteristic data; means for optimizing voice parameters based on the customer attribute data and product characteristic data; means for generating proposal content based on the customer attribute data and product characteristic data; and means for synthesizing the proposal content into speech based on the optimized voice parameters.
[0148] Some or all of the appendices 2 to 4, which are dependent on the minutes generation method described in appendix 1 above, may also be dependent on the minutes generation device of appendix 15 and the recording medium of appendix 16 in the same dependent relationship as between appendix 1 and appendices 2 to 4. Furthermore, some or all of the appendices 6 to 8, which are dependent on the real-time customer service support method of appendix 5, may also be dependent on the real-time customer service support device of appendix 17 and the recording medium of appendix 18 in the same dependent relationship as between appendix 5 and appendices 6 to 8. Furthermore, some or all of the appendices 10 to 14, which are dependent on the voice dialogue method of appendix 9, may also be dependent on the voice dialogue method device of appendix 19 and the recording medium of appendix 20 in the same dependent relationship as between appendix 9 and appendices 10 to 14. In addition, some or all of the configurations described as appendices can be realized in various hardware, software, various recording means for recording software, or systems, without departing from each of the embodiments described above.
[0149] 1A Processor 2A Memory 3A Input / Output Interface 4A Peripheral Circuits 5A Bus
Claims
1. A method for generating meeting minutes, characterized in that one or more computers acquire video data and audio data, analyze emotions from the video data and audio data, generate text data from the audio data, integrate the results of the emotion analysis and the text data in chronological order, and generate meeting minutes that reflect the flow of emotions based on the integrated data.
2. A method for generating meeting minutes according to claim 1, wherein the analysis of emotions is characterized by: recognizing facial expressions from the video data to estimate a first emotion; analyzing voice characteristics from the voice data to estimate a second emotion; and integrating the first emotion and the second emotion to determine the emotional state.
3. A method for generating meeting minutes according to claim 1 or 2, wherein the integration of the emotion analysis results and the character data in chronological order includes linking the emotion analysis results at the timing of each utterance of the utterance of each utterance to each of the utterance texts included in the chronological string data, thereby generating the chronological integrated data.
4. A method for generating meeting minutes according to any one of claims 1 to 3, wherein the video data and the audio data are acquired in real time.
5. A real-time customer service support method characterized by one or more computers acquiring voice and video data of a customer during a call in real time, analyzing emotions from the voice and video data, generating a response policy or question based on the results of the emotion analysis, and presenting the response policy or question to the operator.
6. A real-time customer service support method according to claim 5, wherein generating the response policy includes detecting that a change in the customer's emotional state satisfies predetermined attention conditions, and generating the response policy in response to the detection.
7. The real-time customer service support method according to claim 6, wherein the caution condition includes the customer's emotional state, expressed numerically, changing at a speed exceeding a predetermined threshold.
8. A real-time customer service support method according to any one of claims 5 to 7, wherein generating the question includes determining the necessity of the question based on the emotional state, and generating the question when it is determined that the question is necessary.
9. A voice dialogue method characterized in that one or more computers acquire customer attribute data and product feature data, optimize voice parameters based on the customer attribute data and product feature data, generate proposal content based on the customer attribute data and product feature data, and synthesize the proposal content into speech based on the optimized voice parameters.
10. The voice dialogue method according to claim 9, wherein the customer attribute data includes at least one of the customer's risk preference and hobbies and preferences, and the product feature data includes at least one of the product's risk level and selling points.
11. A voice dialogue method according to claim 9 or 10, wherein optimizing the voice parameters includes adjusting the tone and speed of the voice.
12. A voice dialogue method according to any one of claims 9 to 11, wherein the voice synthesized with the voice parameters is output to the customer.
13. A voice dialogue method according to any one of claims 9 to 12, wherein information indicating the emotional state of a customer is acquired, and the voice parameters are optimized based on the emotional state.
14. A voice dialogue method according to claim 13, wherein the tone of voice is adjusted based on the emotional state.
15. A meeting minutes generation device comprising: means for acquiring video data and audio data; means for analyzing emotions from the video data and audio data; means for generating text data from the audio data; means for integrating the results of the emotion analysis and the text data in chronological order; and means for generating meeting minutes that reflect the flow of emotions based on the integrated data.
16. A recording medium on which a program is recorded that causes a computer to function as: means for acquiring video data and audio data; means for analyzing emotions from the video data and audio data; means for generating text data from the audio data; means for integrating the results of the emotion analysis and the text data in chronological order; and means for generating meeting minutes that reflect the flow of emotions based on the integrated data.
17. A real-time customer service support device comprising: means for acquiring voice data and video data of a customer during a call in real time; means for analyzing emotions from the voice data and video data; means for generating a response policy or question based on the results of the emotion analysis; and means for presenting the response policy or question to the operator.
18. A recording medium on which a program is recorded that causes a computer to function as a means for acquiring voice data and video data of a customer during a call in real time, a means for analyzing emotions from the voice data and video data, a means for generating a response policy or question based on the results of the emotion analysis, and a means for presenting the response policy or question to an operator.
19. A voice dialogue device comprising: means for acquiring customer attribute data and product feature data; means for optimizing voice parameters based on the customer attribute data and product feature data; means for generating proposal content based on the customer attribute data and product feature data; and means for synthesizing the proposal content into speech based on the optimized voice parameters.
20. A recording medium that stores a program causing a computer to function as: means for acquiring customer attribute data and product characteristic data; means for optimizing voice parameters based on the customer attribute data and product characteristic data; means for generating proposal content based on the customer attribute data and product characteristic data; and means for synthesizing the proposal content into speech based on the optimized voice parameters.