System
The system addresses the challenge of real-time intention and emotion recognition in conversations by using cameras and microphones to analyze facial expressions and audio, providing speech plans and feedback for enhanced communication skills.
Patent Information
- Application Number
- JP2024119074
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
AI Technical Summary
Existing systems fail to accurately read and respond to the intentions and emotions of conversation partners in real-time, particularly in business meetings and negotiations, lacking effective feedback for improving speaking skills.
A system that collects facial expressions and gestures using a camera and audio data with a microphone, analyzes this data in real-time to infer intentions and emotions, and provides an appropriate speech plan, along with a self-practice mode for improving communication skills.
Enables smooth and effective interpersonal communication by accurately grasping the conversation partner's intentions and emotions, facilitating improved speaking skills through real-time feedback.
Smart Images

Figure 2026018013000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In interpersonal conversations, it is important to accurately read the other person's intentions and communicate smoothly, but this is an extremely difficult task. In particular, in business meetings, negotiations, counseling, and other situations, it is necessary to quickly and appropriately grasp the other person's subtle changes in facial expressions, gestures, and emotions, and respond accordingly. Furthermore, while practice is essential to improve one's own speaking skills, it is difficult to self-evaluate and identify areas for improvement. For this reason, a system is needed that can estimate the other person's intentions and emotions in real time and propose effective speech plans. [Means for solving the problem]
[0005] To solve this problem, the present invention provides the following means. First, the system includes means for collecting facial expressions and gestures of a conversation partner using a camera. Second, the system includes means for collecting audio data of the conversation partner using a microphone and means for analyzing this data in real time. The analysis means analyzes the collected video and audio data and infers the intentions and emotions of the conversation partner. The system then provides means for displaying an appropriate speech plan to the user based on the inferred intentions and emotions. The system also includes means for organizing the collected conversation content in a tree diagram or the like and displaying it to the user, means for providing an interface that allows the user to select a self-practice mode, and means for analyzing the user's speech data and providing feedback for improving speech. This allows the user to properly understand the intentions and emotions of the conversation partner, enabling smooth communication and improving their own speech skills.
[0006] A "camera" is a device for taking images, and in the present invention is a device used to collect facial expressions and gestures of a conversation partner.
[0007] A "microphone" is a device for recording sound, and in the present invention is a device used to collect sound data of a conversation partner.
[0008] "Video data" refers to data of images or videos captured by a camera, and is data used in the present invention to analyze facial expressions and gestures.
[0009] "Voice data" refers to data of speech recorded by a microphone, and in the present invention, this data is used to analyze the word choice and intonation of a conversation partner.
[0010] "Analysis means" refers to functions and software for processing collected video and audio data and inferring the intentions and emotions of the person being spoken to.
[0011] "Intention" refers to what the other person is thinking and what they want to communicate.
[0012] "Emotions" indicate the mental state and feelings of the person you are talking to, and can be inferred from facial expressions, tone of voice, etc.
[0013] An "utterance plan" is a suggestion or guideline that specifically shows the user what to say next and how to proceed with the conversation.
[0014] A "tree diagram" is a diagram that visually organizes conversation content in a hierarchical structure, and is used to indicate the next step a user should take in the conversation.
[0015] The "self-practice mode" is a mode in which the user can practice speaking by themselves, and in the present invention, can be selected and started through the interface.
[0016] An "interface" is a function or device that enables interaction between the user and the system, and in the present invention, it is used to select a self-practice mode and provide various feedback.
[0017] "Feedback" refers to an evaluation of the user's speech and suggestions for improvement, and is provided to improve speech skills. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram illustrating a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] The present invention is a system that estimates the intentions and emotions of a conversation partner from their facial expressions, gestures, and voice, and presents an appropriate speech plan to the user, in order to facilitate interpersonal communication. Below, we will create a program for this system and explain its processing in detail.
[0040] System Configuration
[0041] The system consists of a terminal device (conversation prediction glasses) and a server device that performs data analysis. The terminal device is equipped with a camera and microphone, and the server device is equipped with an AI analysis model.
[0042] Program processing flow
[0043] Data collection
[0044] The device has a built-in camera to capture the facial expressions and gestures of the person you are talking to, and a microphone to collect audio data. When a conversation begins, the camera collects the video data of the person you are talking to, and the microphone collects the audio data in real time, and then sends the data to a server.
[0045] Data analysis
[0046] The server uses facial recognition technology to extract facial features of the conversation partner from the received video data and also analyzes gestures. At the same time, it converts the speech content into text using speech recognition technology from the audio data and estimates changes in emotion through intonation analysis. This data is input into an AI model, which then infers the other person's intentions and emotions.
[0047] Generate an utterance plan
[0048] The server generates a speech plan based on the analysis results, according to the intentions and emotions expressed by the conversation partner. For example, if a customer shows interest in a particular product, the server generates a speech plan to explain the details of that product and recommends it to the user. The generated speech plan is sent to the terminal.
[0049] User Feedback
[0050] The device displays the speech plan sent from the server on its display and provides the user with appropriate speech content. By continuing the conversation according to the displayed content, the user can communicate appropriately with the other person's intentions and emotions.
[0051] Organizing the conversation
[0052] Furthermore, the server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the device. The device displays this tree diagram, providing visual guidance on the next step in the conversation for the user.
[0053] Self-Practice Mode
[0054] The device has a self-practice mode, allowing users to practice speaking by themselves. In this mode, the server provides speaking samples, analyzes the user's speech data, and provides evaluation and feedback, allowing users to improve their speaking skills.
[0055] Specific examples
[0056] For example, a salesperson puts on these conversation-reading glasses when negotiating with a customer. After the negotiation begins, the glasses' camera captures the customer's facial expressions and gestures, and the microphone collects what the customer says, and these data are sent to a server. The server analyzes the data and infers that the customer is interested in price. The server then generates a speech plan, such as "This product has excellent cost performance and will offer great benefits in the long run," and sends it to the device. The device then displays this on its screen, and the salesperson continues to explain the plan to the customer. In this way, the conversation progresses smoothly, enabling responses to be made that meet the customer's needs.
[0057] This system makes it possible to accurately grasp the intentions and emotions of the person you are talking to, enabling smooth communication, which in turn makes business negotiations, counseling, and even everyday conversations more effective.
[0058] The processing flow will be explained below.
[0059] Step 1: Your device is ready for the camera and microphone
[0060] Before the conversation begins, the device turns on the camera and microphone to prepare for data collection.
[0061] Show the user a ready notification.
[0062] Step 2: The device collects video and audio data from the person you are talking to.
[0063] The camera continuously captures the facial expressions and gestures of the person you are talking to.
[0064] The microphone records the voice of the person you are talking to in real time.
[0065] The collected data is sent to the server in real time.
[0066] Step 3: The server analyzes the video data
[0067] The server analyzes the received video data and extracts facial expression features of the conversation partner.
[0068] Use facial recognition technology to estimate emotions from facial expressions.
[0069] Gesture analysis is performed to infer intentions from hand and body movements.
[0070] Step 4: The server analyzes the audio data
[0071] The server analyzes the received voice data and converts the spoken content into text using voice recognition technology.
[0072] The text of the speech is analyzed, and emotions are inferred from word choice and intonation.
[0073] Step 5: The server estimates the other person's intentions and emotions
[0074] The features extracted from the video and audio data are input into the AI model.
[0075] The AI model estimates the overall intention and emotions of the conversation partner.
[0076] Step 6: The server generates an appropriate utterance plan
[0077] Based on the estimation results, the AI generates the optimal speech plan for the user.
[0078] A speech plan includes specific phrases and next steps in the conversation.
[0079] Step 7: The server sends the speech plan to the device.
[0080] The server transmits the generated speech plan to the terminal.
[0081] Optimize and transmit data to ensure stable communication and prevent delays.
[0082] Step 8: The device displays the speech plan to the user.
[0083] The terminal displays the speech plan received from the server on the display.
[0084] The user continues the conversation by following what is displayed.
[0085] Step 9: The server organizes the conversation
[0086] As the conversation progresses, the server organizes the user's comments and the analysis results.
[0087] Structure the conversation content as a tree diagram or flowchart.
[0088] Step 10: The server sends the organized conversation to the device.
[0089] The organized conversation content and the next steps to take are sent to the terminal.
[0090] Help users make good decisions.
[0091] Step 11: The terminal displays the tree diagram to the user
[0092] The terminal displays the tree diagram or flowchart received from the server.
[0093] Check what the user should say next and the flow.
[0094] Step 12: User selects and starts self-practice mode
[0095] The user selects the self-practice mode through the terminal interface.
[0096] An interface provides the user with practice mode settings.
[0097] Step 13: The server provides and analyzes speech samples
[0098] The server provides speech samples for use in self-practice mode.
[0099] Data spoken by users is collected in real time and analyzed.
[0100] Step 14: The server generates feedback and sends it to the device
[0101] The server evaluates the user's speech data and generates suggestions for improvement and appropriate feedback.
[0102] The generated feedback is sent to the device.
[0103] Step 15: The device displays feedback to the user
[0104] The terminal displays the feedback received from the server to the user.
[0105] The user improves their speaking skills based on the feedback.
[0106] Through these processing steps, users can properly understand the intentions and emotions of their conversation partners and communicate smoothly. In addition, by utilizing the self-practice mode, users can effectively improve their speaking skills.
[0107] Example 1
[0108] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0109] Interpersonal communication requires rapid and accurate understanding of the intentions and emotions of the person you are speaking with, but the current technology for doing so is insufficient. It is particularly difficult to respond appropriately to changes in the other person's intentions and emotions in situations such as business negotiations and counseling. There is also a lack of effective feedback systems for improving speaking skills through self-practice. A system that solves these issues is needed.
[0110] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0111] In this invention, the server includes means for collecting facial expressions and gestures of a conversation partner using a camera, means for collecting voice data of the conversation partner using a microphone, means for transmitting the collected video and voice data to the server in real time, means for extracting facial features from the received video data using face recognition technology and performing gesture analysis, means for converting the speech content into text from the voice data using voice recognition technology and estimating emotional changes through intonation analysis, and means for generating and displaying an appropriate speech plan for the user based on the estimated intentions and emotions. This makes it possible to accurately grasp the intentions and emotions of the conversation partner and achieve smooth communication.
[0112] A "camera" is a device for collecting video data, and in this invention is used to capture the facial expressions and gestures of a conversation partner.
[0113] A "microphone" is a device for collecting voice data, and in the present invention is used to clearly capture the voice of a conversation partner.
[0114] A "server" is a computer system for analyzing and processing data, and in the present invention, it has the role of analyzing video data and audio data and providing speech plans to users.
[0115] "Facial expression features" are a set of feature values that indicate the emotions and intentions of a conversation partner, extracted using face recognition technology.
[0116] "Gesture analysis" is the process of analyzing a conversation partner's hand movements and gestures to infer non-verbal intentions and emotions.
[0117] "Speech recognition technology" is a technology that analyzes voice data and converts it into text data, and is used in the present invention to analyze the content of speech from a conversation partner.
[0118] "Intonation analysis" is a technology that analyzes the intonation and rhythm of voice data to estimate changes in emotion.
[0119] An "utterance plan" is a proposal of utterance content to support the progress of a conversation, generated based on analyzed data.
[0120] A "tree diagram" is a diagram that organizes the development of a conversation in a hierarchical structure to visually show the progress of the conversation.
[0121] The "self-practice mode" is a function that allows the user to practice speaking on their own, and includes providing speaking samples and feedback.
[0122] The present invention is a system that estimates the intentions and emotions of a conversation partner from their facial expressions, gestures, and voice, and presents an appropriate speech plan to the user, in order to facilitate interpersonal communication. Below, we will create a program for this system and explain its processing in detail.
[0123] System Configuration
[0124] The main components of this system are a terminal device (conversation prediction glasses) and a server device that performs data analysis. The terminal device is equipped with a camera and microphone, and the server device is equipped with an AI analysis model.
[0125] Program processing flow and specific functions
[0126] Data collection
[0127] The device has a built-in camera for capturing facial expressions and gestures of the conversation partner, and a microphone for capturing audio data. When a conversation begins, the camera captures the conversation partner's video data and the microphone captures audio data in real time, and these data are then sent to a server. The hardware used is a built-in camera (e.g., an HD camera) and a highly sensitive microphone.
[0128] Data analysis
[0129] The server uses facial recognition technology to extract facial features of the conversation partner from the received video data and also performs gesture analysis. At the same time, it uses speech recognition technology to convert the speech into text from the audio data and estimates emotional changes through intonation analysis. Software libraries such as TensorFlow and OpenCV are used for these analyses. For example, Google's Speech-to-Text API is sometimes used to analyze audio data.
[0130] Generate an utterance plan
[0131] The server generates a speech plan based on the analysis results, taking into account the intentions and emotions of the conversation partner. This is done using a pre-trained generative AI model. For example, if a customer shows interest in a particular product, a speech plan to explain the details of that product is generated. The generated speech plan is sent to the device.
[0132] User Feedback
[0133] The device displays the speech plan sent from the server on its display and provides the user with appropriate speech content. By continuing the conversation according to the displayed content, the user can communicate appropriately with the other person's intentions and emotions.
[0134] Organizing the conversation
[0135] Furthermore, the server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the device. The device displays this tree diagram, providing visual guidance on the next step in the conversation for the user.
[0136] Self-Practice Mode
[0137] The device has a self-practice mode, allowing users to practice speaking by themselves. In this mode, the server provides speaking samples, analyzes the user's speech data, and provides evaluation and feedback, allowing users to improve their speaking skills.
[0138] Specific examples
[0139] For example, a salesperson puts on these conversation-reading glasses when negotiating with a customer. After the negotiation begins, the glasses' camera captures the customer's facial expressions and gestures, and the microphone collects what the customer says, and these data are sent to a server. The server analyzes the data and infers that the customer is interested in price. The server then generates a speech plan, such as "This product has excellent cost performance and will offer great benefits in the long run," and sends it to the device. The device then displays this on its screen, and the salesperson continues to explain the plan to the customer. In this way, the conversation progresses smoothly, enabling responses to be made that meet the customer's needs.
[0140] Example prompts for generative AI models
[0141] Input data: Customer facial expression information, voice data
[0142] AI model: Estimating conversation partner's intentions and emotions
[0143] Output: It is assumed that the customer is interested in the price. Generate the utterance plan "This product offers excellent value for money and offers great benefits in the long run."
[0144] As described above, the present invention is a system for understanding the intentions and emotions of a conversation partner and supporting appropriate communication, and is particularly effective in situations such as business negotiations and counseling.
[0145] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0146] Step 1:
[0147] Data collection
[0148] The device uses a camera and microphone to capture facial expressions, gestures, and voice data from the person you're talking to. The camera captures high-resolution video, while the microphone suppresses background noise to capture clear audio.
[0149] input:
[0150] Video and audio data of the person you are talking to
[0151] output:
[0152] Collected video and audio data
[0153] Specific behavior:
[0154] When the user begins interacting with a customer, the device's camera captures the customer's facial expressions and gestures, and the microphone begins recording what the customer says.
[0155] Step 2:
[0156] Data transmission
[0157] The device transmits the collected video and audio data to the server in real time over a fast and secure network.
[0158] input:
[0159] Collected video and audio data
[0160] output:
[0161] Data sent to the server
[0162] Specific behavior:
[0163] The device encrypts the collected data and sends it to a server using Wi-Fi or 5G networks.
[0164] Step 3:
[0165] Data analysis
[0166] The server extracts facial features from the received video data using facial recognition technology and also performs gesture analysis. It also converts the speech into text using speech recognition technology and estimates emotional changes through intonation analysis. TensorFlow and OpenCV are used as software libraries.
[0167] input:
[0168] Video and audio data sent to the server
[0169] output:
[0170] Extracted facial features, textualized speech, and estimated emotional changes
[0171] Specific behavior:
[0172] The server uses TensorFlow to analyze the video data and extract facial expression and gesture features. At the same time, Google's Speech-to-Text API converts the audio data into text, and the emotion analysis API analyzes emotions from intonation.
[0173] Step 4:
[0174] Generate an utterance plan
[0175] The server generates a speech plan based on the analysis results, taking into account the intentions and emotions of the conversation partner. Using a generative AI model, it constructs an appropriate speech plan that corresponds to the estimated intentions and emotions.
[0176] input:
[0177] Extracted facial features, textualized speech, and estimated emotional changes
[0178] output:
[0179] Generated utterance plan
[0180] Specific behavior:
[0181] If the server's AI model determines that the customer is interested in price, it generates a speech plan such as, "This product has excellent cost performance."
[0182] Step 5:
[0183] Sending a speech plan
[0184] The server then sends the generated speech plan to the device, again in real time, allowing the user to receive immediate feedback.
[0185] input:
[0186] Generated utterance plan
[0187] output:
[0188] Speech plan sent to the device
[0189] Specific behavior:
[0190] The speech plan generated by the server is immediately sent to the device, and a notification is displayed on the device's display.
[0191] Step 6:
[0192] User Feedback
[0193] The device displays the received speech plan on its display and provides the user with appropriate speech content, which the user can refer to as they proceed with the conversation.
[0194] input:
[0195] Speech plan sent to the device
[0196] output:
[0197] The speech plan presented to the user
[0198] Specific behavior:
[0199] The terminal display will show "This product has excellent cost performance," and the user will convey this statement to the customer.
[0200] Step 7:
[0201] Organizing the conversation
[0202] The server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the terminal.
[0203] input:
[0204] Analyzed conversation content
[0205] output:
[0206] Speech content organized as a tree diagram
[0207] Specific behavior:
[0208] The server analyzes the flow of conversation, and if there is a high level of "interest in price," it generates a node for "additional information related to price" and displays it in the form of a tree diagram on the terminal display.
[0209] Step 8:
[0210] Self-Practice Mode
[0211] The device has a self-practice mode, allowing users to practice speaking by themselves. In this mode, the server provides speaking samples and analyzes the user's speech data to provide evaluation and feedback.
[0212] input:
[0213] Speech samples entered by the user
[0214] output:
[0215] Ratings and Feedback
[0216] Specific behavior:
[0217] When a user selects practice mode, the server sends sample utterances, such as "greeting practice," to the device. As the user speaks, the audio is analyzed and pronunciation and intonation are evaluated in real time.
[0218] (Application example 1)
[0219] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0220] Existing conversation support systems have difficulty understanding the intentions and emotions of the target, and are therefore unable to present appropriate speech plans or response methods for achieving effective communication. Furthermore, in security services, where it is necessary to quickly understand the intentions and emotions of visitors and passersby and respond appropriately, the lack of such technology poses a serious problem.
[0221] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0222] In this invention, the server includes means for collecting facial expressions and gestures of a target using a camera, means for collecting audio data of the target using a microphone, means for analyzing the collected video and audio data to estimate the intention and emotion of the target, means for displaying an appropriate speech plan to the user based on the estimated intention and emotion, and means for suggesting a response method to the user based on the estimated intention and emotion, thereby enabling security guards to quickly grasp the intention and emotion of visitors and passersby and to determine an appropriate response method.
[0223] A "camera" is a photographic device for capturing facial expressions and gestures of a subject.
[0224] A "microphone" is a recording device for collecting audio data of a subject.
[0225] "Video data" refers to image and video information collected using a camera.
[0226] "Audio data" is sound information collected using a microphone.
[0227] "Facial expression" refers to emotions and intentions that can be read from the movement of facial muscles.
[0228] A "gesture" is an expression of intent that can be interpreted through hand or body movements.
[0229] "Intention" refers to the purpose or direction of an action that a subject has.
[0230] "Emotion" refers to a subject's state of mind or mood.
[0231] A "speech plan" is a plan for providing appropriate conversation content.
[0232] A "display" is a display device for presenting visual information to a user.
[0233] A "server" is a computer system that analyzes collected data and provides results.
[0234] A "tree diagram" is a diagram that visually organizes the content and development of a conversation.
[0235] The "self-practice mode" is a mode in which the user can practice speaking by himself.
[0236] "Feedback" refers to evaluation of the user's speech and advice for improvement.
[0237] This invention is an interpersonal communication support system specialized for security applications, and employs an approach to infer the intentions and emotions of a conversation partner from their facial expressions, gestures, and voice, and provide appropriate response methods.
[0238] System Configuration
[0239] The system includes smart glasses (terminal devices) and a server device that performs data analysis. The smart glasses are equipped with a camera and microphone, and the server device is equipped with an AI analysis model.
[0240] Data collection
[0241] The terminal device is equipped with a camera to capture the facial expressions and gestures of the visitor and a microphone to collect audio data. When a security guard comes into contact with a visitor, the camera collects the subject's video data and the microphone collects audio data in real time, and these data are then sent to a server.
[0242] Data analysis
[0243] The server uses facial recognition technology to extract facial features from the received video data and also analyzes gestures. At the same time, it converts the speech into text using speech recognition technology and estimates emotional changes through intonation analysis. This data is input into an AI model, which then infers the subject's intentions and emotions.
[0244] Generate speech plans and responses
[0245] Based on the analysis results, the server generates a speech plan and appropriate response method according to the visitor's intentions and emotions. For example, if the visitor seems nervous, the server generates advice recommending polite responses and sends it to the terminal device. The generated speech plan and response method are then displayed on the smart glasses' display.
[0246] User Feedback
[0247] The terminal device displays the speech plan and response method sent from the server on its display and provides the security guard with appropriate response content. By responding to the visitor according to the displayed content, the security guard can appropriately deal with the visitor's intentions and emotions.
[0248] Organizing the conversation
[0249] Furthermore, the server analyzes the progress of the conversation in real time, organizes the content and development of the conversation into a tree diagram, and sends it to the terminal device, which displays this tree diagram and provides visual guidance to the security guard on the next steps in the conversation.
[0250] Self-Practice Mode
[0251] The terminal device is equipped with a self-practice mode, allowing security guards to practice speaking by themselves. In this mode, the server provides speaking samples, analyzes the security guards' speech data, and provides evaluation and feedback, allowing the security guards to improve their speaking skills.
[0252] Examples and prompts
[0253] Specific examples
[0254] A security guard wears smart glasses at the entrance to a facility and handles visitors. If the visitor's facial expression appears stiff and tense, the camera and microphone collect information and the information is analyzed by a server. As a result of the analysis, the smart glasses' display displays advice such as, "It is assumed that the visitor is nervous. Please handle the visitor politely and confirm the purpose of their visit."
[0255] Prompt Sentence Examples
[0256] "Advice to display when a visitor's facial expression appears tense:
[0257] We suspect the visitor may be nervous. Please be polite and confirm the purpose of their visit."
[0258] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0259] Step 1:
[0260] Data collection
[0261] The device (smart glasses) uses a camera to collect video data of the subject's facial expressions and gestures, and simultaneously collects audio data using a microphone. The input is video data from the camera and audio data from the microphone, and the output is the collected raw video and audio data.
[0262] Step 2:
[0263] Data transmission
[0264] The terminal transmits the collected video and audio data to the server in real time. The input is the collected video and audio data, and the output is the data transmitted to the server.
[0265] Step 3:
[0266] Facial recognition and facial expression feature extraction
[0267] The server uses facial recognition technology to detect the target face from the received video data and extract facial expression features. The input is the video data, and the output is the extracted facial expression features. The software used is a facial recognition library such as OpenCV.
[0268] Step 4:
[0269] Gesture Analysis
[0270] The server analyzes hand and body movements from video data to detect gestures. The input is video data, and the output is detected gesture data. The software used is an AI analysis model.
[0271] Step 5:
[0272] Converting audio data to text
[0273] The server converts the voice data into text data using speech recognition technology. The input is voice data and the output is text data. The software used is a speech recognition library (e.g., Google Cloud Speech-to-Text).
[0274] Step 6:
[0275] Intonation analysis
[0276] The server analyzes the intonation of the voice data and estimates emotional changes. The input is the voice data, and the output is estimated emotional data. The software used is an AI analysis model.
[0277] Step 7:
[0278] Intention and emotion estimation
[0279] The server inputs the extracted facial features, gesture data, text data, and emotion data into an AI model to comprehensively estimate the subject's intention and emotion. The input is each feature data, and the output is estimated intention and emotion information.
[0280] Step 8:
[0281] Generate speech plans and responses
[0282] The server generates a speech plan and a response method for the user based on the estimated intention and emotion. For example, if the visitor is nervous, it generates advice including an appropriate response method. The input is the intention and emotion information, and the output is the speech plan and a response method.
[0283] Step 9:
[0284] User Feedback
[0285] The terminal displays the speech plan and response method sent from the server on a display and provides them to the security guard. The input is the speech plan and response method, and the output is the content displayed on the display. The security guard responds to the visitor according to the displayed content.
[0286] Step 10:
[0287] Organizing the conversation
[0288] The server analyzes ongoing conversations in real time and organizes them into a tree diagram. The input is real-time conversation data, and the output is a tree diagram of the organized conversations. The terminal displays this tree diagram to help security guards understand the next steps.
[0289] Step 11:
[0290] Self-Practice Mode
[0291] The terminal provides an interface that allows the user to select self-practice mode. The server provides speech samples, analyzes the user's speech data, and provides evaluation and feedback. The input is the user's speech data, and the output is evaluation and feedback, allowing the security guard to improve their speaking skills.
[0292] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0293] This invention is a system that combines conversation prediction glasses with an emotion engine to more accurately recognize the emotions and intentions of both the user and the conversation partner and present appropriate speech plans. Below, we will create a program for this system and explain its processing in detail.
[0294] System Configuration
[0295] The system includes a terminal device (conversation prediction glasses), a server device that performs data analysis, and an emotion engine with emotion recognition capabilities. The terminal device is equipped with a camera and microphone, and the server device is equipped with an AI analysis model and emotion engine.
[0296] Program processing flow
[0297] Data collection
[0298] The device uses a camera and microphone to collect facial expressions, gestures, and voice data of the user and the conversation partner. When a conversation begins, the camera collects video data of the conversation partner and the user, and the microphone collects voice data in real time, and then transmits this data to a server.
[0299] Data analysis
[0300] The server analyzes the received video and audio data and extracts facial and audio features of the conversation partner and the user. First, facial recognition technology is used to infer emotions from facial expressions, followed by gesture analysis. Furthermore, speech recognition technology is used to convert the speech into text from the audio data, and emotional changes are inferred through intonation analysis. This data is input into an AI model and emotion engine, which infers the intentions and emotions of the conversation partner and the user.
[0301] Generate an utterance plan
[0302] The server generates a speech plan based on the analysis results, taking into account the emotional state and intentions of the conversation partner and the user. For example, the server may generate a speech plan that provides detailed explanations based on a topic that the conversation partner is interested in, or suggest relaxing topics to reduce the stress the user is feeling. The generated speech plan is then sent to the device.
[0303] User Feedback
[0304] The device displays the speech plan sent from the server on its display and provides the user with appropriate speech content. By continuing the conversation according to the displayed content, the user can communicate appropriately with the other person's intentions and emotions.
[0305] Organizing the conversation
[0306] Furthermore, the server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the device. The device displays this tree diagram, providing visual guidance on the next step in the conversation for the user.
[0307] Self-Practice Mode
[0308] The device offers a self-practice mode, allowing users to practice speaking by themselves. In this mode, the server provides speech samples, collects and analyzes the user's speech data in real time, and generates and sends feedback to the device based on the analysis results.
[0309] Specific examples
[0310] For example, a sales representative puts on these conversation-predicting glasses when negotiating with a customer. After the negotiation begins, the glasses' camera captures the facial expressions and gestures of both the customer and the sales representative, while the microphone collects speech, and these data are sent to a server. The server analyzes the data and infers that the customer is beginning to show interest in the product but is concerned about the price. The emotion engine then confirms that the sales representative is explaining with confidence. The server generates a speech plan, such as "This product has excellent cost performance and will offer great benefits in the long run," and sends it to the device. The device displays this on its screen, and the sales representative explains based on it. In this way, the conversation progresses smoothly, enabling responses that meet the customer's needs.
[0311] This system allows both the user and the conversation partner to accurately understand each other's intentions and emotions, enabling smooth and effective communication, which will enable more effective business negotiations, counseling, and everyday conversations.
[0312] The processing flow will be explained below.
[0313] Step 1: Your device is ready for the camera and microphone
[0314] Before the conversation begins, the device turns on the camera and microphone to prepare for data collection.
[0315] Show the user a ready notification.
[0316] Step 2: The device collects video and audio data of the conversation partner and the user.
[0317] The camera continuously captures the facial expressions and gestures of the conversation partner and the user.
[0318] The microphone records the voices of the conversation partner and the user in real time.
[0319] The collected data is sent to the server in real time.
[0320] Step 3: The server analyzes the video data of the person you are talking to.
[0321] The server analyzes the received video data and extracts facial features of the conversation partner.
[0322] Use facial recognition technology to estimate emotions from facial expressions.
[0323] Gesture analysis is performed to infer intentions from hand and body movements.
[0324] Step 4: The server analyzes the voice data of the other party.
[0325] The server analyzes the received voice data and converts the spoken content into text using voice recognition technology.
[0326] The text of the speech is analyzed, and emotions are inferred from word choice and intonation.
[0327] Step 5: The server analyzes the user's video data.
[0328] The server analyzes the received video data and extracts the user's facial expression features.
[0329] Use facial recognition technology to estimate emotions from facial expressions.
[0330] Gesture analysis is performed to estimate the current state and intentions from hand and body movements.
[0331] Step 6: The server analyzes the user's voice data
[0332] The server analyzes the received voice data and converts the spoken content into text using voice recognition technology.
[0333] The system analyzes the text of the speech and estimates the user's emotions from word choice and intonation.
[0334] Step 7: The server integrates and estimates the intentions and emotions of the conversation partner and the user.
[0335] The features extracted from the video and audio data are input into the AI model and emotion engine.
[0336] AI models and emotion engines estimate the overall intent and emotions of the conversation partner and the user.
[0337] Step 8: The server generates an appropriate utterance plan
[0338] Based on the estimation results, the AI generates the optimal speech plan for the user.
[0339] A speech plan includes specific phrases and next steps in the conversation.
[0340] Step 9: The server sends the speech plan to the device.
[0341] The server transmits the generated speech plan to the terminal.
[0342] Optimize and transmit data to ensure stable communication and prevent delays.
[0343] Step 10: The device displays the speech plan to the user.
[0344] The terminal displays the speech plan received from the server on the display.
[0345] The user continues the conversation by following what is displayed.
[0346] Step 11: The server organizes the conversation
[0347] As the conversation progresses, the server organizes the user's comments and the analysis results.
[0348] Structure the conversation content as a tree diagram or flowchart.
[0349] Step 12: The server sends the organized conversation to the device.
[0350] The organized conversation content and the next steps to take are sent to the terminal.
[0351] Help users make good decisions.
[0352] Step 13: The terminal displays the tree diagram to the user
[0353] The terminal displays the tree diagram or flowchart received from the server.
[0354] Check what the user should say next and the flow.
[0355] Step 14: User selects and starts self-practice mode
[0356] The user selects the self-practice mode through the terminal interface.
[0357] An interface provides the user with practice mode settings.
[0358] Step 15: The server provides and analyzes speech samples.
[0359] The server provides speech samples for use in self-practice mode.
[0360] Data spoken by users is collected in real time and analyzed.
[0361] Step 16: The server generates feedback and sends it to the device
[0362] The server evaluates the user's speech data and generates suggestions for improvement and appropriate feedback.
[0363] The generated feedback is sent to the device.
[0364] Step 17: The device displays feedback to the user
[0365] The terminal displays the feedback received from the server to the user.
[0366] The user improves their speaking skills based on the feedback.
[0367] Through these processing steps, users can accurately understand their own and their conversation partner's emotions and intentions, enabling smoother and more effective communication. Furthermore, by utilizing the self-practice mode, users can effectively improve their speaking skills.
[0368] Example 2
[0369] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0370] In conversation, users often have difficulty accurately reading the intentions and emotions of the other person, making smooth communication difficult. They also often lack appropriate responses to the anxiety and stress they feel. Furthermore, even when practicing by themselves, the lack of concrete feedback makes it difficult to improve their speech.
[0371] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0372] In this invention, the server includes means for collecting facial expressions and gestures of the conversation partner and the user using a camera, means for collecting voice data of the conversation partner and the user using a microphone, means for analyzing the collected video and voice data to estimate the intentions and emotions of the conversation partner and the user, means for displaying an appropriate speech plan to the user based on the estimated intentions and emotions, means for organizing the progress of the conversation in a tree diagram based on the analysis results and displaying it to the user, and means for providing an interface that allows the user to select a self-practice mode and receive feedback for improving their speech. This allows the user to accurately understand the intentions and emotions of the conversation partner and take appropriate measures. Furthermore, the self-practice mode makes it easier to improve speech, thereby improving the overall quality of communication.
[0373] A "camera" is an optical device for collecting video data.
[0374] A "conversational partner" refers to the other person with whom the user is interacting.
[0375] "User" refers to a person who uses this system to have a conversation.
[0376] "Facial expressions" refer to facial movements and expressions, and are important non-verbal elements for indicating emotions and intentions.
[0377] "Gestures" refer to hand and body movements and are non-verbal expressions used to convey intentions and emotions.
[0378] A "microphone" is an acoustic device for collecting audio data.
[0379] "Audio data" refers to collected sound information, including speech content and intonation.
[0380] "Collect" refers to obtaining and storing data.
[0381] "Analysis" is the act of extracting and analyzing information based on collected data.
[0382] "Intention" refers to the purpose or thoughts of the conversation partner or user.
[0383] "Emotion" refers to the emotional state felt by a conversation partner or user.
[0384] A "speech plan" is a specific speech content suggested by the system based on a specific intention or emotion.
[0385] "Display" is the act of enabling a user to visually confirm information.
[0386] A "tree diagram" is a diagram that visually shows the progress of a conversation and the relationship between topics.
[0387] "Interface" refers to the operating screen or means by which a user interacts with a system.
[0388] "Feedback" refers to information that provides suggestions for improvement or evaluation of a user's actions or results.
[0389] The present invention is a system that includes a terminal device (conversation prediction glasses) worn by a user, a server device that performs data analysis, and an emotion engine with emotion recognition functionality. Specific embodiments for implementing this system are described below.
[0390] terminal device
[0391] The terminal device is equipped with a camera and a microphone. These hardware components are used to collect video and audio data of the user and the conversation partner in real time. Specifically, the camera captures facial expressions and gestures, and the microphone collects audio data of the conversation. The terminal device has the function of transmitting the collected data to a server device via a network.
[0392] Server device
[0393] The server device is equipped with an AI model and emotion engine for analyzing the received data. From the video data, facial recognition technology is used to extract facial expression features and perform gesture analysis. For the audio data, speech is converted into text using speech recognition technology, and emotional changes are estimated through intonation analysis. This data is input into the AI model and emotion engine, which ultimately estimates the intentions and emotions of the conversation partner and the user.
[0394] Providing Feedback
[0395] The server generates a speech plan based on the analysis results and transmits it to the terminal device. The terminal device displays this speech plan on a display and provides it to the user. The user can continue the conversation based on the displayed speech plan, and can respond appropriately to the intentions and feelings of the conversation partner.
[0396] Organizing the conversation
[0397] The server analyzes the progress of the conversation in real time and organizes the content and development of the conversation into a tree diagram. This tree diagram is sent to the terminal device, helping the user visually grasp the next flow of the conversation.
[0398] Self-Practice Mode
[0399] The terminal device is equipped with an interface that allows the user to select a self-practice mode, in which the server collects the user's speech, analyzes it in real time, and generates feedback, which helps the user improve their speech.
[0400] Specific examples
[0401] For example, a sales representative puts on conversation-predicting glasses when negotiating with a customer. When the negotiation begins, the device's camera captures the facial expressions and gestures of both the customer and the sales representative, and the microphone collects the audio of the conversation. This data is sent to a server, which uses analysis technology to infer the customer's emotions and intentions. Based on the analysis results, a speech plan such as "This product has excellent cost performance and will provide great benefits in the long run" is generated and displayed on the device. The user can explain the situation based on the displayed speech plan, allowing the negotiation to proceed smoothly.
[0402] Prompt Sentence Examples
[0403] "Using conversation-reading glasses, please generate a specific scenario in which a salesperson is negotiating with a customer. Include a process for reading the customer's emotions and interests and proposing an appropriate speech plan."
[0404] This system accurately recognizes the intentions and emotions of both the user and the conversation partner, enabling accurate and smooth communication, promoting effective dialogue in business negotiations, counseling, and everyday conversations.
[0405] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0406] Step 1:
[0407] The terminal collects video and audio of the conversation partner and the user in real time.
[0408] Specific operation: The device's camera captures the face, facial expressions, and gestures of the user and the person in conversation, and the microphone collects audio. The input is the video and audio data acquired through the camera and microphone, and the output is the collected data packets.
[0409] Step 2:
[0410] The terminal transmits the collected video and audio data to the server.
[0411] Specific operation: Video data is packetized as image frames, and audio data is packetized as audio clips, and then sent to the server over the network. The input is the collected data packets, and the output is the server to which the data was sent.
[0412] Step 3:
[0413] The server analyzes the received video data and extracts facial expression features.
[0414] Specific operation: Using facial recognition technology in the server, the system identifies the facial parts of the conversation partner and the user and calculates facial features such as smile, sadness, surprise, etc. The input is the received video data, and the output is the extracted facial features.
[0415] Step 4:
[0416] The server performs gesture analysis and infers emotions and intentions from the collected gesture data.
[0417] Specific operation: Using the server's gesture recognition algorithm, hand movements and body poses are analyzed to estimate emotional states such as interest or tension. The input is the received video data and existing gesture data, and the output is the estimated emotion or intention.
[0418] Step 5:
[0419] The server analyzes the voice data and converts the spoken content into text.
[0420] How it works: The speech recognition system analyzes recorded speech and converts it into text data. It then detects emotional changes through intonation analysis. The input is the received speech data, and the output is the text of the speech and an indicator of emotional changes.
[0421] Step 6:
[0422] The server uses AI models and emotion engines to infer the intentions and emotions of the user and their conversation partner.
[0423] How it works: The extracted features are input into an AI model based on previous research and training data to estimate emotions and intentions. An emotion engine is used to further improve accuracy. The inputs are facial features, gesture data, text data of spoken content, and indicators of emotional changes, and the output is estimated intentions and emotions.
[0424] Step 7:
[0425] Based on the analysis results, the server generates a speech plan that is appropriate for the emotions and intentions of the user and their conversation partner.
[0426] How it works: Using a generative AI model, it automatically determines the appropriate utterance content for the situation (e.g., question, explanation, relaxed topic, etc.) and constructs it as an utterance plan. The input is the estimated intent and emotion, and the output is the generated utterance plan.
[0427] Step 8:
[0428] The server transmits the generated speech plan to the terminal.
[0429] Specific operation: The text information generated as a speech plan is packetized and sent to the terminal via the network. The input is the generated speech plan, and the output is the terminal to which it was sent.
[0430] Step 9:
[0431] The terminal displays the received speech plan on the display.
[0432] Specific operation: The display module in the terminal displays the text of the speech plan on the screen, providing a visual for the user. The input is the received speech plan, and the output is the speech plan displayed on the display.
[0433] Step 10:
[0434] The user continues the conversation according to the content displayed on the display.
[0435] Specific operation: The user refers to the displayed speech plan and speaks at the appropriate time to progress the conversation. The input is the speech plan displayed on the display, and the output is the ongoing conversation.
[0436] Step 11:
[0437] The server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the terminal.
[0438] Specific operation: The collected data is analyzed one after another, and data for visualizing the conversation topics and progress as a tree diagram is generated and sent to the device. The input is the ongoing conversation data and analysis results, and the output is data organized as a tree diagram.
[0439] Step 12:
[0440] The terminal displays a tree diagram showing the progress of the conversation on its display.
[0441] Specific operation: The display module in the terminal displays the tree diagram on the screen, providing the user with a visual of the next conversation flow. The input is the tree diagram data, and the output is the tree diagram displayed on the display.
[0442] Step 13:
[0443] The terminal provides a self-practice mode for the user to practice speaking.
[0444] Specific behavior: The device displays a practice interface and presents speech samples to the user. The input is the user's selection, and the output is the self-practice mode interface.
[0445] Step 14:
[0446] The server collects and analyzes user speech data in real time.
[0447] Specific operations: Receives collected speech data from the device, performs voice analysis and intonation analysis, generates feedback based on the analysis results, and sends it to the device. The input is the collected speech data, and the output is the generated feedback.
[0448] (Application example 2)
[0449] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0450] Conventional dialogue support systems have difficulty accurately grasping the emotions and intentions of their conversation partners and providing appropriate speech plans in real time. Furthermore, in brick-and-mortar stores, staff are required to read customers' emotions and respond effectively, but current technology makes this difficult. Furthermore, feedback and visualization of conversation content using visual devices are insufficient, limiting the ability to improve users' dialogue skills.
[0451] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0452] In this invention, the server includes means for collecting facial expressions and gestures of a conversation partner using a camera, means for collecting voice data of the conversation partner using a microphone, means for analyzing the collected video and voice data to estimate the intentions and emotions of the conversation partner, means for displaying an appropriate speech plan to the user based on the estimated intentions and emotions, and means for generating an utterance plan based on the analysis results and displaying it on the user's visual device. This allows for accurate understanding of the conversation partner's emotions and intentions, enabling real-time customer service in physical stores. Furthermore, by using the visual device to provide appropriate feedback to the user and visualize the conversation content, it is expected that the user's dialogue skills will improve.
[0453] A "camera" is an optical device for collecting video data.
[0454] A "conversational partner" refers to a person other than the user with whom the conversation is taking place.
[0455] "Facial expression" refers to information including emotions and intentions expressed through the movement of facial muscles.
[0456] A "gesture" is a communication action using body and hand movements.
[0457] A "microphone" is an acoustic device for collecting audio data.
[0458] "Audio data" is digital information that digitizes audio.
[0459] "Analysis" refers to the process of extracting patterns and information from collected data.
[0460] "Intention" refers to the purpose or thoughts of the person you are speaking with.
[0461] "Emotions" are information that indicates the feelings and psychological state of the person you are talking to.
[0462] "Inference" is the process of predicting intentions and emotions from collected data.
[0463] "Utterance plan" refers to the content and plan of what the user will say.
[0464] "User" refers to a person using the smart glasses or system.
[0465] "Display" is the act of presenting information to a visual device.
[0466] A "visual device" is a device that allows a user to receive information visually.
[0467] This invention relates to a system that uses a camera, a microphone, a server, and a visual device to analyze the emotions and intentions of a conversation partner and presents an appropriate speech plan to the user. A specific implementation method of this system will be described below.
[0468] Hardware and software used
[0469] The system uses the following main hardware and software:
[0470] Camera: An optical device for collecting video data in real time, such as a camera mounted on smart glasses.
[0471] Microphone: Acoustic equipment for collecting audio data, such as a microphone in smart glasses.
[0472] Server: A high-performance computing device that performs data analysis and emotion recognition. It uses OpenCV for face recognition, Google Cloud Speech-to-Text API for voice recognition, and TensorFlow and Keras for the emotion engine.
[0473] Visual device: A device for presenting information to a user, for example, the display of smart glasses.
[0474] Data collection
[0475] First, a user puts on the smart glasses and starts a conversation with a customer. The camera captures the facial expressions and gestures of the person they are talking to, and the microphone captures audio data. This data is then sent to the server in real time.
[0476] Data analysis
[0477] The server analyzes the received video and audio data. Facial features are extracted from the video data using OpenCV, and emotions are estimated using TensorFlow and Keras. The audio data is converted to text using the Google Cloud Speech-to-Text API, and emotional changes are estimated through intonation analysis. This allows the intentions and emotions of the person being spoken to be identified.
[0478] Generate an utterance plan
[0479] The server generates a speech plan based on the analysis results. For example, if a customer is interested in a particular product but is concerned about the price, the server creates a speech plan such as, "This product offers excellent value for money and offers great benefits in the long run."
[0480] Display and Feedback
[0481] The generated speech plan is displayed on the smart glasses' display, and the user responds appropriately to the customer based on the displayed content. In this way, the conversation progresses smoothly and the customer's needs can be met.
[0482] Specific examples
[0483] For example, if you are explaining about a new smartphone in a physical store, the server may infer that the customer is interested in its features but is concerned about the price. The smart glasses will display a speech plan such as, "This smartphone uses the latest battery technology and will last all day with everyday use," and the user will explain based on that.
[0484] Example prompts to input to the generative AI model
[0485] "Generate utterance plans that recognize customer emotions and provide detailed information about products they're interested in. For example, create an utterance plan for when a customer is interested in the features of a new smartphone and has questions about the details, but is also concerned about the price."
[0486] This system will enable real-time customer service in physical stores by accurately understanding the emotions and intentions of the person you are talking to. Furthermore, by using visual devices to provide appropriate feedback to users and visualize the content of the conversation, it is expected that users' conversation skills will improve.
[0487] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0488] Step 1:
[0489] Data collection
[0490] The device uses a camera and microphone to collect facial expressions, gestures, and voice data of the conversation partner and the user. Specifically, the camera captures video data of the conversation partner and the user in real time, and the microphone collects voice data. This collected data is then sent from the device to a server.
[0491] Input: Camera video, microphone audio
[0492] Output: Video data, audio data
[0493] Step 2:
[0494] Video Data Analysis
[0495] The server analyzes the received video data and extracts facial features of the user and the conversation partner. It uses facial recognition technology to identify facial expressions and estimate emotions based on these. Specifically, it uses OpenCV for facial recognition and inputs the features into an emotion recognition model built with TensorFlow and Keras.
[0496] Input: Video data
[0497] Output: Facial features, emotion estimation results
[0498] Step 3:
[0499] Voice data analysis
[0500] The server analyzes the received voice data, converts it into text using speech recognition technology, and estimates emotional changes through intonation analysis. Specifically, the server converts the voice data into text using the Google Cloud Speech-to-Text API and then performs intonation analysis.
[0501] Input: Audio data
[0502] Output: Text data, emotion change estimation results
[0503] Step 4:
[0504] Integrated analysis of intentions and emotions
[0505] The server integrates the facial expression features, emotion estimation results, text data, and emotion change estimation results to comprehensively estimate the intentions and emotions of the conversation partner and the user.The server uses an emotion engine to comprehensively analyze the data and estimate the intentions.
[0506] Input: Facial features, emotion estimation results, text data, emotion change estimation results
[0507] Output: Intention estimation result, emotion estimation result
[0508] Step 5:
[0509] Generate an utterance plan
[0510] Based on the analysis results, the server generates a speech plan that matches the emotional state and intentions of the conversation partner and the user. Specifically, it uses a generative AI model to create an appropriate speech plan.
[0511] Input: Intention estimation result, emotion estimation result
[0512] Output: Utterance plan
[0513] Step 6:
[0514] Viewing the utterance plan
[0515] The user's device (smart glasses) displays the speech plan sent from the server on its display, allowing the user to proceed with the conversation according to the content displayed on the display. Specifically, text and icons are displayed on the display.
[0516] Input: Utterance plan
[0517] Output: Display
[0518] Step 7:
[0519] Feedback and conversation organization
[0520] The server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the terminal. The user's terminal displays this tree diagram, providing visual guidance on the next step in the conversation.
[0521] Input: Conversation progress data
[0522] Output: Tree diagram (conversation development diagram)
[0523] Step 8:
[0524] Self-Practice Mode
[0525] The device provides a self-practice mode, allowing users to practice speaking by themselves. The server provides speech samples, collects and analyzes the user's speech data in real time, and generates and sends feedback to the device based on the analysis results.
[0526] Input: User utterance data (self-practice mode)
[0527] Output: Feedback
[0528] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0529] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0530] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0531] [Second embodiment]
[0532] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0533] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0534] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0535] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0536] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0537] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0538] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0539] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0540] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0541] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0542] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0543] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0544] The present invention is a system that estimates the intentions and emotions of a conversation partner from their facial expressions, gestures, and voice, and presents an appropriate speech plan to the user, in order to facilitate interpersonal communication. Below, we will create a program for this system and explain its processing in detail.
[0545] System Configuration
[0546] The system consists of a terminal device (conversation prediction glasses) and a server device that performs data analysis. The terminal device is equipped with a camera and microphone, and the server device is equipped with an AI analysis model.
[0547] Program processing flow
[0548] Data collection
[0549] The device has a built-in camera to capture the facial expressions and gestures of the person you are talking to, and a microphone to collect audio data. When a conversation begins, the camera collects the video data of the person you are talking to, and the microphone collects the audio data in real time, and then sends the data to a server.
[0550] Data analysis
[0551] The server uses facial recognition technology to extract facial features of the conversation partner from the received video data and also analyzes gestures. At the same time, it converts the speech content into text using speech recognition technology from the audio data and estimates changes in emotion through intonation analysis. This data is input into an AI model, which then infers the other person's intentions and emotions.
[0552] Generate an utterance plan
[0553] The server generates a speech plan based on the analysis results, according to the intentions and emotions expressed by the conversation partner. For example, if a customer shows interest in a particular product, the server generates a speech plan to explain the details of that product and recommends it to the user. The generated speech plan is sent to the terminal.
[0554] User Feedback
[0555] The device displays the speech plan sent from the server on its display and provides the user with appropriate speech content. By continuing the conversation according to the displayed content, the user can communicate appropriately with the other person's intentions and emotions.
[0556] Organizing the conversation
[0557] Furthermore, the server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the device. The device displays this tree diagram, providing visual guidance on the next step in the conversation for the user.
[0558] Self-Practice Mode
[0559] The device has a self-practice mode, allowing users to practice speaking by themselves. In this mode, the server provides speaking samples, analyzes the user's speech data, and provides evaluation and feedback, allowing users to improve their speaking skills.
[0560] Specific examples
[0561] For example, a salesperson puts on these conversation-reading glasses when negotiating with a customer. After the negotiation begins, the glasses' camera captures the customer's facial expressions and gestures, and the microphone collects what the customer says, and these data are sent to a server. The server analyzes the data and infers that the customer is interested in price. The server then generates a speech plan, such as "This product has excellent cost performance and will offer great benefits in the long run," and sends it to the device. The device then displays this on its screen, and the salesperson continues to explain the plan to the customer. In this way, the conversation progresses smoothly, enabling responses to be made that meet the customer's needs.
[0562] This system makes it possible to accurately grasp the intentions and emotions of the person you are talking to, enabling smooth communication, which in turn makes business negotiations, counseling, and even everyday conversations more effective.
[0563] The processing flow will be explained below.
[0564] Step 1: Your device is ready for the camera and microphone
[0565] Before the conversation begins, the device turns on the camera and microphone to prepare for data collection.
[0566] Show the user a ready notification.
[0567] Step 2: The device collects video and audio data from the person you are talking to.
[0568] The camera continuously captures the facial expressions and gestures of the person you are talking to.
[0569] The microphone records the voice of the person you are talking to in real time.
[0570] The collected data is sent to the server in real time.
[0571] Step 3: The server analyzes the video data
[0572] The server analyzes the received video data and extracts facial expression features of the conversation partner.
[0573] Use facial recognition technology to estimate emotions from facial expressions.
[0574] Gesture analysis is performed to infer intentions from hand and body movements.
[0575] Step 4: The server analyzes the audio data
[0576] The server analyzes the received voice data and converts the spoken content into text using voice recognition technology.
[0577] The text of the speech is analyzed, and emotions are inferred from word choice and intonation.
[0578] Step 5: The server estimates the other person's intentions and emotions
[0579] The features extracted from the video and audio data are input into the AI model.
[0580] The AI model estimates the overall intention and emotions of the conversation partner.
[0581] Step 6: The server generates an appropriate utterance plan
[0582] Based on the estimation results, the AI generates the optimal speech plan for the user.
[0583] A speech plan includes specific phrases and next steps in the conversation.
[0584] Step 7: The server sends the speech plan to the device.
[0585] The server transmits the generated speech plan to the terminal.
[0586] Optimize and transmit data to ensure stable communication and prevent delays.
[0587] Step 8: The device displays the speech plan to the user.
[0588] The terminal displays the speech plan received from the server on the display.
[0589] The user continues the conversation by following what is displayed.
[0590] Step 9: The server organizes the conversation
[0591] As the conversation progresses, the server organizes the user's comments and the analysis results.
[0592] Structure the conversation content as a tree diagram or flowchart.
[0593] Step 10: The server sends the organized conversation to the device.
[0594] The organized conversation content and the next steps to take are sent to the terminal.
[0595] Help users make good decisions.
[0596] Step 11: The terminal displays the tree diagram to the user
[0597] The terminal displays the tree diagram or flowchart received from the server.
[0598] Check what the user should say next and the flow.
[0599] Step 12: User selects and starts self-practice mode
[0600] The user selects the self-practice mode through the terminal interface.
[0601] An interface provides the user with practice mode settings.
[0602] Step 13: The server provides and analyzes speech samples
[0603] The server provides speech samples for use in self-practice mode.
[0604] Data spoken by users is collected in real time and analyzed.
[0605] Step 14: The server generates feedback and sends it to the device
[0606] The server evaluates the user's speech data and generates suggestions for improvement and appropriate feedback.
[0607] The generated feedback is sent to the device.
[0608] Step 15: The device displays feedback to the user
[0609] The terminal displays the feedback received from the server to the user.
[0610] The user improves their speaking skills based on the feedback.
[0611] Through these processing steps, users can properly understand the intentions and emotions of their conversation partners and communicate smoothly. In addition, by utilizing the self-practice mode, users can effectively improve their speaking skills.
[0612] Example 1
[0613] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0614] Interpersonal communication requires rapid and accurate understanding of the intentions and emotions of the person you are speaking with, but the current technology for doing so is insufficient. It is particularly difficult to respond appropriately to changes in the other person's intentions and emotions in situations such as business negotiations and counseling. There is also a lack of effective feedback systems for improving speaking skills through self-practice. A system that solves these issues is needed.
[0615] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0616] In this invention, the server includes means for collecting facial expressions and gestures of a conversation partner using a camera, means for collecting voice data of the conversation partner using a microphone, means for transmitting the collected video and voice data to the server in real time, means for extracting facial features from the received video data using face recognition technology and performing gesture analysis, means for converting the speech content into text from the voice data using voice recognition technology and estimating emotional changes through intonation analysis, and means for generating and displaying an appropriate speech plan for the user based on the estimated intentions and emotions. This makes it possible to accurately grasp the intentions and emotions of the conversation partner and achieve smooth communication.
[0617] A "camera" is a device for collecting video data, and in this invention is used to capture the facial expressions and gestures of a conversation partner.
[0618] A "microphone" is a device for collecting voice data, and in the present invention is used to clearly capture the voice of a conversation partner.
[0619] A "server" is a computer system for analyzing and processing data, and in the present invention, it has the role of analyzing video data and audio data and providing speech plans to users.
[0620] "Facial expression features" are a set of feature values that indicate the emotions and intentions of a conversation partner, extracted using face recognition technology.
[0621] "Gesture analysis" is the process of analyzing a conversation partner's hand movements and gestures to infer non-verbal intentions and emotions.
[0622] "Speech recognition technology" is a technology that analyzes voice data and converts it into text data, and is used in the present invention to analyze the content of speech from a conversation partner.
[0623] "Intonation analysis" is a technology that analyzes the intonation and rhythm of voice data to estimate changes in emotion.
[0624] An "utterance plan" is a proposal of utterance content to support the progress of a conversation, generated based on analyzed data.
[0625] A "tree diagram" is a diagram that organizes the development of a conversation in a hierarchical structure to visually show the progress of the conversation.
[0626] The "self-practice mode" is a function that allows the user to practice speaking on their own, and includes providing speaking samples and feedback.
[0627] The present invention is a system that estimates the intentions and emotions of a conversation partner from their facial expressions, gestures, and voice, and presents an appropriate speech plan to the user, in order to facilitate interpersonal communication. Below, we will create a program for this system and explain its processing in detail.
[0628] System Configuration
[0629] The main components of this system are a terminal device (conversation prediction glasses) and a server device that performs data analysis. The terminal device is equipped with a camera and microphone, and the server device is equipped with an AI analysis model.
[0630] Program processing flow and specific functions
[0631] Data collection
[0632] The device has a built-in camera for capturing facial expressions and gestures of the conversation partner, and a microphone for capturing audio data. When a conversation begins, the camera captures the conversation partner's video data and the microphone captures audio data in real time, and these data are then sent to a server. The hardware used is a built-in camera (e.g., an HD camera) and a highly sensitive microphone.
[0633] Data analysis
[0634] The server uses facial recognition technology to extract facial features of the conversation partner from the received video data and also performs gesture analysis. At the same time, it uses speech recognition technology to convert the speech into text from the audio data and estimates emotional changes through intonation analysis. Software libraries such as TensorFlow and OpenCV are used for these analyses. For example, Google's Speech-to-Text API is sometimes used to analyze audio data.
[0635] Generate an utterance plan
[0636] The server generates a speech plan based on the analysis results, taking into account the intentions and emotions of the conversation partner. This is done using a pre-trained generative AI model. For example, if a customer shows interest in a particular product, a speech plan to explain the details of that product is generated. The generated speech plan is sent to the device.
[0637] User Feedback
[0638] The device displays the speech plan sent from the server on its display and provides the user with appropriate speech content. By continuing the conversation according to the displayed content, the user can communicate appropriately with the other person's intentions and emotions.
[0639] Organizing the conversation
[0640] Furthermore, the server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the device. The device displays this tree diagram, providing visual guidance on the next step in the conversation for the user.
[0641] Self-Practice Mode
[0642] The device has a self-practice mode, allowing users to practice speaking by themselves. In this mode, the server provides speaking samples, analyzes the user's speech data, and provides evaluation and feedback, allowing users to improve their speaking skills.
[0643] Specific examples
[0644] For example, a salesperson puts on these conversation-reading glasses when negotiating with a customer. After the negotiation begins, the glasses' camera captures the customer's facial expressions and gestures, and the microphone collects what the customer says, and these data are sent to a server. The server analyzes the data and infers that the customer is interested in price. The server then generates a speech plan, such as "This product has excellent cost performance and will offer great benefits in the long run," and sends it to the device. The device then displays this on its screen, and the salesperson continues to explain the plan to the customer. In this way, the conversation progresses smoothly, enabling responses to be made that meet the customer's needs.
[0645] Example prompts for generative AI models
[0646] Input data: Customer facial expression information, voice data
[0647] AI model: Estimating conversation partner's intentions and emotions
[0648] Output: It is assumed that the customer is interested in the price. Generate the utterance plan "This product offers excellent value for money and offers great benefits in the long run."
[0649] As described above, the present invention is a system for understanding the intentions and emotions of a conversation partner and supporting appropriate communication, and is particularly effective in situations such as business negotiations and counseling.
[0650] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0651] Step 1:
[0652] Data collection
[0653] The device uses a camera and microphone to capture facial expressions, gestures, and voice data from the person you're talking to. The camera captures high-resolution video, while the microphone suppresses background noise to capture clear audio.
[0654] input:
[0655] Video and audio data of the person you are talking to
[0656] output:
[0657] Collected video and audio data
[0658] Specific behavior:
[0659] When the user begins interacting with a customer, the device's camera captures the customer's facial expressions and gestures, and the microphone begins recording what the customer says.
[0660] Step 2:
[0661] Data transmission
[0662] The device transmits the collected video and audio data to the server in real time over a fast and secure network.
[0663] input:
[0664] Collected video and audio data
[0665] output:
[0666] Data sent to the server
[0667] Specific behavior:
[0668] The device encrypts the collected data and sends it to a server using Wi-Fi or 5G networks.
[0669] Step 3:
[0670] Data analysis
[0671] The server extracts facial features from the received video data using facial recognition technology and also performs gesture analysis. It also converts the speech into text using speech recognition technology and estimates emotional changes through intonation analysis. TensorFlow and OpenCV are used as software libraries.
[0672] input:
[0673] Video and audio data sent to the server
[0674] output:
[0675] Extracted facial features, textualized speech, and estimated emotional changes
[0676] Specific behavior:
[0677] The server uses TensorFlow to analyze the video data and extract facial expression and gesture features. At the same time, Google's Speech-to-Text API converts the audio data into text, and the emotion analysis API analyzes emotions from intonation.
[0678] Step 4:
[0679] Generate an utterance plan
[0680] The server generates a speech plan based on the analysis results, taking into account the intentions and emotions of the conversation partner. Using a generative AI model, it constructs an appropriate speech plan that corresponds to the estimated intentions and emotions.
[0681] input:
[0682] Extracted facial features, textualized speech, and estimated emotional changes
[0683] output:
[0684] Generated utterance plan
[0685] Specific behavior:
[0686] If the server's AI model determines that the customer is interested in price, it generates a speech plan such as, "This product has excellent cost performance."
[0687] Step 5:
[0688] Sending a speech plan
[0689] The server then sends the generated speech plan to the device, again in real time, allowing the user to receive immediate feedback.
[0690] input:
[0691] Generated utterance plan
[0692] output:
[0693] Speech plan sent to the device
[0694] Specific behavior:
[0695] The speech plan generated by the server is immediately sent to the device, and a notification is displayed on the device's display.
[0696] Step 6:
[0697] User Feedback
[0698] The device displays the received speech plan on its display and provides the user with appropriate speech content, which the user can refer to as they proceed with the conversation.
[0699] input:
[0700] Speech plan sent to the device
[0701] output:
[0702] The speech plan presented to the user
[0703] Specific behavior:
[0704] The terminal display will show "This product has excellent cost performance," and the user will convey this statement to the customer.
[0705] Step 7:
[0706] Organizing the conversation
[0707] The server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the terminal.
[0708] input:
[0709] Analyzed conversation content
[0710] output:
[0711] Speech content organized as a tree diagram
[0712] Specific behavior:
[0713] The server analyzes the flow of conversation, and if there is a high level of "interest in price," it generates a node for "additional information related to price" and displays it in the form of a tree diagram on the terminal display.
[0714] Step 8:
[0715] Self-Practice Mode
[0716] The device has a self-practice mode, allowing users to practice speaking by themselves. In this mode, the server provides speaking samples and analyzes the user's speech data to provide evaluation and feedback.
[0717] input:
[0718] Speech samples entered by the user
[0719] output:
[0720] Ratings and Feedback
[0721] Specific behavior:
[0722] When a user selects practice mode, the server sends sample utterances, such as "greeting practice," to the device. As the user speaks, the audio is analyzed and pronunciation and intonation are evaluated in real time.
[0723] (Application example 1)
[0724] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0725] Existing conversation support systems have difficulty understanding the intentions and emotions of the target, and are therefore unable to present appropriate speech plans or response methods for achieving effective communication. Furthermore, in security services, where it is necessary to quickly understand the intentions and emotions of visitors and passersby and respond appropriately, the lack of such technology poses a serious problem.
[0726] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0727] In this invention, the server includes means for collecting facial expressions and gestures of a target using a camera, means for collecting audio data of the target using a microphone, means for analyzing the collected video and audio data to estimate the intention and emotion of the target, means for displaying an appropriate speech plan to the user based on the estimated intention and emotion, and means for suggesting a response method to the user based on the estimated intention and emotion, thereby enabling security guards to quickly grasp the intention and emotion of visitors and passersby and to determine an appropriate response method.
[0728] A "camera" is a photographic device for capturing facial expressions and gestures of a subject.
[0729] A "microphone" is a recording device for collecting audio data of a subject.
[0730] "Video data" refers to image and video information collected using a camera.
[0731] "Audio data" is sound information collected using a microphone.
[0732] "Facial expression" refers to emotions and intentions that can be read from the movement of facial muscles.
[0733] A "gesture" is an expression of intent that can be interpreted through hand or body movements.
[0734] "Intention" refers to the purpose or direction of an action that a subject has.
[0735] "Emotion" refers to a subject's state of mind or mood.
[0736] A "speech plan" is a plan for providing appropriate conversation content.
[0737] A "display" is a display device for presenting visual information to a user.
[0738] A "server" is a computer system that analyzes collected data and provides results.
[0739] A "tree diagram" is a diagram that visually organizes the content and development of a conversation.
[0740] The "self-practice mode" is a mode in which the user can practice speaking by himself.
[0741] "Feedback" refers to evaluation of the user's speech and advice for improvement.
[0742] This invention is an interpersonal communication support system specialized for security applications, and employs an approach to infer the intentions and emotions of a conversation partner from their facial expressions, gestures, and voice, and provide appropriate response methods.
[0743] System Configuration
[0744] The system includes smart glasses (terminal devices) and a server device that performs data analysis. The smart glasses are equipped with a camera and microphone, and the server device is equipped with an AI analysis model.
[0745] Data collection
[0746] The terminal device is equipped with a camera to capture the facial expressions and gestures of the visitor and a microphone to collect audio data. When a security guard comes into contact with a visitor, the camera collects the subject's video data and the microphone collects audio data in real time, and these data are then sent to a server.
[0747] Data analysis
[0748] The server uses facial recognition technology to extract facial features from the received video data and also analyzes gestures. At the same time, it converts the speech into text using speech recognition technology and estimates emotional changes through intonation analysis. This data is input into an AI model, which then infers the subject's intentions and emotions.
[0749] Generate speech plans and responses
[0750] Based on the analysis results, the server generates a speech plan and appropriate response method according to the visitor's intentions and emotions. For example, if the visitor seems nervous, the server generates advice recommending polite responses and sends it to the terminal device. The generated speech plan and response method are then displayed on the smart glasses' display.
[0751] User Feedback
[0752] The terminal device displays the speech plan and response method sent from the server on its display and provides the security guard with appropriate response content. By responding to the visitor according to the displayed content, the security guard can appropriately deal with the visitor's intentions and emotions.
[0753] Organizing the conversation
[0754] Furthermore, the server analyzes the progress of the conversation in real time, organizes the content and development of the conversation into a tree diagram, and sends it to the terminal device, which displays this tree diagram and provides visual guidance to the security guard on the next steps in the conversation.
[0755] Self-Practice Mode
[0756] The terminal device is equipped with a self-practice mode, allowing security guards to practice speaking by themselves. In this mode, the server provides speaking samples, analyzes the security guards' speech data, and provides evaluation and feedback, allowing the security guards to improve their speaking skills.
[0757] Examples and prompts
[0758] Specific examples
[0759] A security guard wears smart glasses at the entrance to a facility and handles visitors. If the visitor's facial expression appears stiff and tense, the camera and microphone collect information and the information is analyzed by a server. As a result of the analysis, the smart glasses' display displays advice such as, "It is assumed that the visitor is nervous. Please handle the visitor politely and confirm the purpose of their visit."
[0760] Prompt Sentence Examples
[0761] "Advice to display when a visitor's facial expression appears tense:
[0762] We suspect the visitor may be nervous. Please be polite and confirm the purpose of their visit."
[0763] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0764] Step 1:
[0765] Data collection
[0766] The device (smart glasses) uses a camera to collect video data of the subject's facial expressions and gestures, and simultaneously collects audio data using a microphone. The input is video data from the camera and audio data from the microphone, and the output is the collected raw video and audio data.
[0767] Step 2:
[0768] Data transmission
[0769] The terminal transmits the collected video and audio data to the server in real time. The input is the collected video and audio data, and the output is the data transmitted to the server.
[0770] Step 3:
[0771] Facial recognition and facial expression feature extraction
[0772] The server uses facial recognition technology to detect the target face from the received video data and extract facial expression features. The input is the video data, and the output is the extracted facial expression features. The software used is a facial recognition library such as OpenCV.
[0773] Step 4:
[0774] Gesture Analysis
[0775] The server analyzes hand and body movements from video data to detect gestures. The input is video data, and the output is detected gesture data. The software used is an AI analysis model.
[0776] Step 5:
[0777] Converting audio data to text
[0778] The server converts the voice data into text data using speech recognition technology. The input is voice data and the output is text data. The software used is a speech recognition library (e.g., Google Cloud Speech-to-Text).
[0779] Step 6:
[0780] Intonation analysis
[0781] The server analyzes the intonation of the voice data and estimates emotional changes. The input is the voice data, and the output is estimated emotional data. The software used is an AI analysis model.
[0782] Step 7:
[0783] Intention and emotion estimation
[0784] The server inputs the extracted facial features, gesture data, text data, and emotion data into an AI model to comprehensively estimate the subject's intention and emotion. The input is each feature data, and the output is estimated intention and emotion information.
[0785] Step 8:
[0786] Generate speech plans and responses
[0787] The server generates a speech plan and a response method for the user based on the estimated intention and emotion. For example, if the visitor is nervous, it generates advice including an appropriate response method. The input is the intention and emotion information, and the output is the speech plan and a response method.
[0788] Step 9:
[0789] User Feedback
[0790] The terminal displays the speech plan and response method sent from the server on a display and provides them to the security guard. The input is the speech plan and response method, and the output is the content displayed on the display. The security guard responds to the visitor according to the displayed content.
[0791] Step 10:
[0792] Organizing the conversation
[0793] The server analyzes ongoing conversations in real time and organizes them into a tree diagram. The input is real-time conversation data, and the output is a tree diagram of the organized conversations. The terminal displays this tree diagram to help security guards understand the next steps.
[0794] Step 11:
[0795] Self-Practice Mode
[0796] The terminal provides an interface that allows the user to select self-practice mode. The server provides speech samples, analyzes the user's speech data, and provides evaluation and feedback. The input is the user's speech data, and the output is evaluation and feedback, allowing the security guard to improve their speaking skills.
[0797] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0798] This invention is a system that combines conversation prediction glasses with an emotion engine to more accurately recognize the emotions and intentions of both the user and the conversation partner and present appropriate speech plans. Below, we will create a program for this system and explain its processing in detail.
[0799] System Configuration
[0800] The system includes a terminal device (conversation prediction glasses), a server device that performs data analysis, and an emotion engine with emotion recognition capabilities. The terminal device is equipped with a camera and microphone, and the server device is equipped with an AI analysis model and emotion engine.
[0801] Program processing flow
[0802] Data collection
[0803] The device uses a camera and microphone to collect facial expressions, gestures, and voice data of the user and the conversation partner. When a conversation begins, the camera collects video data of the conversation partner and the user, and the microphone collects voice data in real time, and then transmits this data to a server.
[0804] Data analysis
[0805] The server analyzes the received video and audio data and extracts facial and audio features of the conversation partner and the user. First, facial recognition technology is used to infer emotions from facial expressions, followed by gesture analysis. Furthermore, speech recognition technology is used to convert the speech into text from the audio data, and emotional changes are inferred through intonation analysis. This data is input into an AI model and emotion engine, which infers the intentions and emotions of the conversation partner and the user.
[0806] Generate an utterance plan
[0807] The server generates a speech plan based on the analysis results, taking into account the emotional state and intentions of the conversation partner and the user. For example, the server may generate a speech plan that provides detailed explanations based on a topic that the conversation partner is interested in, or suggest relaxing topics to reduce the stress the user is feeling. The generated speech plan is then sent to the device.
[0808] User Feedback
[0809] The device displays the speech plan sent from the server on its display and provides the user with appropriate speech content. By continuing the conversation according to the displayed content, the user can communicate appropriately with the other person's intentions and emotions.
[0810] Organizing the conversation
[0811] Furthermore, the server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the device. The device displays this tree diagram, providing visual guidance on the next step in the conversation for the user.
[0812] Self-Practice Mode
[0813] The device offers a self-practice mode, allowing users to practice speaking by themselves. In this mode, the server provides speech samples, collects and analyzes the user's speech data in real time, and generates and sends feedback to the device based on the analysis results.
[0814] Specific examples
[0815] For example, a sales representative puts on these conversation-predicting glasses when negotiating with a customer. After the negotiation begins, the glasses' camera captures the facial expressions and gestures of both the customer and the sales representative, while the microphone collects speech, and these data are sent to a server. The server analyzes the data and infers that the customer is beginning to show interest in the product but is concerned about the price. The emotion engine then confirms that the sales representative is explaining with confidence. The server generates a speech plan, such as "This product has excellent cost performance and will offer great benefits in the long run," and sends it to the device. The device displays this on its screen, and the sales representative explains based on it. In this way, the conversation progresses smoothly, enabling responses that meet the customer's needs.
[0816] This system allows both the user and the conversation partner to accurately understand each other's intentions and emotions, enabling smooth and effective communication, which will enable more effective business negotiations, counseling, and everyday conversations.
[0817] The processing flow will be explained below.
[0818] Step 1: Your device is ready for the camera and microphone
[0819] Before the conversation begins, the device turns on the camera and microphone to prepare for data collection.
[0820] Show the user a ready notification.
[0821] Step 2: The device collects video and audio data of the conversation partner and the user.
[0822] The camera continuously captures the facial expressions and gestures of the conversation partner and the user.
[0823] The microphone records the voices of the conversation partner and the user in real time.
[0824] The collected data is sent to the server in real time.
[0825] Step 3: The server analyzes the video data of the person you are talking to.
[0826] The server analyzes the received video data and extracts facial features of the conversation partner.
[0827] Use facial recognition technology to estimate emotions from facial expressions.
[0828] Gesture analysis is performed to infer intentions from hand and body movements.
[0829] Step 4: The server analyzes the voice data of the other party.
[0830] The server analyzes the received voice data and converts the spoken content into text using voice recognition technology.
[0831] The text of the speech is analyzed, and emotions are inferred from word choice and intonation.
[0832] Step 5: The server analyzes the user's video data.
[0833] The server analyzes the received video data and extracts the user's facial expression features.
[0834] Use facial recognition technology to estimate emotions from facial expressions.
[0835] Gesture analysis is performed to estimate the current state and intentions from hand and body movements.
[0836] Step 6: The server analyzes the user's voice data
[0837] The server analyzes the received voice data and converts the spoken content into text using voice recognition technology.
[0838] The system analyzes the text of the speech and estimates the user's emotions from word choice and intonation.
[0839] Step 7: The server integrates and estimates the intentions and emotions of the conversation partner and the user.
[0840] The features extracted from the video and audio data are input into the AI model and emotion engine.
[0841] AI models and emotion engines estimate the overall intent and emotions of the conversation partner and the user.
[0842] Step 8: The server generates an appropriate utterance plan
[0843] Based on the estimation results, the AI generates the optimal speech plan for the user.
[0844] A speech plan includes specific phrases and next steps in the conversation.
[0845] Step 9: The server sends the speech plan to the device.
[0846] The server transmits the generated speech plan to the terminal.
[0847] Optimize and transmit data to ensure stable communication and prevent delays.
[0848] Step 10: The device displays the speech plan to the user.
[0849] The terminal displays the speech plan received from the server on the display.
[0850] The user continues the conversation by following what is displayed.
[0851] Step 11: The server organizes the conversation
[0852] As the conversation progresses, the server organizes the user's comments and the analysis results.
[0853] Structure the conversation content as a tree diagram or flowchart.
[0854] Step 12: The server sends the organized conversation to the device.
[0855] The organized conversation content and the next steps to take are sent to the terminal.
[0856] Help users make good decisions.
[0857] Step 13: The terminal displays the tree diagram to the user
[0858] The terminal displays the tree diagram or flowchart received from the server.
[0859] Check what the user should say next and the flow.
[0860] Step 14: User selects and starts self-practice mode
[0861] The user selects the self-practice mode through the terminal interface.
[0862] An interface provides the user with practice mode settings.
[0863] Step 15: The server provides and analyzes speech samples.
[0864] The server provides speech samples for use in self-practice mode.
[0865] Data spoken by users is collected in real time and analyzed.
[0866] Step 16: The server generates feedback and sends it to the device
[0867] The server evaluates the user's speech data and generates suggestions for improvement and appropriate feedback.
[0868] The generated feedback is sent to the device.
[0869] Step 17: The device displays feedback to the user
[0870] The terminal displays the feedback received from the server to the user.
[0871] The user improves their speaking skills based on the feedback.
[0872] Through these processing steps, users can accurately understand their own and their conversation partner's emotions and intentions, enabling smoother and more effective communication. Furthermore, by utilizing the self-practice mode, users can effectively improve their speaking skills.
[0873] Example 2
[0874] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0875] In conversation, users often have difficulty accurately reading the intentions and emotions of the other person, making smooth communication difficult. They also often lack appropriate responses to the anxiety and stress they feel. Furthermore, even when practicing by themselves, the lack of concrete feedback makes it difficult to improve their speech.
[0876] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0877] In this invention, the server includes means for collecting facial expressions and gestures of the conversation partner and the user using a camera, means for collecting voice data of the conversation partner and the user using a microphone, means for analyzing the collected video and voice data to estimate the intentions and emotions of the conversation partner and the user, means for displaying an appropriate speech plan to the user based on the estimated intentions and emotions, means for organizing the progress of the conversation in a tree diagram based on the analysis results and displaying it to the user, and means for providing an interface that allows the user to select a self-practice mode and receive feedback for improving their speech. This allows the user to accurately understand the intentions and emotions of the conversation partner and take appropriate measures. Furthermore, the self-practice mode makes it easier to improve speech, thereby improving the overall quality of communication.
[0878] A "camera" is an optical device for collecting video data.
[0879] A "conversational partner" refers to the other person with whom the user is interacting.
[0880] "User" refers to a person who uses this system to have a conversation.
[0881] "Facial expressions" refer to facial movements and expressions, and are important non-verbal elements for indicating emotions and intentions.
[0882] "Gestures" refer to hand and body movements and are non-verbal expressions used to convey intentions and emotions.
[0883] A "microphone" is an acoustic device for collecting audio data.
[0884] "Audio data" refers to collected sound information, including speech content and intonation.
[0885] "Collect" refers to obtaining and storing data.
[0886] "Analysis" is the act of extracting and analyzing information based on collected data.
[0887] "Intention" refers to the purpose or thoughts of the conversation partner or user.
[0888] "Emotion" refers to the emotional state felt by a conversation partner or user.
[0889] A "speech plan" is a specific speech content suggested by the system based on a specific intention or emotion.
[0890] "Display" is the act of enabling a user to visually confirm information.
[0891] A "tree diagram" is a diagram that visually shows the progress of a conversation and the relationship between topics.
[0892] "Interface" refers to the operating screen or means by which a user interacts with a system.
[0893] "Feedback" refers to information that provides suggestions for improvement or evaluation of a user's actions or results.
[0894] The present invention is a system that includes a terminal device (conversation prediction glasses) worn by a user, a server device that performs data analysis, and an emotion engine with emotion recognition functionality. Specific embodiments for implementing this system are described below.
[0895] terminal device
[0896] The terminal device is equipped with a camera and a microphone. These hardware components are used to collect video and audio data of the user and the conversation partner in real time. Specifically, the camera captures facial expressions and gestures, and the microphone collects audio data of the conversation. The terminal device has the function of transmitting the collected data to a server device via a network.
[0897] Server device
[0898] The server device is equipped with an AI model and emotion engine for analyzing the received data. From the video data, facial recognition technology is used to extract facial expression features and perform gesture analysis. For the audio data, speech is converted into text using speech recognition technology, and emotional changes are estimated through intonation analysis. This data is input into the AI model and emotion engine, which ultimately estimates the intentions and emotions of the conversation partner and the user.
[0899] Providing Feedback
[0900] The server generates a speech plan based on the analysis results and transmits it to the terminal device. The terminal device displays this speech plan on a display and provides it to the user. The user can continue the conversation based on the displayed speech plan, and can respond appropriately to the intentions and feelings of the conversation partner.
[0901] Organizing the conversation
[0902] The server analyzes the progress of the conversation in real time and organizes the content and development of the conversation into a tree diagram. This tree diagram is sent to the terminal device, helping the user visually grasp the next flow of the conversation.
[0903] Self-Practice Mode
[0904] The terminal device is equipped with an interface that allows the user to select a self-practice mode, in which the server collects the user's speech, analyzes it in real time, and generates feedback, which helps the user improve their speech.
[0905] Specific examples
[0906] For example, a sales representative puts on conversation-predicting glasses when negotiating with a customer. When the negotiation begins, the device's camera captures the facial expressions and gestures of both the customer and the sales representative, and the microphone collects the audio of the conversation. This data is sent to a server, which uses analysis technology to infer the customer's emotions and intentions. Based on the analysis results, a speech plan such as "This product has excellent cost performance and will provide great benefits in the long run" is generated and displayed on the device. The user can explain the situation based on the displayed speech plan, allowing the negotiation to proceed smoothly.
[0907] Prompt Sentence Examples
[0908] "Using conversation-reading glasses, please generate a specific scenario in which a salesperson is negotiating with a customer. Include a process for reading the customer's emotions and interests and proposing an appropriate speech plan."
[0909] This system accurately recognizes the intentions and emotions of both the user and the conversation partner, enabling accurate and smooth communication, promoting effective dialogue in business negotiations, counseling, and everyday conversations.
[0910] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0911] Step 1:
[0912] The terminal collects video and audio of the conversation partner and the user in real time.
[0913] Specific operation: The device's camera captures the face, facial expressions, and gestures of the user and the person in conversation, and the microphone collects audio. The input is the video and audio data acquired through the camera and microphone, and the output is the collected data packets.
[0914] Step 2:
[0915] The terminal transmits the collected video and audio data to the server.
[0916] Specific operation: Video data is packetized as image frames, and audio data is packetized as audio clips, and then sent to the server over the network. The input is the collected data packets, and the output is the server to which the data was sent.
[0917] Step 3:
[0918] The server analyzes the received video data and extracts facial expression features.
[0919] Specific operation: Using facial recognition technology in the server, the system identifies the facial parts of the conversation partner and the user and calculates facial features such as smile, sadness, surprise, etc. The input is the received video data, and the output is the extracted facial features.
[0920] Step 4:
[0921] The server performs gesture analysis and infers emotions and intentions from the collected gesture data.
[0922] Specific operation: Using the server's gesture recognition algorithm, hand movements and body poses are analyzed to estimate emotional states such as interest or tension. The input is the received video data and existing gesture data, and the output is the estimated emotion or intention.
[0923] Step 5:
[0924] The server analyzes the voice data and converts the spoken content into text.
[0925] How it works: The speech recognition system analyzes recorded speech and converts it into text data. It then detects emotional changes through intonation analysis. The input is the received speech data, and the output is the text of the speech and an indicator of emotional changes.
[0926] Step 6:
[0927] The server uses AI models and emotion engines to infer the intentions and emotions of the user and their conversation partner.
[0928] How it works: The extracted features are input into an AI model based on previous research and training data to estimate emotions and intentions. An emotion engine is used to further improve accuracy. The inputs are facial features, gesture data, text data of spoken content, and indicators of emotional changes, and the output is estimated intentions and emotions.
[0929] Step 7:
[0930] Based on the analysis results, the server generates a speech plan that is appropriate for the emotions and intentions of the user and their conversation partner.
[0931] How it works: Using a generative AI model, it automatically determines the appropriate utterance content for the situation (e.g., question, explanation, relaxed topic, etc.) and constructs it as an utterance plan. The input is the estimated intent and emotion, and the output is the generated utterance plan.
[0932] Step 8:
[0933] The server transmits the generated speech plan to the terminal.
[0934] Specific operation: The text information generated as a speech plan is packetized and sent to the terminal via the network. The input is the generated speech plan, and the output is the terminal to which it was sent.
[0935] Step 9:
[0936] The terminal displays the received speech plan on the display.
[0937] Specific operation: The display module in the terminal displays the text of the speech plan on the screen, providing a visual for the user. The input is the received speech plan, and the output is the speech plan displayed on the display.
[0938] Step 10:
[0939] The user continues the conversation according to the content displayed on the display.
[0940] Specific operation: The user refers to the displayed speech plan and speaks at the appropriate time to progress the conversation. The input is the speech plan displayed on the display, and the output is the ongoing conversation.
[0941] Step 11:
[0942] The server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the terminal.
[0943] Specific operation: The collected data is analyzed one after another, and data for visualizing the conversation topics and progress as a tree diagram is generated and sent to the device. The input is the ongoing conversation data and analysis results, and the output is data organized as a tree diagram.
[0944] Step 12:
[0945] The terminal displays a tree diagram showing the progress of the conversation on its display.
[0946] Specific operation: The display module in the terminal displays the tree diagram on the screen, providing the user with a visual of the next conversation flow. The input is the tree diagram data, and the output is the tree diagram displayed on the display.
[0947] Step 13:
[0948] The terminal provides a self-practice mode for the user to practice speaking.
[0949] Specific behavior: The device displays a practice interface and presents speech samples to the user. The input is the user's selection, and the output is the self-practice mode interface.
[0950] Step 14:
[0951] The server collects and analyzes user speech data in real time.
[0952] Specific operations: Receives collected speech data from the device, performs voice analysis and intonation analysis, generates feedback based on the analysis results, and sends it to the device. The input is the collected speech data, and the output is the generated feedback.
[0953] (Application example 2)
[0954] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0955] Conventional dialogue support systems have difficulty accurately grasping the emotions and intentions of their conversation partners and providing appropriate speech plans in real time. Furthermore, in brick-and-mortar stores, staff are required to read customers' emotions and respond effectively, but current technology makes this difficult. Furthermore, feedback and visualization of conversation content using visual devices are insufficient, limiting the ability to improve users' dialogue skills.
[0956] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0957] In this invention, the server includes means for collecting facial expressions and gestures of a conversation partner using a camera, means for collecting voice data of the conversation partner using a microphone, means for analyzing the collected video and voice data to estimate the intentions and emotions of the conversation partner, means for displaying an appropriate speech plan to the user based on the estimated intentions and emotions, and means for generating an utterance plan based on the analysis results and displaying it on the user's visual device. This allows for accurate understanding of the conversation partner's emotions and intentions, enabling real-time customer service in physical stores. Furthermore, by using the visual device to provide appropriate feedback to the user and visualize the conversation content, it is expected that the user's dialogue skills will improve.
[0958] A "camera" is an optical device for collecting video data.
[0959] A "conversational partner" refers to a person other than the user with whom the conversation is taking place.
[0960] "Facial expression" refers to information including emotions and intentions expressed through the movement of facial muscles.
[0961] A "gesture" is a communication action using body and hand movements.
[0962] A "microphone" is an acoustic device for collecting audio data.
[0963] "Audio data" is digital information that digitizes audio.
[0964] "Analysis" refers to the process of extracting patterns and information from collected data.
[0965] "Intention" refers to the purpose or thoughts of the person you are speaking with.
[0966] "Emotions" are information that indicates the feelings and psychological state of the person you are talking to.
[0967] "Inference" is the process of predicting intentions and emotions from collected data.
[0968] "Utterance plan" refers to the content and plan of what the user will say.
[0969] "User" refers to a person using the smart glasses or system.
[0970] "Display" is the act of presenting information to a visual device.
[0971] A "visual device" is a device that allows a user to receive information visually.
[0972] This invention relates to a system that uses a camera, a microphone, a server, and a visual device to analyze the emotions and intentions of a conversation partner and presents an appropriate speech plan to the user. A specific implementation method of this system will be described below.
[0973] Hardware and software used
[0974] The system uses the following main hardware and software:
[0975] Camera: An optical device for collecting video data in real time, such as a camera mounted on smart glasses.
[0976] Microphone: Acoustic equipment for collecting audio data, such as a microphone in smart glasses.
[0977] Server: A high-performance computing device that performs data analysis and emotion recognition. It uses OpenCV for face recognition, Google Cloud Speech-to-Text API for voice recognition, and TensorFlow and Keras for the emotion engine.
[0978] Visual device: A device for presenting information to a user, for example, the display of smart glasses.
[0979] Data collection
[0980] First, a user puts on the smart glasses and starts a conversation with a customer. The camera captures the facial expressions and gestures of the person they are talking to, and the microphone captures audio data. This data is then sent to the server in real time.
[0981] Data analysis
[0982] The server analyzes the received video and audio data. Facial features are extracted from the video data using OpenCV, and emotions are estimated using TensorFlow and Keras. The audio data is converted to text using the Google Cloud Speech-to-Text API, and emotional changes are estimated through intonation analysis. This allows the intentions and emotions of the person being spoken to be identified.
[0983] Generate an utterance plan
[0984] The server generates a speech plan based on the analysis results. For example, if a customer is interested in a particular product but is concerned about the price, the server creates a speech plan such as, "This product offers excellent value for money and offers great benefits in the long run."
[0985] Display and Feedback
[0986] The generated speech plan is displayed on the smart glasses' display, and the user responds appropriately to the customer based on the displayed content. In this way, the conversation progresses smoothly and the customer's needs can be met.
[0987] Specific examples
[0988] For example, if you are explaining about a new smartphone in a physical store, the server may infer that the customer is interested in its features but is concerned about the price. The smart glasses will display a speech plan such as, "This smartphone uses the latest battery technology and will last all day with everyday use," and the user will explain based on that.
[0989] Example prompts to input to the generative AI model
[0990] "Generate utterance plans that recognize customer emotions and provide detailed information about products they're interested in. For example, create an utterance plan for when a customer is interested in the features of a new smartphone and has questions about the details, but is also concerned about the price."
[0991] This system will enable real-time customer service in physical stores by accurately understanding the emotions and intentions of the person you are talking to. Furthermore, by using visual devices to provide appropriate feedback to users and visualize the content of the conversation, it is expected that users' conversation skills will improve.
[0992] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0993] Step 1:
[0994] Data collection
[0995] The device uses a camera and microphone to collect facial expressions, gestures, and voice data of the conversation partner and the user. Specifically, the camera captures video data of the conversation partner and the user in real time, and the microphone collects voice data. This collected data is then sent from the device to a server.
[0996] Input: Camera video, microphone audio
[0997] Output: Video data, audio data
[0998] Step 2:
[0999] Video Data Analysis
[1000] The server analyzes the received video data and extracts facial features of the user and the conversation partner. It uses facial recognition technology to identify facial expressions and estimate emotions based on these. Specifically, it uses OpenCV for facial recognition and inputs the features into an emotion recognition model built with TensorFlow and Keras.
[1001] Input: Video data
[1002] Output: Facial features, emotion estimation results
[1003] Step 3:
[1004] Voice data analysis
[1005] The server analyzes the received voice data, converts it into text using speech recognition technology, and estimates emotional changes through intonation analysis. Specifically, the server converts the voice data into text using the Google Cloud Speech-to-Text API and then performs intonation analysis.
[1006] Input: Audio data
[1007] Output: Text data, emotion change estimation results
[1008] Step 4:
[1009] Integrated analysis of intentions and emotions
[1010] The server integrates the facial expression features, emotion estimation results, text data, and emotion change estimation results to comprehensively estimate the intentions and emotions of the conversation partner and the user.The server uses an emotion engine to comprehensively analyze the data and estimate the intentions.
[1011] Input: Facial features, emotion estimation results, text data, emotion change estimation results
[1012] Output: Intention estimation result, emotion estimation result
[1013] Step 5:
[1014] Generate an utterance plan
[1015] Based on the analysis results, the server generates a speech plan that matches the emotional state and intentions of the conversation partner and the user. Specifically, it uses a generative AI model to create an appropriate speech plan.
[1016] Input: Intention estimation result, emotion estimation result
[1017] Output: Utterance plan
[1018] Step 6:
[1019] Viewing the utterance plan
[1020] The user's device (smart glasses) displays the speech plan sent from the server on its display, allowing the user to proceed with the conversation according to the content displayed on the display. Specifically, text and icons are displayed on the display.
[1021] Input: Utterance plan
[1022] Output: Display
[1023] Step 7:
[1024] Feedback and conversation organization
[1025] The server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the terminal. The user's terminal displays this tree diagram, providing visual guidance on the next step in the conversation.
[1026] Input: Conversation progress data
[1027] Output: Tree diagram (conversation development diagram)
[1028] Step 8:
[1029] Self-Practice Mode
[1030] The device provides a self-practice mode, allowing users to practice speaking by themselves. The server provides speech samples, collects and analyzes the user's speech data in real time, and generates and sends feedback to the device based on the analysis results.
[1031] Input: User utterance data (self-practice mode)
[1032] Output: Feedback
[1033] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1034] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1035] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1036] [Third embodiment]
[1037] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1038] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1039] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1040] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1041] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1042] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1043] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1044] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1045] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1046] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1047] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1048] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1049] The present invention is a system that estimates the intentions and emotions of a conversation partner from their facial expressions, gestures, and voice, and presents an appropriate speech plan to the user, in order to facilitate interpersonal communication. Below, we will create a program for this system and explain its processing in detail.
[1050] System Configuration
[1051] The system consists of a terminal device (conversation prediction glasses) and a server device that performs data analysis. The terminal device is equipped with a camera and microphone, and the server device is equipped with an AI analysis model.
[1052] Program processing flow
[1053] Data collection
[1054] The device has a built-in camera to capture the facial expressions and gestures of the person you are talking to, and a microphone to collect audio data. When a conversation begins, the camera collects the video data of the person you are talking to, and the microphone collects the audio data in real time, and then sends the data to a server.
[1055] Data analysis
[1056] The server uses facial recognition technology to extract facial features of the conversation partner from the received video data and also analyzes gestures. At the same time, it converts the speech content into text using speech recognition technology from the audio data and estimates changes in emotion through intonation analysis. This data is input into an AI model, which then infers the other person's intentions and emotions.
[1057] Generate an utterance plan
[1058] The server generates a speech plan based on the analysis results, according to the intentions and emotions expressed by the conversation partner. For example, if a customer shows interest in a particular product, the server generates a speech plan to explain the details of that product and recommends it to the user. The generated speech plan is sent to the terminal.
[1059] User Feedback
[1060] The device displays the speech plan sent from the server on its display and provides the user with appropriate speech content. By continuing the conversation according to the displayed content, the user can communicate appropriately with the other person's intentions and emotions.
[1061] Organizing the conversation
[1062] Furthermore, the server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the device. The device displays this tree diagram, providing visual guidance on the next step in the conversation for the user.
[1063] Self-Practice Mode
[1064] The device has a self-practice mode, allowing users to practice speaking by themselves. In this mode, the server provides speaking samples, analyzes the user's speech data, and provides evaluation and feedback, allowing users to improve their speaking skills.
[1065] Specific examples
[1066] For example, a salesperson puts on these conversation-reading glasses when negotiating with a customer. After the negotiation begins, the glasses' camera captures the customer's facial expressions and gestures, and the microphone collects what the customer says, and these data are sent to a server. The server analyzes the data and infers that the customer is interested in price. The server then generates a speech plan, such as "This product has excellent cost performance and will offer great benefits in the long run," and sends it to the device. The device then displays this on its screen, and the salesperson continues to explain the plan to the customer. In this way, the conversation progresses smoothly, enabling responses to be made that meet the customer's needs.
[1067] This system makes it possible to accurately grasp the intentions and emotions of the person you are talking to, enabling smooth communication, which in turn makes business negotiations, counseling, and even everyday conversations more effective.
[1068] The processing flow will be explained below.
[1069] Step 1: Your device is ready for the camera and microphone
[1070] Before the conversation begins, the device turns on the camera and microphone to prepare for data collection.
[1071] Show the user a ready notification.
[1072] Step 2: The device collects video and audio data from the person you are talking to.
[1073] The camera continuously captures the facial expressions and gestures of the person you are talking to.
[1074] The microphone records the voice of the person you are talking to in real time.
[1075] The collected data is sent to the server in real time.
[1076] Step 3: The server analyzes the video data
[1077] The server analyzes the received video data and extracts facial expression features of the conversation partner.
[1078] Use facial recognition technology to estimate emotions from facial expressions.
[1079] Gesture analysis is performed to infer intentions from hand and body movements.
[1080] Step 4: The server analyzes the audio data
[1081] The server analyzes the received voice data and converts the spoken content into text using voice recognition technology.
[1082] The text of the speech is analyzed, and emotions are inferred from word choice and intonation.
[1083] Step 5: The server estimates the other person's intentions and emotions
[1084] The features extracted from the video and audio data are input into the AI model.
[1085] The AI model estimates the overall intention and emotions of the conversation partner.
[1086] Step 6: The server generates an appropriate utterance plan
[1087] Based on the estimation results, the AI generates the optimal speech plan for the user.
[1088] A speech plan includes specific phrases and next steps in the conversation.
[1089] Step 7: The server sends the speech plan to the device.
[1090] The server transmits the generated speech plan to the terminal.
[1091] Optimize and transmit data to ensure stable communication and prevent delays.
[1092] Step 8: The device displays the speech plan to the user.
[1093] The terminal displays the speech plan received from the server on the display.
[1094] The user continues the conversation by following what is displayed.
[1095] Step 9: The server organizes the conversation
[1096] As the conversation progresses, the server organizes the user's comments and the analysis results.
[1097] Structure the conversation content as a tree diagram or flowchart.
[1098] Step 10: The server sends the organized conversation to the device.
[1099] The organized conversation content and the next steps to take are sent to the terminal.
[1100] Help users make good decisions.
[1101] Step 11: The terminal displays the tree diagram to the user
[1102] The terminal displays the tree diagram or flowchart received from the server.
[1103] Check what the user should say next and the flow.
[1104] Step 12: User selects and starts self-practice mode
[1105] The user selects the self-practice mode through the terminal interface.
[1106] An interface provides the user with practice mode settings.
[1107] Step 13: The server provides and analyzes speech samples
[1108] The server provides speech samples for use in self-practice mode.
[1109] Data spoken by users is collected in real time and analyzed.
[1110] Step 14: The server generates feedback and sends it to the device
[1111] The server evaluates the user's speech data and generates suggestions for improvement and appropriate feedback.
[1112] The generated feedback is sent to the device.
[1113] Step 15: The device displays feedback to the user
[1114] The terminal displays the feedback received from the server to the user.
[1115] The user improves their speaking skills based on the feedback.
[1116] Through these processing steps, users can properly understand the intentions and emotions of their conversation partners and communicate smoothly. In addition, by utilizing the self-practice mode, users can effectively improve their speaking skills.
[1117] Example 1
[1118] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1119] Interpersonal communication requires rapid and accurate understanding of the intentions and emotions of the person you are speaking with, but the current technology for doing so is insufficient. It is particularly difficult to respond appropriately to changes in the other person's intentions and emotions in situations such as business negotiations and counseling. There is also a lack of effective feedback systems for improving speaking skills through self-practice. A system that solves these issues is needed.
[1120] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1121] In this invention, the server includes means for collecting facial expressions and gestures of a conversation partner using a camera, means for collecting voice data of the conversation partner using a microphone, means for transmitting the collected video and voice data to the server in real time, means for extracting facial features from the received video data using face recognition technology and performing gesture analysis, means for converting the speech content into text from the voice data using voice recognition technology and estimating emotional changes through intonation analysis, and means for generating and displaying an appropriate speech plan for the user based on the estimated intentions and emotions. This makes it possible to accurately grasp the intentions and emotions of the conversation partner and achieve smooth communication.
[1122] A "camera" is a device for collecting video data, and in this invention is used to capture the facial expressions and gestures of a conversation partner.
[1123] A "microphone" is a device for collecting voice data, and in the present invention is used to clearly capture the voice of a conversation partner.
[1124] A "server" is a computer system for analyzing and processing data, and in the present invention, it has the role of analyzing video data and audio data and providing speech plans to users.
[1125] "Facial expression features" are a set of feature values that indicate the emotions and intentions of a conversation partner, extracted using face recognition technology.
[1126] "Gesture analysis" is the process of analyzing a conversation partner's hand movements and gestures to infer non-verbal intentions and emotions.
[1127] "Speech recognition technology" is a technology that analyzes voice data and converts it into text data, and is used in the present invention to analyze the content of speech from a conversation partner.
[1128] "Intonation analysis" is a technology that analyzes the intonation and rhythm of voice data to estimate changes in emotion.
[1129] An "utterance plan" is a proposal of utterance content to support the progress of a conversation, generated based on analyzed data.
[1130] A "tree diagram" is a diagram that organizes the development of a conversation in a hierarchical structure to visually show the progress of the conversation.
[1131] The "self-practice mode" is a function that allows the user to practice speaking on their own, and includes providing speaking samples and feedback.
[1132] The present invention is a system that estimates the intentions and emotions of a conversation partner from their facial expressions, gestures, and voice, and presents an appropriate speech plan to the user, in order to facilitate interpersonal communication. Below, we will create a program for this system and explain its processing in detail.
[1133] System Configuration
[1134] The main components of this system are a terminal device (conversation prediction glasses) and a server device that performs data analysis. The terminal device is equipped with a camera and microphone, and the server device is equipped with an AI analysis model.
[1135] Program processing flow and specific functions
[1136] Data collection
[1137] The device has a built-in camera for capturing facial expressions and gestures of the conversation partner, and a microphone for capturing audio data. When a conversation begins, the camera captures the conversation partner's video data and the microphone captures audio data in real time, and these data are then sent to a server. The hardware used is a built-in camera (e.g., an HD camera) and a highly sensitive microphone.
[1138] Data analysis
[1139] The server uses facial recognition technology to extract facial features of the conversation partner from the received video data and also performs gesture analysis. At the same time, it uses speech recognition technology to convert the speech into text from the audio data and estimates emotional changes through intonation analysis. Software libraries such as TensorFlow and OpenCV are used for these analyses. For example, Google's Speech-to-Text API is sometimes used to analyze audio data.
[1140] Generate an utterance plan
[1141] The server generates a speech plan based on the analysis results, taking into account the intentions and emotions of the conversation partner. This is done using a pre-trained generative AI model. For example, if a customer shows interest in a particular product, a speech plan to explain the details of that product is generated. The generated speech plan is sent to the device.
[1142] User Feedback
[1143] The device displays the speech plan sent from the server on its display and provides the user with appropriate speech content. By continuing the conversation according to the displayed content, the user can communicate appropriately with the other person's intentions and emotions.
[1144] Organizing the conversation
[1145] Furthermore, the server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the device. The device displays this tree diagram, providing visual guidance on the next step in the conversation for the user.
[1146] Self-Practice Mode
[1147] The device has a self-practice mode, allowing users to practice speaking by themselves. In this mode, the server provides speaking samples, analyzes the user's speech data, and provides evaluation and feedback, allowing users to improve their speaking skills.
[1148] Specific examples
[1149] For example, a salesperson puts on these conversation-reading glasses when negotiating with a customer. After the negotiation begins, the glasses' camera captures the customer's facial expressions and gestures, and the microphone collects what the customer says, and these data are sent to a server. The server analyzes the data and infers that the customer is interested in price. The server then generates a speech plan, such as "This product has excellent cost performance and will offer great benefits in the long run," and sends it to the device. The device then displays this on its screen, and the salesperson continues to explain the plan to the customer. In this way, the conversation progresses smoothly, enabling responses to be made that meet the customer's needs.
[1150] Example prompts for generative AI models
[1151] Input data: Customer facial expression information, voice data
[1152] AI model: Estimating conversation partner's intentions and emotions
[1153] Output: It is assumed that the customer is interested in the price. Generate the utterance plan "This product offers excellent value for money and offers great benefits in the long run."
[1154] As described above, the present invention is a system for understanding the intentions and emotions of a conversation partner and supporting appropriate communication, and is particularly effective in situations such as business negotiations and counseling.
[1155] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1156] Step 1:
[1157] Data collection
[1158] The device uses a camera and microphone to capture facial expressions, gestures, and voice data from the person you're talking to. The camera captures high-resolution video, while the microphone suppresses background noise to capture clear audio.
[1159] input:
[1160] Video and audio data of the person you are talking to
[1161] output:
[1162] Collected video and audio data
[1163] Specific behavior:
[1164] When the user begins interacting with a customer, the device's camera captures the customer's facial expressions and gestures, and the microphone begins recording what the customer says.
[1165] Step 2:
[1166] Data transmission
[1167] The device transmits the collected video and audio data to the server in real time over a fast and secure network.
[1168] input:
[1169] Collected video and audio data
[1170] output:
[1171] Data sent to the server
[1172] Specific behavior:
[1173] The device encrypts the collected data and sends it to a server using Wi-Fi or 5G networks.
[1174] Step 3:
[1175] Data analysis
[1176] The server extracts facial features from the received video data using facial recognition technology and also performs gesture analysis. It also converts the speech into text using speech recognition technology and estimates emotional changes through intonation analysis. TensorFlow and OpenCV are used as software libraries.
[1177] input:
[1178] Video and audio data sent to the server
[1179] output:
[1180] Extracted facial features, textualized speech, and estimated emotional changes
[1181] Specific behavior:
[1182] The server uses TensorFlow to analyze the video data and extract facial expression and gesture features. At the same time, Google's Speech-to-Text API converts the audio data into text, and the emotion analysis API analyzes emotions from intonation.
[1183] Step 4:
[1184] Generate an utterance plan
[1185] The server generates a speech plan based on the analysis results, taking into account the intentions and emotions of the conversation partner. Using a generative AI model, it constructs an appropriate speech plan that corresponds to the estimated intentions and emotions.
[1186] input:
[1187] Extracted facial features, textualized speech, and estimated emotional changes
[1188] output:
[1189] Generated utterance plan
[1190] Specific behavior:
[1191] If the server's AI model determines that the customer is interested in price, it generates a speech plan such as, "This product has excellent cost performance."
[1192] Step 5:
[1193] Sending a speech plan
[1194] The server then sends the generated speech plan to the device, again in real time, allowing the user to receive immediate feedback.
[1195] input:
[1196] Generated utterance plan
[1197] output:
[1198] Speech plan sent to the device
[1199] Specific behavior:
[1200] The speech plan generated by the server is immediately sent to the device, and a notification is displayed on the device's display.
[1201] Step 6:
[1202] User Feedback
[1203] The device displays the received speech plan on its display and provides the user with appropriate speech content, which the user can refer to as they proceed with the conversation.
[1204] input:
[1205] Speech plan sent to the device
[1206] output:
[1207] The speech plan presented to the user
[1208] Specific behavior:
[1209] The terminal display will show "This product has excellent cost performance," and the user will convey this statement to the customer.
[1210] Step 7:
[1211] Organizing the conversation
[1212] The server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the terminal.
[1213] input:
[1214] Analyzed conversation content
[1215] output:
[1216] Speech content organized as a tree diagram
[1217] Specific behavior:
[1218] The server analyzes the flow of conversation, and if there is a high level of "interest in price," it generates a node for "additional information related to price" and displays it in the form of a tree diagram on the terminal display.
[1219] Step 8:
[1220] Self-Practice Mode
[1221] The device has a self-practice mode, allowing users to practice speaking by themselves. In this mode, the server provides speaking samples and analyzes the user's speech data to provide evaluation and feedback.
[1222] input:
[1223] Speech samples entered by the user
[1224] output:
[1225] Ratings and Feedback
[1226] Specific behavior:
[1227] When a user selects practice mode, the server sends sample utterances, such as "greeting practice," to the device. As the user speaks, the audio is analyzed and pronunciation and intonation are evaluated in real time.
[1228] (Application example 1)
[1229] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1230] Existing conversation support systems have difficulty understanding the intentions and emotions of the target, and are therefore unable to present appropriate speech plans or response methods for achieving effective communication. Furthermore, in security services, where it is necessary to quickly understand the intentions and emotions of visitors and passersby and respond appropriately, the lack of such technology poses a serious problem.
[1231] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1232] In this invention, the server includes means for collecting facial expressions and gestures of a target using a camera, means for collecting audio data of the target using a microphone, means for analyzing the collected video and audio data to estimate the intention and emotion of the target, means for displaying an appropriate speech plan to the user based on the estimated intention and emotion, and means for suggesting a response method to the user based on the estimated intention and emotion, thereby enabling security guards to quickly grasp the intention and emotion of visitors and passersby and to determine an appropriate response method.
[1233] A "camera" is a photographic device for capturing facial expressions and gestures of a subject.
[1234] A "microphone" is a recording device for collecting audio data of a subject.
[1235] "Video data" refers to image and video information collected using a camera.
[1236] "Audio data" is sound information collected using a microphone.
[1237] "Facial expression" refers to emotions and intentions that can be read from the movement of facial muscles.
[1238] A "gesture" is an expression of intent that can be interpreted through hand or body movements.
[1239] "Intention" refers to the purpose or direction of an action that a subject has.
[1240] "Emotion" refers to a subject's state of mind or mood.
[1241] A "speech plan" is a plan for providing appropriate conversation content.
[1242] A "display" is a display device for presenting visual information to a user.
[1243] A "server" is a computer system that analyzes collected data and provides results.
[1244] A "tree diagram" is a diagram that visually organizes the content and development of a conversation.
[1245] The "self-practice mode" is a mode in which the user can practice speaking by himself.
[1246] "Feedback" refers to evaluation of the user's speech and advice for improvement.
[1247] This invention is an interpersonal communication support system specialized for security applications, and employs an approach to infer the intentions and emotions of a conversation partner from their facial expressions, gestures, and voice, and provide appropriate response methods.
[1248] System Configuration
[1249] The system includes smart glasses (terminal devices) and a server device that performs data analysis. The smart glasses are equipped with a camera and microphone, and the server device is equipped with an AI analysis model.
[1250] Data collection
[1251] The terminal device is equipped with a camera to capture the facial expressions and gestures of the visitor and a microphone to collect audio data. When a security guard comes into contact with a visitor, the camera collects the subject's video data and the microphone collects audio data in real time, and these data are then sent to a server.
[1252] Data analysis
[1253] The server uses facial recognition technology to extract facial features from the received video data and also analyzes gestures. At the same time, it converts the speech into text using speech recognition technology and estimates emotional changes through intonation analysis. This data is input into an AI model, which then infers the subject's intentions and emotions.
[1254] Generate speech plans and responses
[1255] Based on the analysis results, the server generates a speech plan and appropriate response method according to the visitor's intentions and emotions. For example, if the visitor seems nervous, the server generates advice recommending polite responses and sends it to the terminal device. The generated speech plan and response method are then displayed on the smart glasses' display.
[1256] User Feedback
[1257] The terminal device displays the speech plan and response method sent from the server on its display and provides the security guard with appropriate response content. By responding to the visitor according to the displayed content, the security guard can appropriately deal with the visitor's intentions and emotions.
[1258] Organizing the conversation
[1259] Furthermore, the server analyzes the progress of the conversation in real time, organizes the content and development of the conversation into a tree diagram, and sends it to the terminal device, which displays this tree diagram and provides visual guidance to the security guard on the next steps in the conversation.
[1260] Self-Practice Mode
[1261] The terminal device is equipped with a self-practice mode, allowing security guards to practice speaking by themselves. In this mode, the server provides speaking samples, analyzes the security guards' speech data, and provides evaluation and feedback, allowing the security guards to improve their speaking skills.
[1262] Examples and prompts
[1263] Specific examples
[1264] A security guard wears smart glasses at the entrance to a facility and handles visitors. If the visitor's facial expression appears stiff and tense, the camera and microphone collect information and the information is analyzed by a server. As a result of the analysis, the smart glasses' display displays advice such as, "It is assumed that the visitor is nervous. Please handle the visitor politely and confirm the purpose of their visit."
[1265] Prompt Sentence Examples
[1266] "Advice to display when a visitor's facial expression appears tense:
[1267] We suspect the visitor may be nervous. Please be polite and confirm the purpose of their visit."
[1268] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1269] Step 1:
[1270] Data collection
[1271] The device (smart glasses) uses a camera to collect video data of the subject's facial expressions and gestures, and simultaneously collects audio data using a microphone. The input is video data from the camera and audio data from the microphone, and the output is the collected raw video and audio data.
[1272] Step 2:
[1273] Data transmission
[1274] The terminal transmits the collected video and audio data to the server in real time. The input is the collected video and audio data, and the output is the data transmitted to the server.
[1275] Step 3:
[1276] Facial recognition and facial expression feature extraction
[1277] The server uses facial recognition technology to detect the target face from the received video data and extract facial expression features. The input is the video data, and the output is the extracted facial expression features. The software used is a facial recognition library such as OpenCV.
[1278] Step 4:
[1279] Gesture Analysis
[1280] The server analyzes hand and body movements from video data to detect gestures. The input is video data, and the output is detected gesture data. The software used is an AI analysis model.
[1281] Step 5:
[1282] Converting audio data to text
[1283] The server converts the voice data into text data using speech recognition technology. The input is voice data and the output is text data. The software used is a speech recognition library (e.g., Google Cloud Speech-to-Text).
[1284] Step 6:
[1285] Intonation analysis
[1286] The server analyzes the intonation of the voice data and estimates emotional changes. The input is the voice data, and the output is estimated emotional data. The software used is an AI analysis model.
[1287] Step 7:
[1288] Intention and emotion estimation
[1289] The server inputs the extracted facial features, gesture data, text data, and emotion data into an AI model to comprehensively estimate the subject's intention and emotion. The input is each feature data, and the output is estimated intention and emotion information.
[1290] Step 8:
[1291] Generate speech plans and responses
[1292] The server generates a speech plan and a response method for the user based on the estimated intention and emotion. For example, if the visitor is nervous, it generates advice including an appropriate response method. The input is the intention and emotion information, and the output is the speech plan and a response method.
[1293] Step 9:
[1294] User Feedback
[1295] The terminal displays the speech plan and response method sent from the server on a display and provides them to the security guard. The input is the speech plan and response method, and the output is the content displayed on the display. The security guard responds to the visitor according to the displayed content.
[1296] Step 10:
[1297] Organizing the conversation
[1298] The server analyzes ongoing conversations in real time and organizes them into a tree diagram. The input is real-time conversation data, and the output is a tree diagram of the organized conversations. The terminal displays this tree diagram to help security guards understand the next steps.
[1299] Step 11:
[1300] Self-Practice Mode
[1301] The terminal provides an interface that allows the user to select self-practice mode. The server provides speech samples, analyzes the user's speech data, and provides evaluation and feedback. The input is the user's speech data, and the output is evaluation and feedback, allowing the security guard to improve their speaking skills.
[1302] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1303] This invention is a system that combines conversation prediction glasses with an emotion engine to more accurately recognize the emotions and intentions of both the user and the conversation partner and present appropriate speech plans. Below, we will create a program for this system and explain its processing in detail.
[1304] System Configuration
[1305] The system includes a terminal device (conversation prediction glasses), a server device that performs data analysis, and an emotion engine with emotion recognition capabilities. The terminal device is equipped with a camera and microphone, and the server device is equipped with an AI analysis model and emotion engine.
[1306] Program processing flow
[1307] Data collection
[1308] The device uses a camera and microphone to collect facial expressions, gestures, and voice data of the user and the conversation partner. When a conversation begins, the camera collects video data of the conversation partner and the user, and the microphone collects voice data in real time, and then transmits this data to a server.
[1309] Data analysis
[1310] The server analyzes the received video and audio data and extracts facial and audio features of the conversation partner and the user. First, facial recognition technology is used to infer emotions from facial expressions, followed by gesture analysis. Furthermore, speech recognition technology is used to convert the speech into text from the audio data, and emotional changes are inferred through intonation analysis. This data is input into an AI model and emotion engine, which infers the intentions and emotions of the conversation partner and the user.
[1311] Generate an utterance plan
[1312] The server generates a speech plan based on the analysis results, taking into account the emotional state and intentions of the conversation partner and the user. For example, the server may generate a speech plan that provides detailed explanations based on a topic that the conversation partner is interested in, or suggest relaxing topics to reduce the stress the user is feeling. The generated speech plan is then sent to the device.
[1313] User Feedback
[1314] The device displays the speech plan sent from the server on its display and provides the user with appropriate speech content. By continuing the conversation according to the displayed content, the user can communicate appropriately with the other person's intentions and emotions.
[1315] Organizing the conversation
[1316] Furthermore, the server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the device. The device displays this tree diagram, providing visual guidance on the next step in the conversation for the user.
[1317] Self-Practice Mode
[1318] The device offers a self-practice mode, allowing users to practice speaking by themselves. In this mode, the server provides speech samples, collects and analyzes the user's speech data in real time, and generates and sends feedback to the device based on the analysis results.
[1319] Specific examples
[1320] For example, a sales representative puts on these conversation-predicting glasses when negotiating with a customer. After the negotiation begins, the glasses' camera captures the facial expressions and gestures of both the customer and the sales representative, while the microphone collects speech, and these data are sent to a server. The server analyzes the data and infers that the customer is beginning to show interest in the product but is concerned about the price. The emotion engine then confirms that the sales representative is explaining with confidence. The server generates a speech plan, such as "This product has excellent cost performance and will offer great benefits in the long run," and sends it to the device. The device displays this on its screen, and the sales representative explains based on it. In this way, the conversation progresses smoothly, enabling responses that meet the customer's needs.
[1321] This system allows both the user and the conversation partner to accurately understand each other's intentions and emotions, enabling smooth and effective communication, which will enable more effective business negotiations, counseling, and everyday conversations.
[1322] The processing flow will be explained below.
[1323] Step 1: Your device is ready for the camera and microphone
[1324] Before the conversation begins, the device turns on the camera and microphone to prepare for data collection.
[1325] Show the user a ready notification.
[1326] Step 2: The device collects video and audio data of the conversation partner and the user.
[1327] The camera continuously captures the facial expressions and gestures of the conversation partner and the user.
[1328] The microphone records the voices of the conversation partner and the user in real time.
[1329] The collected data is sent to the server in real time.
[1330] Step 3: The server analyzes the video data of the person you are talking to.
[1331] The server analyzes the received video data and extracts facial features of the conversation partner.
[1332] Use facial recognition technology to estimate emotions from facial expressions.
[1333] Gesture analysis is performed to infer intentions from hand and body movements.
[1334] Step 4: The server analyzes the voice data of the other party.
[1335] The server analyzes the received voice data and converts the spoken content into text using voice recognition technology.
[1336] The text of the speech is analyzed, and emotions are inferred from word choice and intonation.
[1337] Step 5: The server analyzes the user's video data.
[1338] The server analyzes the received video data and extracts the user's facial expression features.
[1339] Use facial recognition technology to estimate emotions from facial expressions.
[1340] Gesture analysis is performed to estimate the current state and intentions from hand and body movements.
[1341] Step 6: The server analyzes the user's voice data
[1342] The server analyzes the received voice data and converts the spoken content into text using voice recognition technology.
[1343] The system analyzes the text of the speech and estimates the user's emotions from word choice and intonation.
[1344] Step 7: The server integrates and estimates the intentions and emotions of the conversation partner and the user.
[1345] The features extracted from the video and audio data are input into the AI model and emotion engine.
[1346] AI models and emotion engines estimate the overall intent and emotions of the conversation partner and the user.
[1347] Step 8: The server generates an appropriate utterance plan
[1348] Based on the estimation results, the AI generates the optimal speech plan for the user.
[1349] A speech plan includes specific phrases and next steps in the conversation.
[1350] Step 9: The server sends the speech plan to the device.
[1351] The server transmits the generated speech plan to the terminal.
[1352] Optimize and transmit data to ensure stable communication and prevent delays.
[1353] Step 10: The device displays the speech plan to the user.
[1354] The terminal displays the speech plan received from the server on the display.
[1355] The user continues the conversation by following what is displayed.
[1356] Step 11: The server organizes the conversation
[1357] As the conversation progresses, the server organizes the user's comments and the analysis results.
[1358] Structure the conversation content as a tree diagram or flowchart.
[1359] Step 12: The server sends the organized conversation to the device.
[1360] The organized conversation content and the next steps to take are sent to the terminal.
[1361] Help users make good decisions.
[1362] Step 13: The terminal displays the tree diagram to the user
[1363] The terminal displays the tree diagram or flowchart received from the server.
[1364] Check what the user should say next and the flow.
[1365] Step 14: User selects and starts self-practice mode
[1366] The user selects the self-practice mode through the terminal interface.
[1367] An interface provides the user with practice mode settings.
[1368] Step 15: The server provides and analyzes speech samples.
[1369] The server provides speech samples for use in self-practice mode.
[1370] Data spoken by users is collected in real time and analyzed.
[1371] Step 16: The server generates feedback and sends it to the device
[1372] The server evaluates the user's speech data and generates suggestions for improvement and appropriate feedback.
[1373] The generated feedback is sent to the device.
[1374] Step 17: The device displays feedback to the user
[1375] The terminal displays the feedback received from the server to the user.
[1376] The user improves their speaking skills based on the feedback.
[1377] Through these processing steps, users can accurately understand their own and their conversation partner's emotions and intentions, enabling smoother and more effective communication. Furthermore, by utilizing the self-practice mode, users can effectively improve their speaking skills.
[1378] Example 2
[1379] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1380] In conversation, users often have difficulty accurately reading the intentions and emotions of the other person, making smooth communication difficult. They also often lack appropriate responses to the anxiety and stress they feel. Furthermore, even when practicing by themselves, the lack of concrete feedback makes it difficult to improve their speech.
[1381] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1382] In this invention, the server includes means for collecting facial expressions and gestures of the conversation partner and the user using a camera, means for collecting voice data of the conversation partner and the user using a microphone, means for analyzing the collected video and voice data to estimate the intentions and emotions of the conversation partner and the user, means for displaying an appropriate speech plan to the user based on the estimated intentions and emotions, means for organizing the progress of the conversation in a tree diagram based on the analysis results and displaying it to the user, and means for providing an interface that allows the user to select a self-practice mode and receive feedback for improving their speech. This allows the user to accurately understand the intentions and emotions of the conversation partner and take appropriate measures. Furthermore, the self-practice mode makes it easier to improve speech, thereby improving the overall quality of communication.
[1383] A "camera" is an optical device for collecting video data.
[1384] A "conversational partner" refers to the other person with whom the user is interacting.
[1385] "User" refers to a person who uses this system to have a conversation.
[1386] "Facial expressions" refer to facial movements and expressions, and are important non-verbal elements for indicating emotions and intentions.
[1387] "Gestures" refer to hand and body movements and are non-verbal expressions used to convey intentions and emotions.
[1388] A "microphone" is an acoustic device for collecting audio data.
[1389] "Audio data" refers to collected sound information, including speech content and intonation.
[1390] "Collect" refers to obtaining and storing data.
[1391] "Analysis" is the act of extracting and analyzing information based on collected data.
[1392] "Intention" refers to the purpose or thoughts of the conversation partner or user.
[1393] "Emotion" refers to the emotional state felt by a conversation partner or user.
[1394] A "speech plan" is a specific speech content suggested by the system based on a specific intention or emotion.
[1395] "Display" is the act of enabling a user to visually confirm information.
[1396] A "tree diagram" is a diagram that visually shows the progress of a conversation and the relationship between topics.
[1397] "Interface" refers to the operating screen or means by which a user interacts with a system.
[1398] "Feedback" refers to information that provides suggestions for improvement or evaluation of a user's actions or results.
[1399] The present invention is a system that includes a terminal device (conversation prediction glasses) worn by a user, a server device that performs data analysis, and an emotion engine with emotion recognition functionality. Specific embodiments for implementing this system are described below.
[1400] terminal device
[1401] The terminal device is equipped with a camera and a microphone. These hardware components are used to collect video and audio data of the user and the conversation partner in real time. Specifically, the camera captures facial expressions and gestures, and the microphone collects audio data of the conversation. The terminal device has the function of transmitting the collected data to a server device via a network.
[1402] Server device
[1403] The server device is equipped with an AI model and emotion engine for analyzing the received data. From the video data, facial recognition technology is used to extract facial expression features and perform gesture analysis. For the audio data, speech is converted into text using speech recognition technology, and emotional changes are estimated through intonation analysis. This data is input into the AI model and emotion engine, which ultimately estimates the intentions and emotions of the conversation partner and the user.
[1404] Providing Feedback
[1405] The server generates a speech plan based on the analysis results and transmits it to the terminal device. The terminal device displays this speech plan on a display and provides it to the user. The user can continue the conversation based on the displayed speech plan, and can respond appropriately to the intentions and feelings of the conversation partner.
[1406] Organizing the conversation
[1407] The server analyzes the progress of the conversation in real time and organizes the content and development of the conversation into a tree diagram. This tree diagram is sent to the terminal device, helping the user visually grasp the next flow of the conversation.
[1408] Self-Practice Mode
[1409] The terminal device is equipped with an interface that allows the user to select a self-practice mode, in which the server collects the user's speech, analyzes it in real time, and generates feedback, which helps the user improve their speech.
[1410] Specific examples
[1411] For example, a sales representative puts on conversation-predicting glasses when negotiating with a customer. When the negotiation begins, the device's camera captures the facial expressions and gestures of both the customer and the sales representative, and the microphone collects the audio of the conversation. This data is sent to a server, which uses analysis technology to infer the customer's emotions and intentions. Based on the analysis results, a speech plan such as "This product has excellent cost performance and will provide great benefits in the long run" is generated and displayed on the device. The user can explain the situation based on the displayed speech plan, allowing the negotiation to proceed smoothly.
[1412] Prompt Sentence Examples
[1413] "Using conversation-reading glasses, please generate a specific scenario in which a salesperson is negotiating with a customer. Include a process for reading the customer's emotions and interests and proposing an appropriate speech plan."
[1414] This system accurately recognizes the intentions and emotions of both the user and the conversation partner, enabling accurate and smooth communication, promoting effective dialogue in business negotiations, counseling, and everyday conversations.
[1415] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1416] Step 1:
[1417] The terminal collects video and audio of the conversation partner and the user in real time.
[1418] Specific operation: The device's camera captures the face, facial expressions, and gestures of the user and the person in conversation, and the microphone collects audio. The input is the video and audio data acquired through the camera and microphone, and the output is the collected data packets.
[1419] Step 2:
[1420] The terminal transmits the collected video and audio data to the server.
[1421] Specific operation: Video data is packetized as image frames, and audio data is packetized as audio clips, and then sent to the server over the network. The input is the collected data packets, and the output is the server to which the data was sent.
[1422] Step 3:
[1423] The server analyzes the received video data and extracts facial expression features.
[1424] Specific operation: Using facial recognition technology in the server, the system identifies the facial parts of the conversation partner and the user and calculates facial features such as smile, sadness, surprise, etc. The input is the received video data, and the output is the extracted facial features.
[1425] Step 4:
[1426] The server performs gesture analysis and infers emotions and intentions from the collected gesture data.
[1427] Specific operation: Using the server's gesture recognition algorithm, hand movements and body poses are analyzed to estimate emotional states such as interest or tension. The input is the received video data and existing gesture data, and the output is the estimated emotion or intention.
[1428] Step 5:
[1429] The server analyzes the voice data and converts the spoken content into text.
[1430] How it works: The speech recognition system analyzes recorded speech and converts it into text data. It then detects emotional changes through intonation analysis. The input is the received speech data, and the output is the text of the speech and an indicator of emotional changes.
[1431] Step 6:
[1432] The server uses AI models and emotion engines to infer the intentions and emotions of the user and their conversation partner.
[1433] How it works: The extracted features are input into an AI model based on previous research and training data to estimate emotions and intentions. An emotion engine is used to further improve accuracy. The inputs are facial features, gesture data, text data of spoken content, and indicators of emotional changes, and the output is estimated intentions and emotions.
[1434] Step 7:
[1435] Based on the analysis results, the server generates a speech plan that is appropriate for the emotions and intentions of the user and their conversation partner.
[1436] How it works: Using a generative AI model, it automatically determines the appropriate utterance content for the situation (e.g., question, explanation, relaxed topic, etc.) and constructs it as an utterance plan. The input is the estimated intent and emotion, and the output is the generated utterance plan.
[1437] Step 8:
[1438] The server transmits the generated speech plan to the terminal.
[1439] Specific operation: The text information generated as a speech plan is packetized and sent to the terminal via the network. The input is the generated speech plan, and the output is the terminal to which it was sent.
[1440] Step 9:
[1441] The terminal displays the received speech plan on the display.
[1442] Specific operation: The display module in the terminal displays the text of the speech plan on the screen, providing a visual for the user. The input is the received speech plan, and the output is the speech plan displayed on the display.
[1443] Step 10:
[1444] The user continues the conversation according to the content displayed on the display.
[1445] Specific operation: The user refers to the displayed speech plan and speaks at the appropriate time to progress the conversation. The input is the speech plan displayed on the display, and the output is the ongoing conversation.
[1446] Step 11:
[1447] The server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the terminal.
[1448] Specific operation: The collected data is analyzed one after another, and data for visualizing the conversation topics and progress as a tree diagram is generated and sent to the device. The input is the ongoing conversation data and analysis results, and the output is data organized as a tree diagram.
[1449] Step 12:
[1450] The terminal displays a tree diagram showing the progress of the conversation on its display.
[1451] Specific operation: The display module in the terminal displays the tree diagram on the screen, providing the user with a visual of the next conversation flow. The input is the tree diagram data, and the output is the tree diagram displayed on the display.
[1452] Step 13:
[1453] The terminal provides a self-practice mode for the user to practice speaking.
[1454] Specific behavior: The device displays a practice interface and presents speech samples to the user. The input is the user's selection, and the output is the self-practice mode interface.
[1455] Step 14:
[1456] The server collects and analyzes user speech data in real time.
[1457] Specific operations: Receives collected speech data from the device, performs voice analysis and intonation analysis, generates feedback based on the analysis results, and sends it to the device. The input is the collected speech data, and the output is the generated feedback.
[1458] (Application example 2)
[1459] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1460] Conventional dialogue support systems have difficulty accurately grasping the emotions and intentions of their conversation partners and providing appropriate speech plans in real time. Furthermore, in brick-and-mortar stores, staff are required to read customers' emotions and respond effectively, but current technology makes this difficult. Furthermore, feedback and visualization of conversation content using visual devices are insufficient, limiting the ability to improve users' dialogue skills.
[1461] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1462] In this invention, the server includes means for collecting facial expressions and gestures of a conversation partner using a camera, means for collecting voice data of the conversation partner using a microphone, means for analyzing the collected video and voice data to estimate the intentions and emotions of the conversation partner, means for displaying an appropriate speech plan to the user based on the estimated intentions and emotions, and means for generating an utterance plan based on the analysis results and displaying it on the user's visual device. This allows for accurate understanding of the conversation partner's emotions and intentions, enabling real-time customer service in physical stores. Furthermore, by using the visual device to provide appropriate feedback to the user and visualize the conversation content, it is expected that the user's dialogue skills will improve.
[1463] A "camera" is an optical device for collecting video data.
[1464] A "conversational partner" refers to a person other than the user with whom the conversation is taking place.
[1465] "Facial expression" refers to information including emotions and intentions expressed through the movement of facial muscles.
[1466] A "gesture" is a communication action using body and hand movements.
[1467] A "microphone" is an acoustic device for collecting audio data.
[1468] "Audio data" is digital information that digitizes audio.
[1469] "Analysis" refers to the process of extracting patterns and information from collected data.
[1470] "Intention" refers to the purpose or thoughts of the person you are speaking with.
[1471] "Emotions" are information that indicates the feelings and psychological state of the person you are talking to.
[1472] "Inference" is the process of predicting intentions and emotions from collected data.
[1473] "Utterance plan" refers to the content and plan of what the user will say.
[1474] "User" refers to a person using the smart glasses or system.
[1475] "Display" is the act of presenting information to a visual device.
[1476] A "visual device" is a device that allows a user to receive information visually.
[1477] This invention relates to a system that uses a camera, a microphone, a server, and a visual device to analyze the emotions and intentions of a conversation partner and presents an appropriate speech plan to the user. A specific implementation method of this system will be described below.
[1478] Hardware and software used
[1479] The system uses the following main hardware and software:
[1480] Camera: An optical device for collecting video data in real time, such as a camera mounted on smart glasses.
[1481] Microphone: Acoustic equipment for collecting audio data, such as a microphone in smart glasses.
[1482] Server: A high-performance computing device that performs data analysis and emotion recognition. It uses OpenCV for face recognition, Google Cloud Speech-to-Text API for voice recognition, and TensorFlow and Keras for the emotion engine.
[1483] Visual device: A device for presenting information to a user, for example, the display of smart glasses.
[1484] Data collection
[1485] First, a user puts on the smart glasses and starts a conversation with a customer. The camera captures the facial expressions and gestures of the person they are talking to, and the microphone captures audio data. This data is then sent to the server in real time.
[1486] Data analysis
[1487] The server analyzes the received video and audio data. Facial features are extracted from the video data using OpenCV, and emotions are estimated using TensorFlow and Keras. The audio data is converted to text using the Google Cloud Speech-to-Text API, and emotional changes are estimated through intonation analysis. This allows the intentions and emotions of the person being spoken to be identified.
[1488] Generate an utterance plan
[1489] The server generates a speech plan based on the analysis results. For example, if a customer is interested in a particular product but is concerned about the price, the server creates a speech plan such as, "This product offers excellent value for money and offers great benefits in the long run."
[1490] Display and Feedback
[1491] The generated speech plan is displayed on the smart glasses' display, and the user responds appropriately to the customer based on the displayed content. In this way, the conversation progresses smoothly and the customer's needs can be met.
[1492] Specific examples
[1493] For example, if you are explaining about a new smartphone in a physical store, the server may infer that the customer is interested in its features but is concerned about the price. The smart glasses will display a speech plan such as, "This smartphone uses the latest battery technology and will last all day with everyday use," and the user will explain based on that.
[1494] Example prompts to input to the generative AI model
[1495] "Generate utterance plans that recognize customer emotions and provide detailed information about products they're interested in. For example, create an utterance plan for when a customer is interested in the features of a new smartphone and has questions about the details, but is also concerned about the price."
[1496] This system will enable real-time customer service in physical stores by accurately understanding the emotions and intentions of the person you are talking to. Furthermore, by using visual devices to provide appropriate feedback to users and visualize the content of the conversation, it is expected that users' conversation skills will improve.
[1497] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1498] Step 1:
[1499] Data collection
[1500] The device uses a camera and microphone to collect facial expressions, gestures, and voice data of the conversation partner and the user. Specifically, the camera captures video data of the conversation partner and the user in real time, and the microphone collects voice data. This collected data is then sent from the device to a server.
[1501] Input: Camera video, microphone audio
[1502] Output: Video data, audio data
[1503] Step 2:
[1504] Video Data Analysis
[1505] The server analyzes the received video data and extracts facial features of the user and the conversation partner. It uses facial recognition technology to identify facial expressions and estimate emotions based on these. Specifically, it uses OpenCV for facial recognition and inputs the features into an emotion recognition model built with TensorFlow and Keras.
[1506] Input: Video data
[1507] Output: Facial features, emotion estimation results
[1508] Step 3:
[1509] Voice data analysis
[1510] The server analyzes the received voice data, converts it into text using speech recognition technology, and estimates emotional changes through intonation analysis. Specifically, the server converts the voice data into text using the Google Cloud Speech-to-Text API and then performs intonation analysis.
[1511] Input: Audio data
[1512] Output: Text data, emotion change estimation results
[1513] Step 4:
[1514] Integrated analysis of intentions and emotions
[1515] The server integrates the facial expression features, emotion estimation results, text data, and emotion change estimation results to comprehensively estimate the intentions and emotions of the conversation partner and the user.The server uses an emotion engine to comprehensively analyze the data and estimate the intentions.
[1516] Input: Facial features, emotion estimation results, text data, emotion change estimation results
[1517] Output: Intention estimation result, emotion estimation result
[1518] Step 5:
[1519] Generate an utterance plan
[1520] Based on the analysis results, the server generates a speech plan that matches the emotional state and intentions of the conversation partner and the user. Specifically, it uses a generative AI model to create an appropriate speech plan.
[1521] Input: Intention estimation result, emotion estimation result
[1522] Output: Utterance plan
[1523] Step 6:
[1524] Viewing the utterance plan
[1525] The user's device (smart glasses) displays the speech plan sent from the server on its display, allowing the user to proceed with the conversation according to the content displayed on the display. Specifically, text and icons are displayed on the display.
[1526] Input: Utterance plan
[1527] Output: Display
[1528] Step 7:
[1529] Feedback and conversation organization
[1530] The server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the terminal. The user's terminal displays this tree diagram, providing visual guidance on the next step in the conversation.
[1531] Input: Conversation progress data
[1532] Output: Tree diagram (conversation development diagram)
[1533] Step 8:
[1534] Self-Practice Mode
[1535] The device provides a self-practice mode, allowing users to practice speaking by themselves. The server provides speech samples, collects and analyzes the user's speech data in real time, and generates and sends feedback to the device based on the analysis results.
[1536] Input: User utterance data (self-practice mode)
[1537] Output: Feedback
[1538] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1539] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1540] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1541] [Fourth embodiment]
[1542] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1543] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1544] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1545] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1546] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1547] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1548] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1549] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1550] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1551] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1552] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1553] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1554] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1555] The present invention is a system that estimates the intentions and emotions of a conversation partner from their facial expressions, gestures, and voice, and presents an appropriate speech plan to the user, in order to facilitate interpersonal communication. Below, we will create a program for this system and explain its processing in detail.
[1556] System Configuration
[1557] The system consists of a terminal device (conversation prediction glasses) and a server device that performs data analysis. The terminal device is equipped with a camera and microphone, and the server device is equipped with an AI analysis model.
[1558] Program processing flow
[1559] Data collection
[1560] The device has a built-in camera to capture the facial expressions and gestures of the person you are talking to, and a microphone to collect audio data. When a conversation begins, the camera collects the video data of the person you are talking to, and the microphone collects the audio data in real time, and then sends the data to a server.
[1561] Data analysis
[1562] The server uses facial recognition technology to extract facial features of the conversation partner from the received video data and also analyzes gestures. At the same time, it converts the speech content into text using speech recognition technology from the audio data and estimates changes in emotion through intonation analysis. This data is input into an AI model, which then infers the other person's intentions and emotions.
[1563] Generate an utterance plan
[1564] The server generates a speech plan based on the analysis results, according to the intentions and emotions expressed by the conversation partner. For example, if a customer shows interest in a particular product, the server generates a speech plan to explain the details of that product and recommends it to the user. The generated speech plan is sent to the terminal.
[1565] User Feedback
[1566] The device displays the speech plan sent from the server on its display and provides the user with appropriate speech content. By continuing the conversation according to the displayed content, the user can communicate appropriately with the other person's intentions and emotions.
[1567] Organizing the conversation
[1568] Furthermore, the server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the device. The device displays this tree diagram, providing visual guidance on the next step in the conversation for the user.
[1569] Self-Practice Mode
[1570] The device has a self-practice mode, allowing users to practice speaking by themselves. In this mode, the server provides speaking samples, analyzes the user's speech data, and provides evaluation and feedback, allowing users to improve their speaking skills.
[1571] Specific examples
[1572] For example, a salesperson puts on these conversation-reading glasses when negotiating with a customer. After the negotiation begins, the glasses' camera captures the customer's facial expressions and gestures, and the microphone collects what the customer says, and these data are sent to a server. The server analyzes the data and infers that the customer is interested in price. The server then generates a speech plan, such as "This product has excellent cost performance and will offer great benefits in the long run," and sends it to the device. The device then displays this on its screen, and the salesperson continues to explain the plan to the customer. In this way, the conversation progresses smoothly, enabling responses to be made that meet the customer's needs.
[1573] This system makes it possible to accurately grasp the intentions and emotions of the person you are talking to, enabling smooth communication, which in turn makes business negotiations, counseling, and even everyday conversations more effective.
[1574] The processing flow will be explained below.
[1575] Step 1: Your device is ready for the camera and microphone
[1576] Before the conversation begins, the device turns on the camera and microphone to prepare for data collection.
[1577] Show the user a ready notification.
[1578] Step 2: The device collects video and audio data from the person you are talking to.
[1579] The camera continuously captures the facial expressions and gestures of the person you are talking to.
[1580] The microphone records the voice of the person you are talking to in real time.
[1581] The collected data is sent to the server in real time.
[1582] Step 3: The server analyzes the video data
[1583] The server analyzes the received video data and extracts facial expression features of the conversation partner.
[1584] Use facial recognition technology to estimate emotions from facial expressions.
[1585] Gesture analysis is performed to infer intentions from hand and body movements.
[1586] Step 4: The server analyzes the audio data
[1587] The server analyzes the received voice data and converts the spoken content into text using voice recognition technology.
[1588] The text of the speech is analyzed, and emotions are inferred from word choice and intonation.
[1589] Step 5: The server estimates the other person's intentions and emotions
[1590] The features extracted from the video and audio data are input into the AI model.
[1591] The AI model estimates the overall intention and emotions of the conversation partner.
[1592] Step 6: The server generates an appropriate utterance plan
[1593] Based on the estimation results, the AI generates the optimal speech plan for the user.
[1594] A speech plan includes specific phrases and next steps in the conversation.
[1595] Step 7: The server sends the speech plan to the device.
[1596] The server transmits the generated speech plan to the terminal.
[1597] Optimize and transmit data to ensure stable communication and prevent delays.
[1598] Step 8: The device displays the speech plan to the user.
[1599] The terminal displays the speech plan received from the server on the display.
[1600] The user continues the conversation by following what is displayed.
[1601] Step 9: The server organizes the conversation
[1602] As the conversation progresses, the server organizes the user's comments and the analysis results.
[1603] Structure the conversation content as a tree diagram or flowchart.
[1604] Step 10: The server sends the organized conversation to the device.
[1605] The organized conversation content and the next steps to take are sent to the terminal.
[1606] Help users make good decisions.
[1607] Step 11: The terminal displays the tree diagram to the user
[1608] The terminal displays the tree diagram or flowchart received from the server.
[1609] Check what the user should say next and the flow.
[1610] Step 12: User selects and starts self-practice mode
[1611] The user selects the self-practice mode through the terminal interface.
[1612] An interface provides the user with practice mode settings.
[1613] Step 13: The server provides and analyzes speech samples
[1614] The server provides speech samples for use in self-practice mode.
[1615] Data spoken by users is collected in real time and analyzed.
[1616] Step 14: The server generates feedback and sends it to the device
[1617] The server evaluates the user's speech data and generates suggestions for improvement and appropriate feedback.
[1618] The generated feedback is sent to the device.
[1619] Step 15: The device displays feedback to the user
[1620] The terminal displays the feedback received from the server to the user.
[1621] The user improves their speaking skills based on the feedback.
[1622] Through these processing steps, users can properly understand the intentions and emotions of their conversation partners and communicate smoothly. In addition, by utilizing the self-practice mode, users can effectively improve their speaking skills.
[1623] Example 1
[1624] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1625] Interpersonal communication requires rapid and accurate understanding of the intentions and emotions of the person you are speaking with, but the current technology for doing so is insufficient. It is particularly difficult to respond appropriately to changes in the other person's intentions and emotions in situations such as business negotiations and counseling. There is also a lack of effective feedback systems for improving speaking skills through self-practice. A system that solves these issues is needed.
[1626] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1627] In this invention, the server includes means for collecting facial expressions and gestures of a conversation partner using a camera, means for collecting voice data of the conversation partner using a microphone, means for transmitting the collected video and voice data to the server in real time, means for extracting facial features from the received video data using face recognition technology and performing gesture analysis, means for converting the speech content into text from the voice data using voice recognition technology and estimating emotional changes through intonation analysis, and means for generating and displaying an appropriate speech plan for the user based on the estimated intentions and emotions. This makes it possible to accurately grasp the intentions and emotions of the conversation partner and achieve smooth communication.
[1628] A "camera" is a device for collecting video data, and in this invention is used to capture the facial expressions and gestures of a conversation partner.
[1629] A "microphone" is a device for collecting voice data, and in the present invention is used to clearly capture the voice of a conversation partner.
[1630] A "server" is a computer system for analyzing and processing data, and in the present invention, it has the role of analyzing video data and audio data and providing speech plans to users.
[1631] "Facial expression features" are a set of feature values that indicate the emotions and intentions of a conversation partner, extracted using face recognition technology.
[1632] "Gesture analysis" is the process of analyzing a conversation partner's hand movements and gestures to infer non-verbal intentions and emotions.
[1633] "Speech recognition technology" is a technology that analyzes voice data and converts it into text data, and is used in the present invention to analyze the content of speech from a conversation partner.
[1634] "Intonation analysis" is a technology that analyzes the intonation and rhythm of voice data to estimate changes in emotion.
[1635] An "utterance plan" is a proposal of utterance content to support the progress of a conversation, generated based on analyzed data.
[1636] A "tree diagram" is a diagram that organizes the development of a conversation in a hierarchical structure to visually show the progress of the conversation.
[1637] The "self-practice mode" is a function that allows the user to practice speaking on their own, and includes providing speaking samples and feedback.
[1638] The present invention is a system that estimates the intentions and emotions of a conversation partner from their facial expressions, gestures, and voice, and presents an appropriate speech plan to the user, in order to facilitate interpersonal communication. Below, we will create a program for this system and explain its processing in detail.
[1639] System Configuration
[1640] The main components of this system are a terminal device (conversation prediction glasses) and a server device that performs data analysis. The terminal device is equipped with a camera and microphone, and the server device is equipped with an AI analysis model.
[1641] Program processing flow and specific functions
[1642] Data collection
[1643] The device has a built-in camera for capturing facial expressions and gestures of the conversation partner, and a microphone for capturing audio data. When a conversation begins, the camera captures the conversation partner's video data and the microphone captures audio data in real time, and these data are then sent to a server. The hardware used is a built-in camera (e.g., an HD camera) and a highly sensitive microphone.
[1644] Data analysis
[1645] The server uses facial recognition technology to extract facial features of the conversation partner from the received video data and also performs gesture analysis. At the same time, it uses speech recognition technology to convert the speech into text from the audio data and estimates emotional changes through intonation analysis. Software libraries such as TensorFlow and OpenCV are used for these analyses. For example, Google's Speech-to-Text API is sometimes used to analyze audio data.
[1646] Generate an utterance plan
[1647] The server generates a speech plan based on the analysis results, taking into account the intentions and emotions of the conversation partner. This is done using a pre-trained generative AI model. For example, if a customer shows interest in a particular product, a speech plan to explain the details of that product is generated. The generated speech plan is sent to the device.
[1648] User Feedback
[1649] The device displays the speech plan sent from the server on its display and provides the user with appropriate speech content. By continuing the conversation according to the displayed content, the user can communicate appropriately with the other person's intentions and emotions.
[1650] Organizing the conversation
[1651] Furthermore, the server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the device. The device displays this tree diagram, providing visual guidance on the next step in the conversation for the user.
[1652] Self-Practice Mode
[1653] The device has a self-practice mode, allowing users to practice speaking by themselves. In this mode, the server provides speaking samples, analyzes the user's speech data, and provides evaluation and feedback, allowing users to improve their speaking skills.
[1654] Specific examples
[1655] For example, a salesperson puts on these conversation-reading glasses when negotiating with a customer. After the negotiation begins, the glasses' camera captures the customer's facial expressions and gestures, and the microphone collects what the customer says, and these data are sent to a server. The server analyzes the data and infers that the customer is interested in price. The server then generates a speech plan, such as "This product has excellent cost performance and will offer great benefits in the long run," and sends it to the device. The device then displays this on its screen, and the salesperson continues to explain the plan to the customer. In this way, the conversation progresses smoothly, enabling responses to be made that meet the customer's needs.
[1656] Example prompts for generative AI models
[1657] Input data: Customer facial expression information, voice data
[1658] AI model: Estimating conversation partner's intentions and emotions
[1659] Output: It is assumed that the customer is interested in the price. Generate the utterance plan "This product offers excellent value for money and offers great benefits in the long run."
[1660] As described above, the present invention is a system for understanding the intentions and emotions of a conversation partner and supporting appropriate communication, and is particularly effective in situations such as business negotiations and counseling.
[1661] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1662] Step 1:
[1663] Data collection
[1664] The device uses a camera and microphone to capture facial expressions, gestures, and voice data from the person you're talking to. The camera captures high-resolution video, while the microphone suppresses background noise to capture clear audio.
[1665] input:
[1666] Video and audio data of the person you are talking to
[1667] output:
[1668] Collected video and audio data
[1669] Specific behavior:
[1670] When the user begins interacting with a customer, the device's camera captures the customer's facial expressions and gestures, and the microphone begins recording what the customer says.
[1671] Step 2:
[1672] Data transmission
[1673] The device transmits the collected video and audio data to the server in real time over a fast and secure network.
[1674] input:
[1675] Collected video and audio data
[1676] output:
[1677] Data sent to the server
[1678] Specific behavior:
[1679] The device encrypts the collected data and sends it to a server using Wi-Fi or 5G networks.
[1680] Step 3:
[1681] Data analysis
[1682] The server extracts facial features from the received video data using facial recognition technology and also performs gesture analysis. It also converts the speech into text using speech recognition technology and estimates emotional changes through intonation analysis. TensorFlow and OpenCV are used as software libraries.
[1683] input:
[1684] Video and audio data sent to the server
[1685] output:
[1686] Extracted facial features, textualized speech, and estimated emotional changes
[1687] Specific behavior:
[1688] The server uses TensorFlow to analyze the video data and extract facial expression and gesture features. At the same time, Google's Speech-to-Text API converts the audio data into text, and the emotion analysis API analyzes emotions from intonation.
[1689] Step 4:
[1690] Generate an utterance plan
[1691] The server generates a speech plan based on the analysis results, taking into account the intentions and emotions of the conversation partner. Using a generative AI model, it constructs an appropriate speech plan that corresponds to the estimated intentions and emotions.
[1692] input:
[1693] Extracted facial features, textualized speech, and estimated emotional changes
[1694] output:
[1695] Generated utterance plan
[1696] Specific behavior:
[1697] If the server's AI model determines that the customer is interested in price, it generates a speech plan such as, "This product has excellent cost performance."
[1698] Step 5:
[1699] Sending a speech plan
[1700] The server then sends the generated speech plan to the device, again in real time, allowing the user to receive immediate feedback.
[1701] input:
[1702] Generated utterance plan
[1703] output:
[1704] Speech plan sent to the device
[1705] Specific behavior:
[1706] The speech plan generated by the server is immediately sent to the device, and a notification is displayed on the device's display.
[1707] Step 6:
[1708] User Feedback
[1709] The device displays the received speech plan on its display and provides the user with appropriate speech content, which the user can refer to as they proceed with the conversation.
[1710] input:
[1711] Speech plan sent to the device
[1712] output:
[1713] The speech plan presented to the user
[1714] Specific behavior:
[1715] The terminal display will show "This product has excellent cost performance," and the user will convey this statement to the customer.
[1716] Step 7:
[1717] Organizing the conversation
[1718] The server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the terminal.
[1719] input:
[1720] Analyzed conversation content
[1721] output:
[1722] Speech content organized as a tree diagram
[1723] Specific behavior:
[1724] The server analyzes the flow of conversation, and if there is a high level of "interest in price," it generates a node for "additional information related to price" and displays it in the form of a tree diagram on the terminal display.
[1725] Step 8:
[1726] Self-Practice Mode
[1727] The device has a self-practice mode, allowing users to practice speaking by themselves. In this mode, the server provides speaking samples and analyzes the user's speech data to provide evaluation and feedback.
[1728] input:
[1729] Speech samples entered by the user
[1730] output:
[1731] Ratings and Feedback
[1732] Specific behavior:
[1733] When a user selects practice mode, the server sends sample utterances, such as "greeting practice," to the device. As the user speaks, the audio is analyzed and pronunciation and intonation are evaluated in real time.
[1734] (Application example 1)
[1735] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1736] Existing conversation support systems have difficulty understanding the intentions and emotions of the target, and are therefore unable to present appropriate speech plans or response methods for achieving effective communication. Furthermore, in security services, where it is necessary to quickly understand the intentions and emotions of visitors and passersby and respond appropriately, the lack of such technology poses a serious problem.
[1737] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1738] In this invention, the server includes means for collecting facial expressions and gestures of a target using a camera, means for collecting audio data of the target using a microphone, means for analyzing the collected video and audio data to estimate the intention and emotion of the target, means for displaying an appropriate speech plan to the user based on the estimated intention and emotion, and means for suggesting a response method to the user based on the estimated intention and emotion, thereby enabling security guards to quickly grasp the intention and emotion of visitors and passersby and to determine an appropriate response method.
[1739] A "camera" is a photographic device for capturing facial expressions and gestures of a subject.
[1740] A "microphone" is a recording device for collecting audio data of a subject.
[1741] "Video data" refers to image and video information collected using a camera.
[1742] "Audio data" is sound information collected using a microphone.
[1743] "Facial expression" refers to emotions and intentions that can be read from the movement of facial muscles.
[1744] A "gesture" is an expression of intent that can be interpreted through hand or body movements.
[1745] "Intention" refers to the purpose or direction of an action that a subject has.
[1746] "Emotion" refers to a subject's state of mind or mood.
[1747] A "speech plan" is a plan for providing appropriate conversation content.
[1748] A "display" is a display device for presenting visual information to a user.
[1749] A "server" is a computer system that analyzes collected data and provides results.
[1750] A "tree diagram" is a diagram that visually organizes the content and development of a conversation.
[1751] The "self-practice mode" is a mode in which the user can practice speaking by himself.
[1752] "Feedback" refers to evaluation of the user's speech and advice for improvement.
[1753] This invention is an interpersonal communication support system specialized for security applications, and employs an approach to infer the intentions and emotions of a conversation partner from their facial expressions, gestures, and voice, and provide appropriate response methods.
[1754] System Configuration
[1755] The system includes smart glasses (terminal devices) and a server device that performs data analysis. The smart glasses are equipped with a camera and microphone, and the server device is equipped with an AI analysis model.
[1756] Data collection
[1757] The terminal device is equipped with a camera to capture the facial expressions and gestures of the visitor and a microphone to collect audio data. When a security guard comes into contact with a visitor, the camera collects the subject's video data and the microphone collects audio data in real time, and these data are then sent to a server.
[1758] Data analysis
[1759] The server uses facial recognition technology to extract facial features from the received video data and also analyzes gestures. At the same time, it converts the speech into text using speech recognition technology and estimates emotional changes through intonation analysis. This data is input into an AI model, which then infers the subject's intentions and emotions.
[1760] Generate speech plans and responses
[1761] Based on the analysis results, the server generates a speech plan and appropriate response method according to the visitor's intentions and emotions. For example, if the visitor seems nervous, the server generates advice recommending polite responses and sends it to the terminal device. The generated speech plan and response method are then displayed on the smart glasses' display.
[1762] User Feedback
[1763] The terminal device displays the speech plan and response method sent from the server on its display and provides the security guard with appropriate response content. By responding to the visitor according to the displayed content, the security guard can appropriately deal with the visitor's intentions and emotions.
[1764] Organizing the conversation
[1765] Furthermore, the server analyzes the progress of the conversation in real time, organizes the content and development of the conversation into a tree diagram, and sends it to the terminal device, which displays this tree diagram and provides visual guidance to the security guard on the next steps in the conversation.
[1766] Self-Practice Mode
[1767] The terminal device is equipped with a self-practice mode, allowing security guards to practice speaking by themselves. In this mode, the server provides speaking samples, analyzes the security guards' speech data, and provides evaluation and feedback, allowing the security guards to improve their speaking skills.
[1768] Examples and prompts
[1769] Specific examples
[1770] A security guard wears smart glasses at the entrance to a facility and handles visitors. If the visitor's facial expression appears stiff and tense, the camera and microphone collect information and the information is analyzed by a server. As a result of the analysis, the smart glasses' display displays advice such as, "It is assumed that the visitor is nervous. Please handle the visitor politely and confirm the purpose of their visit."
[1771] Prompt Sentence Examples
[1772] "Advice to display when a visitor's facial expression appears tense:
[1773] We suspect the visitor may be nervous. Please be polite and confirm the purpose of their visit."
[1774] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1775] Step 1:
[1776] Data collection
[1777] The device (smart glasses) uses a camera to collect video data of the subject's facial expressions and gestures, and simultaneously collects audio data using a microphone. The input is video data from the camera and audio data from the microphone, and the output is the collected raw video and audio data.
[1778] Step 2:
[1779] Data transmission
[1780] The terminal transmits the collected video and audio data to the server in real time. The input is the collected video and audio data, and the output is the data transmitted to the server.
[1781] Step 3:
[1782] Facial recognition and facial expression feature extraction
[1783] The server uses facial recognition technology to detect the target face from the received video data and extract facial expression features. The input is the video data, and the output is the extracted facial expression features. The software used is a facial recognition library such as OpenCV.
[1784] Step 4:
[1785] Gesture Analysis
[1786] The server analyzes hand and body movements from video data to detect gestures. The input is video data, and the output is detected gesture data. The software used is an AI analysis model.
[1787] Step 5:
[1788] Converting audio data to text
[1789] The server converts the voice data into text data using speech recognition technology. The input is voice data and the output is text data. The software used is a speech recognition library (e.g., Google Cloud Speech-to-Text).
[1790] Step 6:
[1791] Intonation analysis
[1792] The server analyzes the intonation of the voice data and estimates emotional changes. The input is the voice data, and the output is estimated emotional data. The software used is an AI analysis model.
[1793] Step 7:
[1794] Intention and emotion estimation
[1795] The server inputs the extracted facial features, gesture data, text data, and emotion data into an AI model to comprehensively estimate the subject's intention and emotion. The input is each feature data, and the output is estimated intention and emotion information.
[1796] Step 8:
[1797] Generate speech plans and responses
[1798] The server generates a speech plan and a response method for the user based on the estimated intention and emotion. For example, if the visitor is nervous, it generates advice including an appropriate response method. The input is the intention and emotion information, and the output is the speech plan and a response method.
[1799] Step 9:
[1800] User Feedback
[1801] The terminal displays the speech plan and response method sent from the server on a display and provides them to the security guard. The input is the speech plan and response method, and the output is the content displayed on the display. The security guard responds to the visitor according to the displayed content.
[1802] Step 10:
[1803] Organizing the conversation
[1804] The server analyzes ongoing conversations in real time and organizes them into a tree diagram. The input is real-time conversation data, and the output is a tree diagram of the organized conversations. The terminal displays this tree diagram to help security guards understand the next steps.
[1805] Step 11:
[1806] Self-Practice Mode
[1807] The terminal provides an interface that allows the user to select self-practice mode. The server provides speech samples, analyzes the user's speech data, and provides evaluation and feedback. The input is the user's speech data, and the output is evaluation and feedback, allowing the security guard to improve their speaking skills.
[1808] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1809] This invention is a system that combines conversation prediction glasses with an emotion engine to more accurately recognize the emotions and intentions of both the user and the conversation partner and present appropriate speech plans. Below, we will create a program for this system and explain its processing in detail.
[1810] System Configuration
[1811] The system includes a terminal device (conversation prediction glasses), a server device that performs data analysis, and an emotion engine with emotion recognition capabilities. The terminal device is equipped with a camera and microphone, and the server device is equipped with an AI analysis model and emotion engine.
[1812] Program processing flow
[1813] Data collection
[1814] The device uses a camera and microphone to collect facial expressions, gestures, and voice data of the user and the conversation partner. When a conversation begins, the camera collects video data of the conversation partner and the user, and the microphone collects voice data in real time, and then transmits this data to a server.
[1815] Data analysis
[1816] The server analyzes the received video and audio data and extracts facial and audio features of the conversation partner and the user. First, facial recognition technology is used to infer emotions from facial expressions, followed by gesture analysis. Furthermore, speech recognition technology is used to convert the speech into text from the audio data, and emotional changes are inferred through intonation analysis. This data is input into an AI model and emotion engine, which infers the intentions and emotions of the conversation partner and the user.
[1817] Generate an utterance plan
[1818] The server generates a speech plan based on the analysis results, taking into account the emotional state and intentions of the conversation partner and the user. For example, the server may generate a speech plan that provides detailed explanations based on a topic that the conversation partner is interested in, or suggest relaxing topics to reduce the stress the user is feeling. The generated speech plan is then sent to the device.
[1819] User Feedback
[1820] The device displays the speech plan sent from the server on its display and provides the user with appropriate speech content. By continuing the conversation according to the displayed content, the user can communicate appropriately with the other person's intentions and emotions.
[1821] Organizing the conversation
[1822] Furthermore, the server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the device. The device displays this tree diagram, providing visual guidance on the next step in the conversation for the user.
[1823] Self-Practice Mode
[1824] The device offers a self-practice mode, allowing users to practice speaking by themselves. In this mode, the server provides speech samples, collects and analyzes the user's speech data in real time, and generates and sends feedback to the device based on the analysis results.
[1825] Specific examples
[1826] For example, a sales representative puts on these conversation-predicting glasses when negotiating with a customer. After the negotiation begins, the glasses' camera captures the facial expressions and gestures of both the customer and the sales representative, while the microphone collects speech, and these data are sent to a server. The server analyzes the data and infers that the customer is beginning to show interest in the product but is concerned about the price. The emotion engine then confirms that the sales representative is explaining with confidence. The server generates a speech plan, such as "This product has excellent cost performance and will offer great benefits in the long run," and sends it to the device. The device displays this on its screen, and the sales representative explains based on it. In this way, the conversation progresses smoothly, enabling responses that meet the customer's needs.
[1827] This system allows both the user and the conversation partner to accurately understand each other's intentions and emotions, enabling smooth and effective communication, which will enable more effective business negotiations, counseling, and everyday conversations.
[1828] The processing flow will be explained below.
[1829] Step 1: Your device is ready for the camera and microphone
[1830] Before the conversation begins, the device turns on the camera and microphone to prepare for data collection.
[1831] Show the user a ready notification.
[1832] Step 2: The device collects video and audio data of the conversation partner and the user.
[1833] The camera continuously captures the facial expressions and gestures of the conversation partner and the user.
[1834] The microphone records the voices of the conversation partner and the user in real time.
[1835] The collected data is sent to the server in real time.
[1836] Step 3: The server analyzes the video data of the person you are talking to.
[1837] The server analyzes the received video data and extracts facial features of the conversation partner.
[1838] Use facial recognition technology to estimate emotions from facial expressions.
[1839] Gesture analysis is performed to infer intentions from hand and body movements.
[1840] Step 4: The server analyzes the voice data of the other party.
[1841] The server analyzes the received voice data and converts the spoken content into text using voice recognition technology.
[1842] The text of the speech is analyzed, and emotions are inferred from word choice and intonation.
[1843] Step 5: The server analyzes the user's video data.
[1844] The server analyzes the received video data and extracts the user's facial expression features.
[1845] Use facial recognition technology to estimate emotions from facial expressions.
[1846] Gesture analysis is performed to estimate the current state and intentions from hand and body movements.
[1847] Step 6: The server analyzes the user's voice data
[1848] The server analyzes the received voice data and converts the spoken content into text using voice recognition technology.
[1849] The system analyzes the text of the speech and estimates the user's emotions from word choice and intonation.
[1850] Step 7: The server integrates and estimates the intentions and emotions of the conversation partner and the user.
[1851] The features extracted from the video and audio data are input into the AI model and emotion engine.
[1852] AI models and emotion engines estimate the overall intent and emotions of the conversation partner and the user.
[1853] Step 8: The server generates an appropriate utterance plan
[1854] Based on the estimation results, the AI generates the optimal speech plan for the user.
[1855] A speech plan includes specific phrases and next steps in the conversation.
[1856] Step 9: The server sends the speech plan to the device.
[1857] The server transmits the generated speech plan to the terminal.
[1858] Optimize and transmit data to ensure stable communication and prevent delays.
[1859] Step 10: The device displays the speech plan to the user.
[1860] The terminal displays the speech plan received from the server on the display.
[1861] The user continues the conversation by following what is displayed.
[1862] Step 11: The server organizes the conversation
[1863] As the conversation progresses, the server organizes the user's comments and the analysis results.
[1864] Structure the conversation content as a tree diagram or flowchart.
[1865] Step 12: The server sends the organized conversation to the device.
[1866] The organized conversation content and the next steps to take are sent to the terminal.
[1867] Help users make good decisions.
[1868] Step 13: The terminal displays the tree diagram to the user
[1869] The terminal displays the tree diagram or flowchart received from the server.
[1870] Check what the user should say next and the flow.
[1871] Step 14: User selects and starts self-practice mode
[1872] The user selects the self-practice mode through the terminal interface.
[1873] An interface provides the user with practice mode settings.
[1874] Step 15: The server provides and analyzes speech samples.
[1875] The server provides speech samples for use in self-practice mode.
[1876] Data spoken by users is collected in real time and analyzed.
[1877] Step 16: The server generates feedback and sends it to the device
[1878] The server evaluates the user's speech data and generates suggestions for improvement and appropriate feedback.
[1879] The generated feedback is sent to the device.
[1880] Step 17: The device displays feedback to the user
[1881] The terminal displays the feedback received from the server to the user.
[1882] The user improves their speaking skills based on the feedback.
[1883] Through these processing steps, users can accurately understand their own and their conversation partner's emotions and intentions, enabling smoother and more effective communication. Furthermore, by utilizing the self-practice mode, users can effectively improve their speaking skills.
[1884] Example 2
[1885] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1886] In conversation, users often have difficulty accurately reading the intentions and emotions of the other person, making smooth communication difficult. They also often lack appropriate responses to the anxiety and stress they feel. Furthermore, even when practicing by themselves, the lack of concrete feedback makes it difficult to improve their speech.
[1887] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1888] In this invention, the server includes means for collecting facial expressions and gestures of the conversation partner and the user using a camera, means for collecting voice data of the conversation partner and the user using a microphone, means for analyzing the collected video and voice data to estimate the intentions and emotions of the conversation partner and the user, means for displaying an appropriate speech plan to the user based on the estimated intentions and emotions, means for organizing the progress of the conversation in a tree diagram based on the analysis results and displaying it to the user, and means for providing an interface that allows the user to select a self-practice mode and receive feedback for improving their speech. This allows the user to accurately understand the intentions and emotions of the conversation partner and take appropriate measures. Furthermore, the self-practice mode makes it easier to improve speech, thereby improving the overall quality of communication.
[1889] A "camera" is an optical device for collecting video data.
[1890] A "conversational partner" refers to the other person with whom the user is interacting.
[1891] "User" refers to a person who uses this system to have a conversation.
[1892] "Facial expressions" refer to facial movements and expressions, and are important non-verbal elements for indicating emotions and intentions.
[1893] "Gestures" refer to hand and body movements and are non-verbal expressions used to convey intentions and emotions.
[1894] A "microphone" is an acoustic device for collecting audio data.
[1895] "Audio data" refers to collected sound information, including speech content and intonation.
[1896] "Collect" refers to obtaining and storing data.
[1897] "Analysis" is the act of extracting and analyzing information based on collected data.
[1898] "Intention" refers to the purpose or thoughts of the conversation partner or user.
[1899] "Emotion" refers to the emotional state felt by a conversation partner or user.
[1900] A "speech plan" is a specific speech content suggested by the system based on a specific intention or emotion.
[1901] "Display" is the act of enabling a user to visually confirm information.
[1902] A "tree diagram" is a diagram that visually shows the progress of a conversation and the relationship between topics.
[1903] "Interface" refers to the operating screen or means by which a user interacts with a system.
[1904] "Feedback" refers to information that provides suggestions for improvement or evaluation of a user's actions or results.
[1905] The present invention is a system that includes a terminal device (conversation prediction glasses) worn by a user, a server device that performs data analysis, and an emotion engine with emotion recognition functionality. Specific embodiments for implementing this system are described below.
[1906] terminal device
[1907] The terminal device is equipped with a camera and a microphone. These hardware components are used to collect video and audio data of the user and the conversation partner in real time. Specifically, the camera captures facial expressions and gestures, and the microphone collects audio data of the conversation. The terminal device has the function of transmitting the collected data to a server device via a network.
[1908] Server device
[1909] The server device is equipped with an AI model and emotion engine for analyzing the received data. From the video data, facial recognition technology is used to extract facial expression features and perform gesture analysis. For the audio data, speech is converted into text using speech recognition technology, and emotional changes are estimated through intonation analysis. This data is input into the AI model and emotion engine, which ultimately estimates the intentions and emotions of the conversation partner and the user.
[1910] Providing Feedback
[1911] The server generates a speech plan based on the analysis results and transmits it to the terminal device. The terminal device displays this speech plan on a display and provides it to the user. The user can continue the conversation based on the displayed speech plan, and can respond appropriately to the intentions and feelings of the conversation partner.
[1912] Organizing the conversation
[1913] The server analyzes the progress of the conversation in real time and organizes the content and development of the conversation into a tree diagram. This tree diagram is sent to the terminal device, helping the user visually grasp the next flow of the conversation.
[1914] Self-Practice Mode
[1915] The terminal device is equipped with an interface that allows the user to select a self-practice mode, in which the server collects the user's speech, analyzes it in real time, and generates feedback, which helps the user improve their speech.
[1916] Specific examples
[1917] For example, a sales representative puts on conversation-predicting glasses when negotiating with a customer. When the negotiation begins, the device's camera captures the facial expressions and gestures of both the customer and the sales representative, and the microphone collects the audio of the conversation. This data is sent to a server, which uses analysis technology to infer the customer's emotions and intentions. Based on the analysis results, a speech plan such as "This product has excellent cost performance and will provide great benefits in the long run" is generated and displayed on the device. The user can explain the situation based on the displayed speech plan, allowing the negotiation to proceed smoothly.
[1918] Prompt Sentence Examples
[1919] "Using conversation-reading glasses, please generate a specific scenario in which a salesperson is negotiating with a customer. Include a process for reading the customer's emotions and interests and proposing an appropriate speech plan."
[1920] This system accurately recognizes the intentions and emotions of both the user and the conversation partner, enabling accurate and smooth communication, promoting effective dialogue in business negotiations, counseling, and everyday conversations.
[1921] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1922] Step 1:
[1923] The terminal collects video and audio of the conversation partner and the user in real time.
[1924] Specific operation: The device's camera captures the face, facial expressions, and gestures of the user and the person in conversation, and the microphone collects audio. The input is the video and audio data acquired through the camera and microphone, and the output is the collected data packets.
[1925] Step 2:
[1926] The terminal transmits the collected video and audio data to the server.
[1927] Specific operation: Video data is packetized as image frames, and audio data is packetized as audio clips, and then sent to the server over the network. The input is the collected data packets, and the output is the server to which the data was sent.
[1928] Step 3:
[1929] The server analyzes the received video data and extracts facial expression features.
[1930] Specific operation: Using facial recognition technology in the server, the system identifies the facial parts of the conversation partner and the user and calculates facial features such as smile, sadness, surprise, etc. The input is the received video data, and the output is the extracted facial features.
[1931] Step 4:
[1932] The server performs gesture analysis and infers emotions and intentions from the collected gesture data.
[1933] Specific operation: Using the server's gesture recognition algorithm, hand movements and body poses are analyzed to estimate emotional states such as interest or tension. The input is the received video data and existing gesture data, and the output is the estimated emotion or intention.
[1934] Step 5:
[1935] The server analyzes the voice data and converts the spoken content into text.
[1936] How it works: The speech recognition system analyzes recorded speech and converts it into text data. It then detects emotional changes through intonation analysis. The input is the received speech data, and the output is the text of the speech and an indicator of emotional changes.
[1937] Step 6:
[1938] The server uses AI models and emotion engines to infer the intentions and emotions of the user and their conversation partner.
[1939] How it works: The extracted features are input into an AI model based on previous research and training data to estimate emotions and intentions. An emotion engine is used to further improve accuracy. The inputs are facial features, gesture data, text data of spoken content, and indicators of emotional changes, and the output is estimated intentions and emotions.
[1940] Step 7:
[1941] Based on the analysis results, the server generates a speech plan that is appropriate for the emotions and intentions of the user and their conversation partner.
[1942] How it works: Using a generative AI model, it automatically determines the appropriate utterance content for the situation (e.g., question, explanation, relaxed topic, etc.) and constructs it as an utterance plan. The input is the estimated intent and emotion, and the output is the generated utterance plan.
[1943] Step 8:
[1944] The server transmits the generated speech plan to the terminal.
[1945] Specific operation: The text information generated as a speech plan is packetized and sent to the terminal via the network. The input is the generated speech plan, and the output is the terminal to which it was sent.
[1946] Step 9:
[1947] The terminal displays the received speech plan on the display.
[1948] Specific operation: The display module in the terminal displays the text of the speech plan on the screen, providing a visual for the user. The input is the received speech plan, and the output is the speech plan displayed on the display.
[1949] Step 10:
[1950] The user continues the conversation according to the content displayed on the display.
[1951] Specific operation: The user refers to the displayed speech plan and speaks at the appropriate time to progress the conversation. The input is the speech plan displayed on the display, and the output is the ongoing conversation.
[1952] Step 11:
[1953] The server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the terminal.
[1954] Specific operation: The collected data is analyzed one after another, and data for visualizing the conversation topics and progress as a tree diagram is generated and sent to the device. The input is the ongoing conversation data and analysis results, and the output is data organized as a tree diagram.
[1955] Step 12:
[1956] The terminal displays a tree diagram showing the progress of the conversation on its display.
[1957] Specific operation: The display module in the terminal displays the tree diagram on the screen, providing the user with a visual of the next conversation flow. The input is the tree diagram data, and the output is the tree diagram displayed on the display.
[1958] Step 13:
[1959] The terminal provides a self-practice mode for the user to practice speaking.
[1960] Specific behavior: The device displays a practice interface and presents speech samples to the user. The input is the user's selection, and the output is the self-practice mode interface.
[1961] Step 14:
[1962] The server collects and analyzes user speech data in real time.
[1963] Specific operations: Receives collected speech data from the device, performs voice analysis and intonation analysis, generates feedback based on the analysis results, and sends it to the device. The input is the collected speech data, and the output is the generated feedback.
[1964] (Application example 2)
[1965] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1966] Conventional dialogue support systems have difficulty accurately grasping the emotions and intentions of their conversation partners and providing appropriate speech plans in real time. Furthermore, in brick-and-mortar stores, staff are required to read customers' emotions and respond effectively, but current technology makes this difficult. Furthermore, feedback and visualization of conversation content using visual devices are insufficient, limiting the ability to improve users' dialogue skills.
[1967] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1968] In this invention, the server includes means for collecting facial expressions and gestures of a conversation partner using a camera, means for collecting voice data of the conversation partner using a microphone, means for analyzing the collected video and voice data to estimate the intentions and emotions of the conversation partner, means for displaying an appropriate speech plan to the user based on the estimated intentions and emotions, and means for generating an utterance plan based on the analysis results and displaying it on the user's visual device. This allows for accurate understanding of the conversation partner's emotions and intentions, enabling real-time customer service in physical stores. Furthermore, by using the visual device to provide appropriate feedback to the user and visualize the conversation content, it is expected that the user's dialogue skills will improve.
[1969] A "camera" is an optical device for collecting video data.
[1970] A "conversational partner" refers to a person other than the user with whom the conversation is taking place.
[1971] "Facial expression" refers to information including emotions and intentions expressed through the movement of facial muscles.
[1972] A "gesture" is a communication action using body and hand movements.
[1973] A "microphone" is an acoustic device for collecting audio data.
[1974] "Audio data" is digital information that digitizes audio.
[1975] "Analysis" refers to the process of extracting patterns and information from collected data.
[1976] "Intention" refers to the purpose or thoughts of the person you are speaking with.
[1977] "Emotions" are information that indicates the feelings and psychological state of the person you are talking to.
[1978] "Inference" is the process of predicting intentions and emotions from collected data.
[1979] "Utterance plan" refers to the content and plan of what the user will say.
[1980] "User" refers to a person using the smart glasses or system.
[1981] "Display" is the act of presenting information to a visual device.
[1982] A "visual device" is a device that allows a user to receive information visually.
[1983] This invention relates to a system that uses a camera, a microphone, a server, and a visual device to analyze the emotions and intentions of a conversation partner and presents an appropriate speech plan to the user. A specific implementation method of this system will be described below.
[1984] Hardware and software used
[1985] The system uses the following main hardware and software:
[1986] Camera: An optical device for collecting video data in real time, such as a camera mounted on smart glasses.
[1987] Microphone: Acoustic equipment for collecting audio data, such as a microphone in smart glasses.
[1988] Server: A high-performance computing device that performs data analysis and emotion recognition. It uses OpenCV for face recognition, Google Cloud Speech-to-Text API for voice recognition, and TensorFlow and Keras for the emotion engine.
[1989] Visual device: A device for presenting information to a user, for example, the display of smart glasses.
[1990] Data collection
[1991] First, a user puts on the smart glasses and starts a conversation with a customer. The camera captures the facial expressions and gestures of the person they are talking to, and the microphone captures audio data. This data is then sent to the server in real time.
[1992] Data analysis
[1993] The server analyzes the received video and audio data. Facial features are extracted from the video data using OpenCV, and emotions are estimated using TensorFlow and Keras. The audio data is converted to text using the Google Cloud Speech-to-Text API, and emotional changes are estimated through intonation analysis. This allows the intentions and emotions of the person being spoken to be identified.
[1994] Generate an utterance plan
[1995] The server generates a speech plan based on the analysis results. For example, if a customer is interested in a particular product but is concerned about the price, the server creates a speech plan such as, "This product offers excellent value for money and offers great benefits in the long run."
[1996] Display and Feedback
[1997] The generated speech plan is displayed on the smart glasses' display, and the user responds appropriately to the customer based on the displayed content. In this way, the conversation progresses smoothly and the customer's needs can be met.
[1998] Specific examples
[1999] For example, if you are explaining about a new smartphone in a physical store, the server may infer that the customer is interested in its features but is concerned about the price. The smart glasses will display a speech plan such as, "This smartphone uses the latest battery technology and will last all day with everyday use," and the user will explain based on that.
[2000] Example prompts to input to the generative AI model
[2001] "Generate utterance plans that recognize customer emotions and provide detailed information about products they're interested in. For example, create an utterance plan for when a customer is interested in the features of a new smartphone and has questions about the details, but is also concerned about the price."
[2002] This system will enable real-time customer service in physical stores by accurately understanding the emotions and intentions of the person you are talking to. Furthermore, by using visual devices to provide appropriate feedback to users and visualize the content of the conversation, it is expected that users' conversation skills will improve.
[2003] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2004] Step 1:
[2005] Data collection
[2006] The device uses a camera and microphone to collect facial expressions, gestures, and voice data of the conversation partner and the user. Specifically, the camera captures video data of the conversation partner and the user in real time, and the microphone collects voice data. This collected data is then sent from the device to a server.
[2007] Input: Camera video, microphone audio
[2008] Output: Video data, audio data
[2009] Step 2:
[2010] Video Data Analysis
[2011] The server analyzes the received video data and extracts facial features of the user and the conversation partner. It uses facial recognition technology to identify facial expressions and estimate emotions based on these. Specifically, it uses OpenCV for facial recognition and inputs the features into an emotion recognition model built with TensorFlow and Keras.
[2012] Input: Video data
[2013] Output: Facial features, emotion estimation results
[2014] Step 3:
[2015] Voice data analysis
[2016] The server analyzes the received voice data, converts it into text using speech recognition technology, and estimates emotional changes through intonation analysis. Specifically, the server converts the voice data into text using the Google Cloud Speech-to-Text API and then performs intonation analysis.
[2017] Input: Audio data
[2018] Output: Text data, emotion change estimation results
[2019] Step 4:
[2020] Integrated analysis of intentions and emotions
[2021] The server integrates the facial expression features, emotion estimation results, text data, and emotion change estimation results to comprehensively estimate the intentions and emotions of the conversation partner and the user.The server uses an emotion engine to comprehensively analyze the data and estimate the intentions.
[2022] Input: Facial features, emotion estimation results, text data, emotion change estimation results
[2023] Output: Intention estimation result, emotion estimation result
[2024] Step 5:
[2025] Generate an utterance plan
[2026] Based on the analysis results, the server generates a speech plan that matches the emotional state and intentions of the conversation partner and the user. Specifically, it uses a generative AI model to create an appropriate speech plan.
[2027] Input: Intention estimation result, emotion estimation result
[2028] Output: Utterance plan
[2029] Step 6:
[2030] Viewing the utterance plan
[2031] The user's device (smart glasses) displays the speech plan sent from the server on its display, allowing the user to proceed with the conversation according to the content displayed on the display. Specifically, text and icons are displayed on the display.
[2032] Input: Utterance plan
[2033] Output: Display
[2034] Step 7:
[2035] Feedback and conversation organization
[2036] The server analyzes the progress of the conversation in real time, organizes the content and development of the conversation as a tree diagram, and sends it to the terminal. The user's terminal displays this tree diagram, providing visual guidance on the next step in the conversation.
[2037] Input: Conversation progress data
[2038] Output: Tree diagram (conversation development diagram)
[2039] Step 8:
[2040] Self-Practice Mode
[2041] The device provides a self-practice mode, allowing users to practice speaking by themselves. The server provides speech samples, collects and analyzes the user's speech data in real time, and generates and sends feedback to the device based on the analysis results.
[2042] Input: User utterance data (self-practice mode)
[2043] Output: Feedback
[2044] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2045] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2046] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2047] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification p...
Claims
1. means for collecting facial expressions and gestures of a conversation partner using a camera; A means for collecting voice data of a conversation partner using a microphone; A means for analyzing the collected video and audio data to estimate the intentions and emotions of the conversation partner; means for displaying an appropriate utterance plan to a user based on the estimated intention and emotion; A system including:
2. 2. The system according to claim 1, further comprising means for organizing the collected conversation contents into a tree diagram or the like and displaying the same to the user.
3. means for providing an interface through which a user can select a self-practice mode; 10. The system of claim 1, further comprising means for analyzing the user's speech data and providing feedback for speech improvement.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A