system

The system addresses the challenge of accurately analyzing feelings and impressions in communication by generating optimal reply messages, enhancing communication effectiveness and naturalness through user-selected impressions and natural language processing.

JP2026071540APending Publication Date: 2026-04-30SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-17
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Existing communication systems struggle to accurately analyze the feelings and impressions of others, leading to stress and inefficiencies in human interactions, particularly in digital and real-time communication scenarios.

Method used

A system that collects and analyzes communication information using natural language processing and emotion analysis to generate optimal reply messages based on user-selected impressions, ensuring effective and natural communication.

Benefits of technology

Enables users to efficiently convey desired impressions and emotions, improving communication quality by providing contextually appropriate and natural-sounding responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071540000001_ABST
    Figure 2026071540000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Means for collecting communication information, A means of analyzing collected communication information to identify the emotional state of the other party, A means for generating selectable impression options to present to the user based on an identified emotional state, A means of generating the optimal reply message based on the selected impression, The means of sending a reply message to the other party, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005] ,

[0001] The technology of this disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] With the diversification of communication in modern society, people are required to accurately grasp the feelings and impressions of others and find a reply method suitable for the situation. However, many people feel a great deal of stress in doing this manually and face difficulties in conducting effective conversations. To address such a situation, there is a need for a technology that can accurately analyze the feelings and impressions of others and automatically generate an optimal reply thereto.

Means for Solving the Problems

[0005] This invention provides a system that identifies the emotional state of the other party by collecting and analyzing communication information. Based on the analysis results, the system presents the user with multiple impression options, allowing the user to select their preferred impression. Based on this selection, the system automatically generates a reply message and sends it to the other party. Furthermore, by utilizing natural language processing technology, the accuracy of the analysis is improved, and the user can customize the message, thereby achieving more natural and effective communication.

[0006] "Communication information" refers to electronic messages and audio data exchanged between users.

[0007] "Means of collection" refers to hardware and software used to collect and store communication information.

[0008] "Means of analysis" refers to the technologies and processes used to analyze collected communication information and understand the meaning and characteristics of the data.

[0009] "Emotional state" refers to the emotions an individual experiences at a particular moment, and is a state of feeling that can be identified through analysis.

[0010] "Impression options" refer to different patterns of impressions that users can choose from, intended to convey to others.

[0011] "Generating methods" refer to the technologies and processes that create the optimal reply message based on selected impressions and analysis results.

[0012] A "reply message" refers to a form of communication from a user to another user, structured based on analysis results and impression selections.

[0013] "Means of transmission" refers to the technology and processes used to deliver the generated reply message to the recipient.

[0014] "Natural language processing technology" refers to the technology used to understand and process human language using computers.

[0015] The "user interface" refers to the interfaces and tools for users to interact with the system.

Brief Description of Drawings

[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiment for Carrying out the Invention

[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.

[0020] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0021] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0024] [First Embodiment]

[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0037] This invention is implemented as a system that analyzes the emotions and impressions of others through communication tools used by users on a daily basis and generates the optimal response. The system primarily functions between a terminal and a server.

[0038] The device collects data from various communication methods used by the user, such as LINE, social media, email, and voice calls. This device monitors the content of messages and calls sent and received in real time or at regular intervals, encrypts the collected communication information for privacy protection, and sends it to a server. During this process, voice data is also converted to text.

[0039] The server receives data sent from the terminal and analyzes it. This analysis utilizes emotion analysis AI and natural language processing technology. This allows the server to identify the other party's emotional state and understand the impression gained from the conversation. Based on the analysis results, the server has the function of presenting the user with multiple impression options to choose from.

[0040] Next, the user selects the impression they wish to convey from the presented impression options. Once the user has made their selection, the server uses a generative AI model to generate a natural-sounding reply message based on the selected impression. In this process, contextual language generation is performed to ensure the message is natural and appropriate.

[0041] Finally, the device presents the generated reply message to the user, obtains the user's confirmation, and then sends it to the recipient. This entire process allows the user to efficiently convey the desired impression to the recipient, resulting in better communication.

[0042] For example, consider a scenario where a user receives a message from a friend asking, "How are you doing lately?" In this case, the device collects the message, and the server analyzes it. If the analysis detects that the friend is "concerned," the user is presented with options to respond, such as "reassure" or "cheer up." If the user selects "reassure," a corresponding reply such as "I'm fine, how about you?" is generated and sent. In this way, the system of the present invention helps users communicate effectively.

[0043] The following describes the processing flow.

[0044] Step 1:

[0045] The device monitors messages and voice calls sent and received from the communication tools used by the user, and collects necessary communication information in real time. For voice calls, it also performs the process of converting speech to text.

[0046] Step 2:

[0047] The terminal encrypts the collected communication information and implements security measures to prevent data leakage before sending the data to the server.

[0048] Step 3:

[0049] The server analyzes the communication information received from the terminal. Using sentiment analysis AI, it identifies the recipient's emotional state from the message, and uses natural language processing technology to understand the impression the recipient has of the user.

[0050] Step 4:

[0051] The server presents the user with multiple impression options based on the analysis results. These options include choices related to the other person's feelings (e.g., reassurance, encouragement).

[0052] Step 5:

[0053] The user selects the impression they want to give to the other person from a list of impression options presented by the server.

[0054] Step 6:

[0055] The server generates the optimal reply message using a generative AI model based on the user's selection. This reply message is then adjusted to best match the selected impression.

[0056] Step 7:

[0057] The device presents the generated reply message to the user and obtains the user's confirmation and final approval.

[0058] Step 8:

[0059] After the user reviews and approves the reply message, the device sends the message to the recipient.

[0060] Through this series of steps, users can effectively communicate with others and leave the desired impression.

[0061] (Example 1)

[0062] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0063] In modern digital communication, users often struggle to properly understand the other person's emotions and formulate appropriate responses. This problem is particularly pronounced in text-only communication, where the emotional nuances of the other person cannot be accurately captured. As a result, misunderstandings and inefficient communication can occur, negatively impacting the quality of relationships.

[0064] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0065] In this invention, the server includes a device for collecting communication data, a device for analyzing the collected communication data and identifying the emotional state of another person, and a device for creating selectable impression options to display to the user based on the identified emotional state. This enables the user to efficiently understand the other person's emotional state and generate a response message for optimal communication.

[0066] "Communication data" refers to all digital data exchanged between people, such as messages and voice information, that are sent and received through electronic means.

[0067] "Device" refers to hardware, software, or a combination thereof designed to perform a specific function.

[0068] "Analysis" refers to the process of extracting specific information from collected data and using that information to gain new insights.

[0069] "Emotional state" refers to the type and intensity of emotions that the communication partner is presumed to be experiencing, and is often revealed through text analysis.

[0070] "Users" refer to people who use this system to communicate.

[0071] "Selectable impression options" refers to multiple choices presented to the user to select the impression they want to convey to the other person.

[0072] A "response message" refers to a reply generated based on the user's intent, and its content is appropriate to the recipient's intentions and feelings.

[0073] This invention is a system aimed at enabling users to understand the emotions of others through digital communication and to create appropriate response messages. This system primarily operates between a device held by the user and a server located in the cloud.

[0074] The device collects communication data from communication tools that users use daily, such as messaging applications, social networking services, email, and voice calls. The collected data is protected using standard encryption technologies such as AES and RSA to ensure security and privacy. Voice data is converted to text using real-time speech recognition software.

[0075] When the server receives encrypted communication data sent from a terminal, it decrypts it and analyzes the data using sentiment analysis AI and natural language processing technology. In this analysis process, machine learning libraries (e.g., TENSORFLOW® and PyTorch) are used to identify the other party's emotional state and tone of communication.

[0076] Based on the analysis results, the server presents the user with multiple selectable impression options. These options are created based on predefined templates and AI-generated suggestions, and are important for improving the user experience of the product or service.

[0077] The generative AI model used in the analysis process receives prompts based on the user's selection and generates natural, contextually appropriate responses. For example, if a user receives a message from a friend saying "How are you?", they might instruct the model with the prompt, "My friend messaged me 'How are you?' Please generate a reassuring response." In this way, the generated response messages help users communicate quickly and accurately with others.

[0078] Finally, the generated message is presented to the user on their device, where it can be modified and reviewed as needed before being sent to the recipient. This entire process allows users to achieve smoother and more effective communication.

[0079] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0080] Step 1:

[0081] The device collects communication data from various communication methods used by the user. Inputs include messages and voice data. Specifically, the device periodically acquires text messages and voice communication data through APIs and interfaces, and converts voice data to text using speech recognition software. The output is encrypted text data.

[0082] Step 2:

[0083] The terminal securely encrypts the collected communication data and sends it to the server. The input is the text data obtained in step 1. Specifically, the data is protected using encryption technologies such as AES and RSA and uploaded to the server via the internet. The output is the encrypted communication data received by the server.

[0084] Step 3:

[0085] The server decrypts encrypted communication data sent from the terminal and analyzes it using sentiment analysis AI and natural language processing technology. The input is encrypted communication data. Specifically, machine learning algorithms are used to analyze the emotional state and important keywords in the text and identify the emotional state of others. The output is the analyzed emotional state and communication tone information.

[0086] Step 4:

[0087] The server generates and provides the user with selectable impression options based on the analysis results. The input is the emotional state information obtained in step 3. Specifically, it generates impression options (e.g., "reassure," "encourage," etc.) using AI or a predefined rule set and sends them to the terminal. The output is the impression options presented to the user.

[0088] Step 5:

[0089] The user selects their desired impression from the impression options presented by the server. The input is the impression options received from the server. Specifically, the user makes a selection by tapping or clicking on an option, and this selection information is sent to the server. The output is the selected impression information.

[0090] Step 6:

[0091] The server utilizes a generative AI model based on the user's selected impression to generate an appropriate response message. The input is the user's selected impression information. Specifically, the generative AI model is given a prompt, and a process is executed to generate a natural response that takes the context into account. The output is the generated response message.

[0092] Step 7:

[0093] The terminal receives a response message generated from the server and presents it to the user. The input is the generated response message. Specifically, it displays a UI that allows the user to review the response and make corrections if necessary. After obtaining user approval, it sends the response message to the other party. The output is the final response message that is sent.

[0094] (Application Example 1)

[0095] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0096] In the security field in particular, immediate responses tailored to the situation on-site are often required. However, conventional systems have a problem in that communication after anomaly detection relies on manual processes, making it difficult to respond quickly and accurately.

[0097] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0098] In this invention, the server includes means for analyzing communication information and identifying the emotional state of the other party, means for detecting anomalies from the surrounding environment and communicating with people on site in response to those anomalies, and means for analyzing people's emotions in real time and generating appropriate instruction messages. This enables emotionally appropriate communication and rapid response tailored to the situation on site.

[0099] "Communication information" refers to digital data that is sent and received between a user and others in the form of voice data, text messages, image data, etc.

[0100] "Emotional state" refers to the psychological reactions and emotional shifts of the other person obtained through communication, and is usually expressed in categories such as positive, negative, and neutral.

[0101] "Impression options" are choices that allow users to select what kind of impression they want to give to the other person, and are presented to the user based on an analysis of their emotional state.

[0102] A "reply message" is a message generated based on the user's selections, and its purpose is to create the best possible impression on the recipient.

[0103] "Detecting anomalies from the surrounding environment" means discovering unusual situations or events by analyzing environmental data such as sound, video, and vibration.

[0104] "Communicating with people on-site" means communicating with people at the scene via voice or text when an anomaly is detected, and sharing necessary information.

[0105] "Analyzing people's emotions in real time" is a process of instantly processing communication data and understanding their emotional state as a result.

[0106] An "appropriate instruction message" is an instruction or guidance message generated based on an analyzed emotional state, and its content is designed to provide appropriate instructions according to the situation on site.

[0107] This invention relates to a system applied in the security field and provides technology for effective communication with people on site in response to abnormal situations within a building.

[0108] The server utilizes a speech recognition module to convert audio data collected from the field into text data. Specifically, it uses the Google® Cloud Speech-to-Text API to convert speech to text in real time. The transcribed information is then analyzed using natural language processing techniques and used to identify emotional states.

[0109] The server analyzes emotions from video data using Amazon Rekognition and other tools via an emotion analysis module. This allows for the evaluation of the psychological state of people on-site and their individual reactions to abnormal situations. Based on the results of the real-time emotion analysis, prompt sentences are supplied to an AI model to obtain the optimal response.

[0110] The AI ​​model generates appropriate instructional messages based on the user's selected impressions, responding to any unusual situation. These messages are transmitted to people on-site via speakers or monitors. For example, if a suspicious sound is detected during a night patrol and workers are on alert, the server will generate and send a message stating, "Security check in progress. Please remain calm."

[0111] As an example of a prompt, we will use a prompt in the format of, "Create a reassuring message for employees who are feeling anxious." Based on this prompt, the generative AI model will provide flexible instructions that are relevant to the context.

[0112] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0113] Step 1:

[0114] The device collects audio and video data from within the building. Audio data is acquired via a microphone, and video data is acquired via a camera. This raw data is transmitted to the server in real time.

[0115] Step 2:

[0116] The server converts the received audio data into text data using the Google Cloud Speech-to-Text API. The input is audio data, and the output is the corresponding text data. This conversion makes the audio information processable as text.

[0117] Step 3:

[0118] The server analyzes the converted text data using natural language processing techniques to identify the emotional state in the communication. The input is text data, and the output is a category of emotional state (e.g., positive, negative, neutral). This enables data processing that understands the emotions behind the text.

[0119] Step 4:

[0120] The server uses Amazon Rekognition to analyze emotional states from video data. The input is video data, and the output is an evaluation of emotional states based on the video. This process makes it possible to understand people's emotions through visual data.

[0121] Step 5:

[0122] Based on these sentiment analysis results, the server generates prompt sentences to give to the generating AI model. These prompt sentences are guiding sentences designed to elicit an appropriate response corresponding to the analyzed sentiment. The input is the evaluation result of the emotional state, and the output is the prompt sentence.

[0123] Step 6:

[0124] The server uses a generative AI model to generate appropriate instruction messages based on the prompt. The input is the prompt, and the output is the generated instruction message. This message is contextualized by the generative AI model.

[0125] Step 7:

[0126] The terminal transmits the generated instruction message to people on-site as actual voice or text via a speaker or monitor. This ensures that messages that help ensure safety and promote reassurance are delivered. Input is the generated instruction message, and output is the transmitted audio or visual information.

[0127] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0128] This invention is implemented as a system that combines the analysis of communication information with an emotion engine that recognizes the user's own emotions, in order to facilitate smooth communication between the user and the other party. The system functions in cooperation with a terminal and a server to provide an optimal dialogue environment for the user.

[0129] The device is responsible for collecting message and voice data sent and received from the user's communication tools. Furthermore, it analyzes biosignals (e.g., heart rate, voice tone) using an emotion engine to read the user's emotional state. Based on this input data, the emotion engine recognizes the user's current emotions and sends that information to the server.

[0130] The server first receives communication information sent from the terminal and analyzes this information using natural language processing technology. During the analysis, it identifies the other party's emotional state during the communication and understands the impression the user has of that party. The server also takes into account the user's emotional state, which is also transmitted from the terminal. Based on this information, impression choices to be presented to the user are generated, and these choices are optimized to reflect the user's emotions.

[0131] The user selects the impression they want to convey to the other person from a set of impression options provided by the server. This selection process is designed to reflect the user's own emotional state. Once the user completes their selection, the server uses a generative AI model to generate a natural-sounding reply message that takes the selected impression into account. The generated message is displayed to the user and can be customized.

[0132] Ultimately, the device sends a reply message to the recipient that has been approved by the user. This system helps users effectively convey their emotions while leaving a desirable impression on the recipient.

[0133] As a concrete example, consider a scenario where a user is asked by their partner, "How was your day?" If the data collected by the device indicates that the user is feeling "tired," the server provides the user with options such as "I want to relax" or "I want to be encouraged." If the user selects "I want to relax," the server generates and sends a message such as, "I was a little tired today, but it was great chatting!" In this way, the system of the present invention takes the user's own emotions into consideration and provides the necessary support in communication.

[0134] The following describes the processing flow.

[0135] Step 1:

[0136] The device collects text messages and voice data from the communication tools used by the user. It also measures biometric signals such as the user's voice tone and heart rate using an emotion engine, preparing to estimate the user's emotional state.

[0137] Step 2:

[0138] The device encrypts the collected communication information and the user's estimated emotional state when sending them to the server to protect the confidentiality of the data.

[0139] Step 3:

[0140] The server analyzes the communication information received from the terminal and uses natural language processing technology to identify the other party's emotional state and impression. At the same time, it receives the user's emotional state transmitted from the terminal and integrates the information from both.

[0141] Step 4:

[0142] The server generates a selection of impression options to present to the user, based on the user's emotions and the identified emotional state of the other party. These options are optimized to be consistent with the emotions the user is currently feeling.

[0143] Step 5:

[0144] The user selects the impression they want to convey to the other person from a set of impression options presented by the server. This selection corresponds to the user's own emotional state and facilitates appropriate dialogue.

[0145] Step 6:

[0146] The server uses a generative AI model to generate the optimal reply message based on the user's selected impression. This message is adjusted based on the user's emotional state and selected impression.

[0147] Step 7:

[0148] The terminal displays the generated reply message to the user, allowing the user to review the content and customize it as needed.

[0149] Step 8:

[0150] Once the user approves the message, the device sends the message to the recipient, completing the communication.

[0151] This process allows users to recognize their own emotional state and effectively communicate the impression they want to make to others.

[0152] (Example 2)

[0153] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0154] Conventional communication systems struggle to appropriately reflect user emotions and generate messages that create a desirable impression on the recipient. This problem hinders efficient and natural dialogue, especially in communication where emotions are complexly intertwined. Therefore, there is a need for technology that simultaneously considers the user's emotions and the recipient's state to generate the optimal response.

[0155] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0156] In this invention, the server includes means for collecting communication data, means for analyzing the communication data to identify the emotional state of the other party, and means for analyzing the user's biosignals to recognize emotions. This makes it possible to generate a response message that gives the most appropriate impression to the other party while taking the user's emotions into consideration.

[0157] "Communication data" refers to all information transmitted and received through electronic means, and specifically includes text messages and voice data.

[0158] "The other party's emotional state" refers to a description of the other party's emotions and mental state, analyzed from communication data.

[0159] "Biosignals" refer to data that indicates the user's physical state, including physical measurements such as heart rate and voice tone.

[0160] "Impression options" refer to multiple choices presented to a user to select the impression they want to convey to the other party.

[0161] A "response message" refers to a message that facilitates natural conversation, generated with consideration for the emotional state of both the user and the other party.

[0162] A "user interface" refers to an interactive mechanism that allows a user to interact with a computer system and perform various operations.

[0163] This invention is a system designed to facilitate communication between a user and another party. It utilizes communication data and biosignals to analyze emotions and generate optimal response messages. The roles of each component of this system are described below.

[0164] The device serves as a communication tool used by users on a daily basis and plays a role in collecting text messages and voice data. This device, such as a smartphone or PC, acquires this data in real time through dedicated applications and APIs, and also collects biometric signals such as heart rate and voice tone. Wearable devices and dedicated sensors are used to collect these biometric signals.

[0165] The collected data is analyzed by a server. The server possesses powerful processing capabilities and utilizes natural language processing techniques and machine learning algorithms to analyze the communication data. In the future, models such as BERT (Bidirectional Encoder Representations from Transformers) and LSTM (Long Short-Term Memory) may be used. Based on the emotional state of the other party and the user's emotions obtained from the emotion analysis, the server generates impression options for the user.

[0166] The user can review the impression options suggested by the server and select the one they deem most effective. Based on this selection, the server uses a generative AI model to create a natural and harmonious response message. This response message is presented to the user through the user interface and can be customized as needed.

[0167] For example, if a user is asked "How was it?" by their partner, and the device recognizes the user's emotion as "tired," the server will present impressions such as "I want to relax" or "I want to be encouraged." If the user selects "I want to relax," the server will generate a message such as "I'm a little tired today, but it was great chatting!" This process helps users effectively convey their emotions while maintaining natural communication.

[0168] The AI ​​generation model is input with a prompt such as, "If the user's emotional state is fatigued, they have selected 'I want to relax' as the impression they want to give. Based on this, please generate a natural reply message." In this way, the system provides an accurate and flexible conversational experience.

[0169] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0170] Step 1:

[0171] The device collects text messages and voice data in real time from the user's communication tools. Furthermore, it measures the user's heart rate and voice tone through biometric sensors. This data is acquired as input necessary for recognizing emotional states.

[0172] Step 2:

[0173] Inside the device, an emotion engine analyzes the user's emotions based on biosignal data. This analysis uses a machine learning model, with heart rate and voice tone fluctuations serving as foundational data for inferring the user's emotions. The analysis results are output as the user's emotional state and sent to the server.

[0174] Step 3:

[0175] The server receives text messages sent from terminals and analyzes their content using natural language processing (NLP) techniques. Specifically, it performs processes such as text tokenization, sentiment analysis, and keyword extraction to identify the recipient's emotional state. This analysis result serves as foundational data for evaluating the impression a user has of another person.

[0176] Step 4:

[0177] The server generates impression options to present to the user based on the analyzed emotional states of the user and the other party. In this process, optimal options are provided based on past data and statistical models, taking into account the user's communication style and situation. The generated options are then sent from the server to the user.

[0178] Step 5:

[0179] The user reviews the impression options provided by the server and selects the impression they wish to convey. The user's selection is intuitively made through a GUI, and this selection significantly influences the generation of the next response message.

[0180] Step 6:

[0181] The server generates a natural response message using a generative AI model based on the impression selected by the user. In this generation process, the prompt "If the user's emotional state is fatigued, they have selected 'I want to relax' as the impression they want to give to the other party. Based on this, please generate a natural reply message." is used as input data, and the AI ​​model outputs optimized text.

[0182] Step 7:

[0183] The terminal displays the response message received from the server to the user, providing an interface that allows the user to customize the content as needed. The user then makes a final confirmation of the message.

[0184] Step 8:

[0185] The device sends a response message to the other party, acknowledging the user's input. This completes the communication and effectively conveys the user's emotions.

[0186] (Application Example 2)

[0187] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0188] In online communication, it is necessary to appropriately understand users' emotions and provide support for making optimal responses and product suggestions based on those emotions. Conventional systems lack the ability to provide individualized support that takes into account the user's emotional state, resulting in a decline in the quality of dialogue. Furthermore, the lack of systems that offer product suggestions tailored to the user's emotions means that appropriate purchasing support is not being provided.

[0189] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0190] In this invention, the server includes means for collecting communication information, means for analyzing the collected communication information and identifying the emotional state of the other party, and means for generating selectable impression options and product suggestions to present to the user based on the identified emotional state. This enables communication tailored to the user's emotions and optimal product suggestions.

[0191] "Communication information" refers to information such as messages and voice data exchanged between a user and another party.

[0192] "Emotional state" refers to the state of mind and mood of the user or the other party, and is an indicator used to consider the impact on communication by understanding the changes in this state.

[0193] An "impression option" is a set of choices generated to allow a user to select what kind of impression they want to give to the other person.

[0194] A "reply message" is a response that a user sends to another person, optimized based on their emotional state and impression choices.

[0195] A "user interface" refers to the screens and means through which a user operates the system, and to review and customize generated reply messages and product recommendations.

[0196] "Product recommendation" refers to the presentation of product information by suggesting the most suitable product based on the user's emotions and choices.

[0197] "Natural language processing technology" is a general term for technologies used to enable computers to understand and process human language.

[0198] The system for carrying out this invention includes a terminal used by the user and a server that processes information. First, the terminal collects the user's voice data and text messages. This information is analyzed by an emotion engine to identify the user's emotional state. The emotion engine uses biosignals such as the user's heart rate and voice tone as input data to accurately recognize emotions.

[0199] The server receives communication information sent from the terminal and analyzes the information using natural language processing technology. As a result of the analysis, the server identifies the other party's emotional state and understands the impression the user has. Based on this information, the server generates selectable impression options and product recommendations. When the user selects an impression or product from the provided options, the server uses a generative AI model to generate a natural response message corresponding to the selected impression.

[0200] Users can review the generated messages and product suggestions displayed on their device and customize them as needed. The device sends a reply message to the recipient after receiving user approval. One possible application of this system is user support during product purchases in a virtual store. For example, if a user expresses a desire to "relax," the server might suggest a product such as, "This candle promotes relaxation."

[0201] Through this system, users can receive optimal communication and product suggestions that take their emotions into consideration, improving the quality of the interaction. An example of a prompt message is, "Generate a message that provides optimal product suggestions based on the emotion analysis results."

[0202] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0203] Step 1:

[0204] The device collects user voice data and text messages. This data is sent to an emotion engine, which analyzes it along with biosignals (e.g., heart rate, voice tone) necessary to recognize the user's emotional state. The input is the user's raw communication data, and the output is the user's emotional state as recognized by the emotion engine. Machine learning algorithms are applied based on the data to recognize the emotional state.

[0205] Step 2:

[0206] The server receives the user's emotional state and communication information transmitted from the terminal. Based on the received data, natural language processing techniques are used to analyze the other party's emotional state and identify the emotional relationship between the two parties. The input for this step is the user's emotional state and communication information, and the output is the other party's emotional state and what they understood it to be. Text mining and sentiment analysis techniques are used for the analysis.

[0207] Step 3:

[0208] The server generates selectable impression options and product suggestions based on the analysis results. This is done using an AI model that generates options that correspond to the user's emotional state and the other party's emotions. The input is the analysis results from the previous stage, and the output is the impression options and product suggestions. Text generation techniques, including topic modeling, are used to generate the impression options and product suggestions.

[0209] Step 4:

[0210] The user reviews impression options and product suggestions via their device and makes a selection. The user's selection is sent to the server, which uses a generative AI model to generate a natural-sounding reply message based on this selection. The input is the user's selection, and the output is a natural-sounding reply message. Advanced natural language generation techniques, such as transformer models, are used for generation.

[0211] Step 5:

[0212] The terminal displays the generated message to the user, and after the user's final confirmation or customization, sends the approved reply message to the recipient. At this stage, the input consists of the reply message and the user's customizations, while the output is the final response message sent. The sending function utilizes standard communication protocols for data transmission.

[0213] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0214] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0215] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0216] [Second Embodiment]

[0217] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0218] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0219] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0220] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0221] The microphone 238 receives voice signals from the user 20 and accepts instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0222] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0223] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0224] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0225] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0226] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0227] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0228] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0229] This invention is implemented as a system that analyzes the emotions and impressions of others through communication tools used by users on a daily basis and generates the optimal response. The system primarily functions between a terminal and a server.

[0230] The device collects data from various communication methods used by the user, such as LINE, social media, email, and voice calls. This device monitors the content of messages and calls sent and received in real time or at regular intervals, encrypts the collected communication information for privacy protection, and sends it to a server. During this process, voice data is also converted to text.

[0231] The server receives data sent from the terminal and analyzes it. This analysis utilizes emotion analysis AI and natural language processing technology. This allows the server to identify the other party's emotional state and understand the impression gained from the conversation. Based on the analysis results, the server has the function of presenting the user with multiple impression options to choose from.

[0232] Next, the user selects the impression they wish to convey from the presented impression options. Once the user has made their selection, the server uses a generative AI model to generate a natural-sounding reply message based on the selected impression. In this process, contextual language generation is performed to ensure the message is natural and appropriate.

[0233] Finally, the device presents the generated reply message to the user, obtains the user's confirmation, and then sends it to the recipient. This entire process allows the user to efficiently convey the desired impression to the recipient, resulting in better communication.

[0234] For example, consider a scenario where a user receives a message from a friend asking, "How are you doing lately?" In this case, the device collects the message, and the server analyzes it. If the analysis detects that the friend is "concerned," the user is presented with options to respond, such as "reassure" or "cheer up." If the user selects "reassure," a corresponding reply such as "I'm fine, how about you?" is generated and sent. In this way, the system of the present invention helps users communicate effectively.

[0235] The following describes the processing flow.

[0236] Step 1:

[0237] The device monitors messages and voice calls sent and received from the communication tools used by the user, and collects necessary communication information in real time. For voice calls, it also performs the process of converting speech to text.

[0238] Step 2:

[0239] The terminal encrypts the collected communication information and implements security measures to prevent data leakage before sending the data to the server.

[0240] Step 3:

[0241] The server analyzes the communication information received from the terminal. Using sentiment analysis AI, it identifies the recipient's emotional state from the message, and uses natural language processing technology to understand the impression the recipient has of the user.

[0242] Step 4:

[0243] The server presents the user with multiple impression options based on the analysis results. These options include choices related to the other person's feelings (e.g., reassurance, encouragement).

[0244] Step 5:

[0245] The user selects the impression they want to give to the other person from a list of impression options presented by the server.

[0246] Step 6:

[0247] The server generates the optimal reply message using a generative AI model based on the user's selection. This reply message is then adjusted to best match the selected impression.

[0248] Step 7:

[0249] The device presents the generated reply message to the user and obtains the user's confirmation and final approval.

[0250] Step 8:

[0251] After the user reviews and approves the reply message, the device sends the message to the recipient.

[0252] Through this series of steps, users can effectively communicate with others and leave the desired impression.

[0253] (Example 1)

[0254] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0255] In modern digital communication, users often struggle to properly understand the other person's emotions and formulate appropriate responses. This problem is particularly pronounced in text-only communication, where the emotional nuances of the other person cannot be accurately captured. As a result, misunderstandings and inefficient communication can occur, negatively impacting the quality of relationships.

[0256] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0257] In this invention, the server includes a device for collecting communication data, a device for analyzing the collected communication data and identifying the emotional state of another person, and a device for creating selectable impression options to display to the user based on the identified emotional state. This enables the user to efficiently understand the other person's emotional state and generate a response message for optimal communication.

[0258] "Communication data" refers to all digital data exchanged between people, such as messages and voice information, that are sent and received through electronic means.

[0259] "Device" refers to hardware, software, or a combination thereof designed to perform a specific function.

[0260] "Analysis" refers to the process of extracting specific information from collected data and using that information to gain new insights.

[0261] "Emotional state" refers to the type and intensity of emotions that the communication partner is presumed to be experiencing, and is often revealed through text analysis.

[0262] "Users" refer to people who use this system to communicate.

[0263] "Selectable impression options" refers to multiple choices presented to the user to select the impression they want to convey to the other person.

[0264] A "response message" refers to a reply generated based on the user's intent, and its content is appropriate to the recipient's intentions and feelings.

[0265] This invention is a system aimed at enabling users to understand the emotions of others through digital communication and to create appropriate response messages. This system primarily operates between a device held by the user and a server located in the cloud.

[0266] The device collects communication data from communication tools that users use daily, such as messaging applications, social networking services, email, and voice calls. The collected data is protected using standard encryption technologies such as AES and RSA to ensure security and privacy. Voice data is converted to text using real-time speech recognition software.

[0267] When the server receives encrypted communication data sent from a terminal, it decrypts it and analyzes the data using sentiment analysis AI and natural language processing techniques. In this analysis process, machine learning libraries (e.g., TensorFlow and PyTorch) are used to identify the other party's emotional state and tone of communication.

[0268] Based on the analysis results, the server presents the user with multiple selectable impression options. These options are created based on predefined templates and AI-generated suggestions, and are important for improving the user experience of the product or service.

[0269] The generative AI model used in the analysis process receives prompts based on the user's selection and generates natural, contextually appropriate responses. For example, if a user receives a message from a friend saying "How are you?", they might instruct the model with the prompt, "My friend messaged me 'How are you?' Please generate a reassuring response." In this way, the generated response messages help users communicate quickly and accurately with others.

[0270] Finally, the generated message is presented to the user on their device, where it can be modified and reviewed as needed before being sent to the recipient. This entire process allows users to achieve smoother and more effective communication.

[0271] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0272] Step 1:

[0273] The device collects communication data from various communication methods used by the user. Inputs include messages and voice data. Specifically, the device periodically acquires text messages and voice communication data through APIs and interfaces, and converts voice data to text using speech recognition software. The output is encrypted text data.

[0274] Step 2:

[0275] The terminal securely encrypts the collected communication data and sends it to the server. The input is the text data obtained in step 1. Specifically, the data is protected using encryption technologies such as AES and RSA and uploaded to the server via the internet. The output is the encrypted communication data received by the server.

[0276] Step 3:

[0277] The server decrypts encrypted communication data sent from the terminal and analyzes it using sentiment analysis AI and natural language processing technology. The input is encrypted communication data. Specifically, machine learning algorithms are used to analyze the emotional state and important keywords in the text and identify the emotional state of others. The output is the analyzed emotional state and communication tone information.

[0278] Step 4:

[0279] The server generates and provides the user with selectable impression options based on the analysis results. The input is the emotional state information obtained in step 3. Specifically, it generates impression options (e.g., "reassure," "encourage," etc.) using AI or a predefined rule set and sends them to the terminal. The output is the impression options presented to the user.

[0280] Step 5:

[0281] The user selects the desired option from the impression options presented by the server. The input is the impression option received from the server. Specifically, the user makes a selection by tapping or clicking on the option, and the selection information is sent to the server. The output is the selected impression information.

[0282] Step 6:

[0283] The server utilizes a generation AI model based on the impression selected by the user to generate an appropriate response message. The input is the impression information selected by the user. Specifically, a process is executed where a prompt sentence is given to the generation AI model to generate a natural response considering the context. The output is the generated response message.

[0284] Step 7:

[0285] The terminal receives the response message generated by the server and presents it to the user. The input is the generated response message. Specifically, a UI is displayed where the user can confirm the response and make corrections if necessary. After obtaining the user's approval, the response message is sent to the recipient. The output is the finally sent response message.

[0286] (Application Example 1)

[0287] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0288] Especially in the field of security, immediate responses according to the on-site situation may be required. However, in conventional systems, communication after anomaly detection depends on human hands, and there is a problem that it is difficult to make a quick and accurate response.

[0289] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following respective means.

[0290] In this invention, the server includes means for analyzing communication information and identifying the emotional state of the other party, means for detecting anomalies from the surrounding environment and communicating with people on site in response to those anomalies, and means for analyzing people's emotions in real time and generating appropriate instruction messages. This enables emotionally appropriate communication and rapid response tailored to the situation on site.

[0291] "Communication information" refers to digital data that is sent and received between a user and others in the form of voice data, text messages, image data, etc.

[0292] "Emotional state" refers to the psychological reactions and emotional shifts of the other person obtained through communication, and is usually expressed in categories such as positive, negative, and neutral.

[0293] "Impression options" are choices that allow users to select what kind of impression they want to give to the other person, and are presented to the user based on an analysis of their emotional state.

[0294] A "reply message" is a message generated based on the user's selections, and its purpose is to create the best possible impression on the recipient.

[0295] "Detecting anomalies from the surrounding environment" means discovering unusual situations or events by analyzing environmental data such as sound, video, and vibration.

[0296] "Communicating with people on-site" means communicating with people at the scene via voice or text when an anomaly is detected, and sharing necessary information.

[0297] "Analyzing people's emotions in real time" is a process of instantly processing communication data and understanding their emotional state as a result.

[0298] An "appropriate instruction message" is an instruction or guidance message generated based on an analyzed emotional state, and its content is designed to provide appropriate instructions according to the situation on site.

[0299] This invention relates to a system applied in the security field and provides technology for effective communication with people on site in response to abnormal situations within a building.

[0300] The server utilizes a speech recognition module to convert audio data collected from the field into text data. Specifically, it uses the Google Cloud Speech-to-Text API to convert speech to text in real time. The transcribed information is then analyzed using natural language processing techniques to identify emotional states.

[0301] The server analyzes emotions from video data using Amazon Rekognition and other tools via an emotion analysis module. This allows for the evaluation of the psychological state of people on-site and their individual reactions to abnormal situations. Based on the results of the real-time emotion analysis, prompt sentences are supplied to an AI model to obtain the optimal response.

[0302] The AI ​​model generates appropriate instructional messages based on the user's selected impressions, responding to any unusual situation. These messages are transmitted to people on-site via speakers or monitors. For example, if a suspicious sound is detected during a night patrol and workers are on alert, the server will generate and send a message stating, "Security check in progress. Please remain calm."

[0303] As an example of a prompt, we will use a prompt in the format of, "Create a reassuring message for employees who are feeling anxious." Based on this prompt, the generative AI model will provide flexible instructions that are relevant to the context.

[0304] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0305] Step 1:

[0306] The terminal collects audio and video data inside the building. The audio data is obtained through a microphone, and the video data is obtained through a camera. Those raw data are sent to the server in real time.

[0307] Step 2:

[0308] The server converts the received audio data into text data using the Google Cloud Speech-to-Text API. The input is the audio data, and the output is the corresponding text data. Through this conversion, the audio information becomes processable as text.

[0309] Step 3:

[0310] The server analyzes the converted text data using natural language processing techniques to identify the emotional state in communication. The input is the text data, and the output is the category of the emotional state (e.g., positive, negative, neutral). Through this, data processing for understanding the emotion behind the text is performed.

[0311] Step 4:

[0312] The server also analyzes the emotional state from the video data using Amazon Rekognition. The input is the video data, and the output is the evaluation result of the emotional state based on the video. Through this process, it becomes possible to grasp people's emotions through visual data.

[0313] Step 5:

[0314] The server generates a prompt sentence to be given to the generative AI model based on these emotion analysis results. The prompt sentence is a guiding sentence for obtaining an appropriate response according to the analyzed emotion. The input is the evaluation result of the emotional state, and the output is the prompt sentence.

[0315] Step 6:

[0316] The server uses a generative AI model to generate appropriate instruction messages based on the prompt. The input is the prompt, and the output is the generated instruction message. This message is contextualized by the generative AI model.

[0317] Step 7:

[0318] The terminal transmits the generated instruction message to people on-site as actual voice or text via a speaker or monitor. This ensures that messages that help ensure safety and promote reassurance are delivered. Input is the generated instruction message, and output is the transmitted audio or visual information.

[0319] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0320] This invention is implemented as a system that combines the analysis of communication information with an emotion engine that recognizes the user's own emotions, in order to facilitate smooth communication between the user and the other party. The system functions in cooperation with a terminal and a server to provide an optimal dialogue environment for the user.

[0321] The device is responsible for collecting message and voice data sent and received from the user's communication tools. Furthermore, it analyzes biosignals (e.g., heart rate, voice tone) using an emotion engine to read the user's emotional state. Based on this input data, the emotion engine recognizes the user's current emotions and sends that information to the server.

[0322] The server first receives communication information sent from the terminal and analyzes this information using natural language processing technology. During the analysis, it identifies the other party's emotional state during the communication and understands the impression the user has of that party. The server also takes into account the user's emotional state, which is also transmitted from the terminal. Based on this information, impression choices to be presented to the user are generated, and these choices are optimized to reflect the user's emotions.

[0323] The user selects the impression they want to convey to the other person from a set of impression options provided by the server. This selection process is designed to reflect the user's own emotional state. Once the user completes their selection, the server uses a generative AI model to generate a natural-sounding reply message that takes the selected impression into account. The generated message is displayed to the user and can be customized.

[0324] Ultimately, the device sends a reply message to the recipient that has been approved by the user. This system helps users effectively convey their emotions while leaving a desirable impression on the recipient.

[0325] As a concrete example, consider a scenario where a user is asked by their partner, "How was your day?" If the data collected by the device indicates that the user is feeling "tired," the server provides the user with options such as "I want to relax" or "I want to be encouraged." If the user selects "I want to relax," the server generates and sends a message such as, "I was a little tired today, but it was great chatting!" In this way, the system of the present invention takes the user's own emotions into consideration and provides the necessary support in communication.

[0326] The following describes the processing flow.

[0327] Step 1:

[0328] The device collects text messages and voice data from the communication tools used by the user. It also measures biometric signals such as the user's voice tone and heart rate using an emotion engine, preparing to estimate the user's emotional state.

[0329] Step 2:

[0330] The device encrypts the collected communication information and the user's estimated emotional state when sending them to the server to protect the confidentiality of the data.

[0331] Step 3:

[0332] The server analyzes the communication information received from the terminal and uses natural language processing technology to identify the other party's emotional state and impression. At the same time, it receives the user's emotional state transmitted from the terminal and integrates the information from both.

[0333] Step 4:

[0334] The server generates a selection of impression options to present to the user, based on the user's emotions and the identified emotional state of the other party. These options are optimized to be consistent with the emotions the user is currently feeling.

[0335] Step 5:

[0336] The user selects the impression they want to convey to the other person from a set of impression options presented by the server. This selection corresponds to the user's own emotional state and facilitates appropriate dialogue.

[0337] Step 6:

[0338] The server uses a generative AI model to generate the optimal reply message based on the user's selected impression. This message is adjusted based on the user's emotional state and selected impression.

[0339] Step 7:

[0340] The terminal displays the generated reply message to the user, allowing the user to review the content and customize it as needed.

[0341] Step 8:

[0342] Once the user approves the message, the device sends the message to the recipient, completing the communication.

[0343] This process allows users to recognize their own emotional state and effectively communicate the impression they want to make to others.

[0344] (Example 2)

[0345] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0346] Conventional communication systems struggle to appropriately reflect user emotions and generate messages that create a desirable impression on the recipient. This problem hinders efficient and natural dialogue, especially in communication where emotions are complexly intertwined. Therefore, there is a need for technology that simultaneously considers the user's emotions and the recipient's state to generate the optimal response.

[0347] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0348] In this invention, the server includes means for collecting communication data, means for analyzing the communication data to identify the emotional state of the other party, and means for analyzing the user's biosignals to recognize emotions. This makes it possible to generate a response message that gives the most appropriate impression to the other party while taking the user's emotions into consideration.

[0349] "Communication data" refers to all information transmitted and received through electronic means, and specifically includes text messages and voice data.

[0350] "The other party's emotional state" refers to a description of the other party's emotions and mental state, analyzed from communication data.

[0351] "Biosignals" refer to data that indicates the user's physical state, including physical measurements such as heart rate and voice tone.

[0352] "Impression options" refer to multiple choices presented to a user to select the impression they want to convey to the other party.

[0353] A "response message" refers to a message that facilitates natural conversation, generated with consideration for the emotional state of both the user and the other party.

[0354] A "user interface" refers to an interactive mechanism that allows a user to interact with a computer system and perform various operations.

[0355] This invention is a system designed to facilitate communication between a user and another party. It utilizes communication data and biosignals to analyze emotions and generate optimal response messages. The roles of each component of this system are described below.

[0356] The device serves as a communication tool used by users on a daily basis and plays a role in collecting text messages and voice data. This device, such as a smartphone or PC, acquires this data in real time through dedicated applications and APIs, and also collects biometric signals such as heart rate and voice tone. Wearable devices and dedicated sensors are used to collect these biometric signals.

[0357] The collected data is analyzed by a server. The server possesses powerful processing capabilities and utilizes natural language processing techniques and machine learning algorithms to analyze the communication data. In the future, models such as BERT (Bidirectional Encoder Representations from Transformers) and LSTM (Long Short-Term Memory) may be used. Based on the emotional state of the other party and the user's emotions obtained from the emotion analysis, the server generates impression options for the user.

[0358] The user can review the impression options suggested by the server and select the one they deem most effective. Based on this selection, the server uses a generative AI model to create a natural and harmonious response message. This response message is presented to the user through the user interface and can be customized as needed.

[0359] For example, if a user is asked "How was it?" by their partner, and the device recognizes the user's emotion as "tired," the server will present impressions such as "I want to relax" or "I want to be encouraged." If the user selects "I want to relax," the server will generate a message such as "I'm a little tired today, but it was great chatting!" This process helps users effectively convey their emotions while maintaining natural communication.

[0360] The AI ​​generation model is input with a prompt such as, "If the user's emotional state is fatigued, they have selected 'I want to relax' as the impression they want to give. Based on this, please generate a natural reply message." In this way, the system provides an accurate and flexible conversational experience.

[0361] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0362] Step 1:

[0363] The device collects text messages and voice data in real time from the user's communication tools. Furthermore, it measures the user's heart rate and voice tone through biometric sensors. This data is acquired as input necessary for recognizing emotional states.

[0364] Step 2:

[0365] Inside the device, an emotion engine analyzes the user's emotions based on biosignal data. This analysis uses a machine learning model, with heart rate and voice tone fluctuations serving as foundational data for inferring the user's emotions. The analysis results are output as the user's emotional state and sent to the server.

[0366] Step 3:

[0367] The server receives text messages sent from terminals and analyzes their content using natural language processing (NLP) techniques. Specifically, it performs processes such as text tokenization, sentiment analysis, and keyword extraction to identify the recipient's emotional state. This analysis result serves as foundational data for evaluating the impression a user has of another person.

[0368] Step 4:

[0369] The server generates impression options to present to the user based on the analyzed emotional states of the user and the other party. In this process, optimal options are provided based on past data and statistical models, taking into account the user's communication style and situation. The generated options are then sent from the server to the user.

[0370] Step 5:

[0371] The user reviews the impression options provided by the server and selects the impression they wish to convey. The user's selection is intuitively made through a GUI, and this selection significantly influences the generation of the next response message.

[0372] Step 6:

[0373] The server generates a natural response message using a generative AI model based on the impression selected by the user. In this generation process, the prompt "If the user's emotional state is fatigued, they have selected 'I want to relax' as the impression they want to give to the other party. Based on this, please generate a natural reply message." is used as input data, and the AI ​​model outputs optimized text.

[0374] Step 7:

[0375] The terminal displays the response message received from the server to the user, providing an interface that allows the user to customize the content as needed. The user then makes a final confirmation of the message.

[0376] Step 8:

[0377] The device sends a response message to the other party, acknowledging the user's input. This completes the communication and effectively conveys the user's emotions.

[0378] (Application Example 2)

[0379] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0380] In online communication, it is necessary to appropriately understand users' emotions and provide support for making optimal responses and product suggestions based on those emotions. Conventional systems lack the ability to provide individualized support that takes into account the user's emotional state, resulting in a decline in the quality of dialogue. Furthermore, the lack of systems that offer product suggestions tailored to the user's emotions means that appropriate purchasing support is not being provided.

[0381] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0382] In this invention, the server includes means for collecting communication information, means for analyzing the collected communication information and identifying the emotional state of the other party, and means for generating selectable impression options and product suggestions to present to the user based on the identified emotional state. This enables communication tailored to the user's emotions and optimal product suggestions.

[0383] "Communication information" refers to information such as messages and voice data exchanged between a user and another party.

[0384] "Emotional state" refers to the state of mind and mood of the user or the other party, and is an indicator used to consider the impact on communication by understanding the changes in this state.

[0385] An "impression option" is a set of choices generated to allow a user to select what kind of impression they want to give to the other person.

[0386] A "reply message" is a response that a user sends to another person, optimized based on their emotional state and impression choices.

[0387] A "user interface" refers to the screens and means through which a user operates the system, and to review and customize generated reply messages and product recommendations.

[0388] "Product recommendation" refers to the presentation of product information by suggesting the most suitable product based on the user's emotions and choices.

[0389] "Natural language processing technology" is a general term for technologies used to enable computers to understand and process human language.

[0390] The system for carrying out this invention includes a terminal used by the user and a server that processes information. First, the terminal collects the user's voice data and text messages. This information is analyzed by an emotion engine to identify the user's emotional state. The emotion engine uses biosignals such as the user's heart rate and voice tone as input data to accurately recognize emotions.

[0391] The server receives communication information sent from the terminal and analyzes the information using natural language processing technology. As a result of the analysis, the server identifies the other party's emotional state and understands the impression the user has. Based on this information, the server generates selectable impression options and product recommendations. When the user selects an impression or product from the provided options, the server uses a generative AI model to generate a natural response message corresponding to the selected impression.

[0392] Users can review the generated messages and product suggestions displayed on their device and customize them as needed. The device sends a reply message to the recipient after receiving user approval. One possible application of this system is user support during product purchases in a virtual store. For example, if a user expresses a desire to "relax," the server might suggest a product such as, "This candle promotes relaxation."

[0393] Through this system, users can receive optimal communication and product suggestions that take their emotions into consideration, improving the quality of the interaction. An example of a prompt message is, "Generate a message that provides optimal product suggestions based on the emotion analysis results."

[0394] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0395] Step 1:

[0396] The device collects user voice data and text messages. This data is sent to an emotion engine, which analyzes it along with biosignals (e.g., heart rate, voice tone) necessary to recognize the user's emotional state. The input is the user's raw communication data, and the output is the user's emotional state as recognized by the emotion engine. Machine learning algorithms are applied based on the data to recognize the emotional state.

[0397] Step 2:

[0398] The server receives the user's emotional state and communication information transmitted from the terminal. Based on the received data, natural language processing techniques are used to analyze the other party's emotional state and identify the emotional relationship between the two parties. The input for this step is the user's emotional state and communication information, and the output is the other party's emotional state and what they understood it to be. Text mining and sentiment analysis techniques are used for the analysis.

[0399] Step 3:

[0400] The server generates selectable impression options and product suggestions based on the analysis results. This is done using an AI model that generates options that correspond to the user's emotional state and the other party's emotions. The input is the analysis results from the previous stage, and the output is the impression options and product suggestions. Text generation techniques, including topic modeling, are used to generate the impression options and product suggestions.

[0401] Step 4:

[0402] The user reviews impression options and product suggestions via their device and makes a selection. The user's selection is sent to the server, which uses a generative AI model to generate a natural-sounding reply message based on this selection. The input is the user's selection, and the output is a natural-sounding reply message. Advanced natural language generation techniques, such as transformer models, are used for generation.

[0403] Step 5:

[0404] The terminal displays the generated message to the user, and after the user's final confirmation or customization, sends the approved reply message to the recipient. At this stage, the input consists of the reply message and the user's customizations, while the output is the final response message sent. The sending function utilizes standard communication protocols for data transmission.

[0405] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0406] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0407] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0408] [Third Embodiment]

[0409] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0410] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0411] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0412] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0413] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0414] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0415] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0416] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0417] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0418] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0419] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0420] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0421] This invention is implemented as a system that analyzes the emotions and impressions of others through communication tools used by users on a daily basis and generates the optimal response. The system primarily functions between a terminal and a server.

[0422] The device collects data from various communication methods used by the user, such as LINE, social media, email, and voice calls. This device monitors the content of messages and calls sent and received in real time or at regular intervals, encrypts the collected communication information for privacy protection, and sends it to a server. During this process, voice data is also converted to text.

[0423] The server receives data sent from the terminal and analyzes it. This analysis utilizes emotion analysis AI and natural language processing technology. This allows the server to identify the other party's emotional state and understand the impression gained from the conversation. Based on the analysis results, the server has the function of presenting the user with multiple impression options to choose from.

[0424] Next, the user selects the impression they wish to convey from the presented impression options. Once the user has made their selection, the server uses a generative AI model to generate a natural-sounding reply message based on the selected impression. In this process, contextual language generation is performed to ensure the message is natural and appropriate.

[0425] Finally, the device presents the generated reply message to the user, obtains the user's confirmation, and then sends it to the recipient. This entire process allows the user to efficiently convey the desired impression to the recipient, resulting in better communication.

[0426] For example, consider a scenario where a user receives a message from a friend asking, "How are you doing lately?" In this case, the device collects the message, and the server analyzes it. If the analysis detects that the friend is "concerned," the user is presented with options to respond, such as "reassure" or "cheer up." If the user selects "reassure," a corresponding reply such as "I'm fine, how about you?" is generated and sent. In this way, the system of the present invention helps users communicate effectively.

[0427] The following describes the processing flow.

[0428] Step 1:

[0429] The device monitors messages and voice calls sent and received from the communication tools used by the user, and collects necessary communication information in real time. For voice calls, it also performs the process of converting speech to text.

[0430] Step 2:

[0431] The terminal encrypts the collected communication information and implements security measures to prevent data leakage before sending the data to the server.

[0432] Step 3:

[0433] The server analyzes the communication information received from the terminal. Using sentiment analysis AI, it identifies the recipient's emotional state from the message, and uses natural language processing technology to understand the impression the recipient has of the user.

[0434] Step 4:

[0435] The server presents the user with multiple impression options based on the analysis results. These options include choices related to the other person's feelings (e.g., reassurance, encouragement).

[0436] Step 5:

[0437] The user selects the impression they want to give to the other person from a list of impression options presented by the server.

[0438] Step 6:

[0439] The server generates the optimal reply message using a generative AI model based on the user's selection. This reply message is then adjusted to best match the selected impression.

[0440] Step 7:

[0441] The device presents the generated reply message to the user and obtains the user's confirmation and final approval.

[0442] Step 8:

[0443] After the user reviews and approves the reply message, the device sends the message to the recipient.

[0444] Through this series of steps, users can effectively communicate with others and leave the desired impression.

[0445] (Example 1)

[0446] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0447] In modern digital communication, users often struggle to properly understand the other person's emotions and formulate appropriate responses. This problem is particularly pronounced in text-only communication, where the emotional nuances of the other person cannot be accurately captured. As a result, misunderstandings and inefficient communication can occur, negatively impacting the quality of relationships.

[0448] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0449] In this invention, the server includes a device for collecting communication data, a device for analyzing the collected communication data and identifying the emotional state of another person, and a device for creating selectable impression options to display to the user based on the identified emotional state. This enables the user to efficiently understand the other person's emotional state and generate a response message for optimal communication.

[0450] "Communication data" refers to all digital data exchanged between people, such as messages and voice information, that are sent and received through electronic means.

[0451] "Device" refers to hardware, software, or a combination thereof designed to perform a specific function.

[0452] "Analysis" refers to the process of extracting specific information from collected data and using that information to gain new insights.

[0453] "Emotional state" refers to the type and intensity of emotions that the communication partner is presumed to be experiencing, and is often revealed through text analysis.

[0454] "Users" refer to people who use this system to communicate.

[0455] "Selectable impression options" refers to multiple choices presented to the user to select the impression they want to convey to the other person.

[0456] A "response message" refers to a reply generated based on the user's intent, and its content is appropriate to the recipient's intentions and feelings.

[0457] This invention is a system aimed at enabling users to understand the emotions of others through digital communication and to create appropriate response messages. This system primarily operates between a device held by the user and a server located in the cloud.

[0458] The device collects communication data from communication tools that users use daily, such as messaging applications, social networking services, email, and voice calls. The collected data is protected using standard encryption technologies such as AES and RSA to ensure security and privacy. Voice data is converted to text using real-time speech recognition software.

[0459] When the server receives encrypted communication data sent from a terminal, it decrypts it and analyzes the data using sentiment analysis AI and natural language processing techniques. In this analysis process, machine learning libraries (e.g., TensorFlow and PyTorch) are used to identify the other party's emotional state and tone of communication.

[0460] Based on the analysis results, the server presents the user with multiple selectable impression options. These options are created based on predefined templates and AI-generated suggestions, and are important for improving the user experience of the product or service.

[0461] The generative AI model used in the analysis process receives prompts based on the user's selection and generates natural, contextually appropriate responses. For example, if a user receives a message from a friend saying "How are you?", they might instruct the model with the prompt, "My friend messaged me 'How are you?' Please generate a reassuring response." In this way, the generated response messages help users communicate quickly and accurately with others.

[0462] Finally, the generated message is presented to the user on their device, where it can be modified and reviewed as needed before being sent to the recipient. This entire process allows users to achieve smoother and more effective communication.

[0463] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0464] Step 1:

[0465] The device collects communication data from various communication methods used by the user. Inputs include messages and voice data. Specifically, the device periodically acquires text messages and voice communication data through APIs and interfaces, and converts voice data to text using speech recognition software. The output is encrypted text data.

[0466] Step 2:

[0467] The terminal securely encrypts the collected communication data and sends it to the server. The input is the text data obtained in step 1. Specifically, the data is protected using encryption technologies such as AES and RSA and uploaded to the server via the internet. The output is the encrypted communication data received by the server.

[0468] Step 3:

[0469] The server decrypts encrypted communication data sent from the terminal and analyzes it using sentiment analysis AI and natural language processing technology. The input is encrypted communication data. Specifically, machine learning algorithms are used to analyze the emotional state and important keywords in the text and identify the emotional state of others. The output is the analyzed emotional state and communication tone information.

[0470] Step 4:

[0471] The server generates and provides the user with selectable impression options based on the analysis results. The input is the emotional state information obtained in step 3. Specifically, it generates impression options (e.g., "reassure," "encourage," etc.) using AI or a predefined rule set and sends them to the terminal. The output is the impression options presented to the user.

[0472] Step 5:

[0473] The user selects their desired impression from the impression options presented by the server. The input is the impression options received from the server. Specifically, the user makes a selection by tapping or clicking on an option, and this selection information is sent to the server. The output is the selected impression information.

[0474] Step 6:

[0475] The server utilizes a generative AI model based on the user's selected impression to generate an appropriate response message. The input is the user's selected impression information. Specifically, the generative AI model is given a prompt, and a process is executed to generate a natural response that takes the context into account. The output is the generated response message.

[0476] Step 7:

[0477] The terminal receives a response message generated from the server and presents it to the user. The input is the generated response message. Specifically, it displays a UI that allows the user to review the response and make corrections if necessary. After obtaining user approval, it sends the response message to the other party. The output is the final response message that is sent.

[0478] (Application Example 1)

[0479] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0480] In the security field in particular, immediate responses tailored to the situation on-site are often required. However, conventional systems have a problem in that communication after anomaly detection relies on manual processes, making it difficult to respond quickly and accurately.

[0481] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0482] In this invention, the server includes means for analyzing communication information and identifying the emotional state of the other party, means for detecting anomalies from the surrounding environment and communicating with people on site in response to those anomalies, and means for analyzing people's emotions in real time and generating appropriate instruction messages. This enables emotionally appropriate communication and rapid response tailored to the situation on site.

[0483] "Communication information" refers to digital data that is sent and received between a user and others in the form of voice data, text messages, image data, etc.

[0484] "Emotional state" refers to the psychological reactions and emotional shifts of the other person obtained through communication, and is usually expressed in categories such as positive, negative, and neutral.

[0485] "Impression options" are choices that allow users to select what kind of impression they want to give to the other person, and are presented to the user based on an analysis of their emotional state.

[0486] A "reply message" is a message generated based on the user's selections, and its purpose is to create the best possible impression on the recipient.

[0487] "Detecting anomalies from the surrounding environment" means discovering unusual situations or events by analyzing environmental data such as sound, video, and vibration.

[0488] "Communicating with people on-site" means communicating with people at the scene via voice or text when an anomaly is detected, and sharing necessary information.

[0489] "Analyzing people's emotions in real time" is a process of instantly processing communication data and understanding their emotional state as a result.

[0490] An "appropriate instruction message" is an instruction or guidance message generated based on an analyzed emotional state, and its content is designed to provide appropriate instructions according to the situation on site.

[0491] This invention relates to a system applied in the security field and provides technology for effective communication with people on site in response to abnormal situations within a building.

[0492] The server utilizes a speech recognition module to convert audio data collected from the field into text data. Specifically, it uses the Google Cloud Speech-to-Text API to convert speech to text in real time. The transcribed information is then analyzed using natural language processing techniques to identify emotional states.

[0493] The server analyzes emotions from video data using Amazon Rekognition and other tools via an emotion analysis module. This allows for the evaluation of the psychological state of people on-site and their individual reactions to abnormal situations. Based on the results of the real-time emotion analysis, prompt sentences are supplied to an AI model to obtain the optimal response.

[0494] The AI ​​model generates appropriate instructional messages based on the user's selected impressions, responding to any unusual situation. These messages are transmitted to people on-site via speakers or monitors. For example, if a suspicious sound is detected during a night patrol and workers are on alert, the server will generate and send a message stating, "Security check in progress. Please remain calm."

[0495] As an example of a prompt, we will use a prompt in the format of, "Create a reassuring message for employees who are feeling anxious." Based on this prompt, the generative AI model will provide flexible instructions that are relevant to the context.

[0496] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0497] Step 1:

[0498] The device collects audio and video data from within the building. Audio data is acquired via a microphone, and video data is acquired via a camera. This raw data is transmitted to the server in real time.

[0499] Step 2:

[0500] The server converts the received audio data into text data using the Google Cloud Speech-to-Text API. The input is audio data, and the output is the corresponding text data. This conversion makes the audio information processable as text.

[0501] Step 3:

[0502] The server analyzes the converted text data using natural language processing techniques to identify the emotional state in the communication. The input is text data, and the output is a category of emotional state (e.g., positive, negative, neutral). This enables data processing that understands the emotions behind the text.

[0503] Step 4:

[0504] The server uses Amazon Rekognition to analyze emotional states from video data. The input is video data, and the output is an evaluation of emotional states based on the video. This process makes it possible to understand people's emotions through visual data.

[0505] Step 5:

[0506] Based on these sentiment analysis results, the server generates prompt sentences to give to the generating AI model. These prompt sentences are guiding sentences designed to elicit an appropriate response corresponding to the analyzed sentiment. The input is the evaluation result of the emotional state, and the output is the prompt sentence.

[0507] Step 6:

[0508] The server uses a generative AI model to generate appropriate instruction messages based on the prompt. The input is the prompt, and the output is the generated instruction message. This message is contextualized by the generative AI model.

[0509] Step 7:

[0510] The terminal transmits the generated instruction message to people on-site as actual voice or text via a speaker or monitor. This ensures that messages that help ensure safety and promote reassurance are delivered. Input is the generated instruction message, and output is the transmitted audio or visual information.

[0511] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0512] This invention is implemented as a system that combines the analysis of communication information with an emotion engine that recognizes the user's own emotions, in order to facilitate smooth communication between the user and the other party. The system functions in cooperation with a terminal and a server to provide an optimal dialogue environment for the user.

[0513] The device is responsible for collecting message and voice data sent and received from the user's communication tools. Furthermore, it analyzes biosignals (e.g., heart rate, voice tone) using an emotion engine to read the user's emotional state. Based on this input data, the emotion engine recognizes the user's current emotions and sends that information to the server.

[0514] The server first receives communication information sent from the terminal and analyzes this information using natural language processing technology. During the analysis, it identifies the other party's emotional state during the communication and understands the impression the user has of that party. The server also takes into account the user's emotional state, which is also transmitted from the terminal. Based on this information, impression choices to be presented to the user are generated, and these choices are optimized to reflect the user's emotions.

[0515] The user selects the impression they want to convey to the other person from a set of impression options provided by the server. This selection process is designed to reflect the user's own emotional state. Once the user completes their selection, the server uses a generative AI model to generate a natural-sounding reply message that takes the selected impression into account. The generated message is displayed to the user and can be customized.

[0516] Ultimately, the device sends a reply message to the recipient that has been approved by the user. This system helps users effectively convey their emotions while leaving a desirable impression on the recipient.

[0517] As a concrete example, consider a scenario where a user is asked by their partner, "How was your day?" If the data collected by the device indicates that the user is feeling "tired," the server provides the user with options such as "I want to relax" or "I want to be encouraged." If the user selects "I want to relax," the server generates and sends a message such as, "I was a little tired today, but it was great chatting!" In this way, the system of the present invention takes the user's own emotions into consideration and provides the necessary support in communication.

[0518] The following describes the processing flow.

[0519] Step 1:

[0520] The device collects text messages and voice data from the communication tools used by the user. It also measures biometric signals such as the user's voice tone and heart rate using an emotion engine, preparing to estimate the user's emotional state.

[0521] Step 2:

[0522] The device encrypts the collected communication information and the user's estimated emotional state when sending them to the server to protect the confidentiality of the data.

[0523] Step 3:

[0524] The server analyzes the communication information received from the terminal and uses natural language processing technology to identify the other party's emotional state and impression. At the same time, it receives the user's emotional state transmitted from the terminal and integrates the information from both.

[0525] Step 4:

[0526] The server generates a selection of impression options to present to the user, based on the user's emotions and the identified emotional state of the other party. These options are optimized to be consistent with the emotions the user is currently feeling.

[0527] Step 5:

[0528] The user selects the impression they want to convey to the other person from a set of impression options presented by the server. This selection corresponds to the user's own emotional state and facilitates appropriate dialogue.

[0529] Step 6:

[0530] The server uses a generative AI model to generate the optimal reply message based on the user's selected impression. This message is adjusted based on the user's emotional state and selected impression.

[0531] Step 7:

[0532] The terminal displays the generated reply message to the user, allowing the user to review the content and customize it as needed.

[0533] Step 8:

[0534] Once the user approves the message, the device sends the message to the recipient, completing the communication.

[0535] This process allows users to recognize their own emotional state and effectively communicate the impression they want to make to others.

[0536] (Example 2)

[0537] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0538] Conventional communication systems struggle to appropriately reflect user emotions and generate messages that create a desirable impression on the recipient. This problem hinders efficient and natural dialogue, especially in communication where emotions are complexly intertwined. Therefore, there is a need for technology that simultaneously considers the user's emotions and the recipient's state to generate the optimal response.

[0539] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0540] In this invention, the server includes means for collecting communication data, means for analyzing the communication data to identify the emotional state of the other party, and means for analyzing the user's biosignals to recognize emotions. This makes it possible to generate a response message that gives the most appropriate impression to the other party while taking the user's emotions into consideration.

[0541] "Communication data" refers to all information transmitted and received through electronic means, and specifically includes text messages and voice data.

[0542] "The other party's emotional state" refers to a description of the other party's emotions and mental state, analyzed from communication data.

[0543] "Biosignals" refer to data that indicates the user's physical state, including physical measurements such as heart rate and voice tone.

[0544] "Impression options" refer to multiple choices presented to a user to select the impression they want to convey to the other party.

[0545] A "response message" refers to a message that facilitates natural conversation, generated with consideration for the emotional state of both the user and the other party.

[0546] A "user interface" refers to an interactive mechanism that allows a user to interact with a computer system and perform various operations.

[0547] This invention is a system designed to facilitate communication between a user and another party. It utilizes communication data and biosignals to analyze emotions and generate optimal response messages. The roles of each component of this system are described below.

[0548] The device serves as a communication tool used by users on a daily basis and plays a role in collecting text messages and voice data. This device, such as a smartphone or PC, acquires this data in real time through dedicated applications and APIs, and also collects biometric signals such as heart rate and voice tone. Wearable devices and dedicated sensors are used to collect these biometric signals.

[0549] The collected data is analyzed by a server. The server possesses powerful processing capabilities and utilizes natural language processing techniques and machine learning algorithms to analyze the communication data. In the future, models such as BERT (Bidirectional Encoder Representations from Transformers) and LSTM (Long Short-Term Memory) may be used. Based on the emotional state of the other party and the user's emotions obtained from the emotion analysis, the server generates impression options for the user.

[0550] The user can review the impression options suggested by the server and select the one they deem most effective. Based on this selection, the server uses a generative AI model to create a natural and harmonious response message. This response message is presented to the user through the user interface and can be customized as needed.

[0551] For example, if a user is asked "How was it?" by their partner, and the device recognizes the user's emotion as "tired," the server will present impressions such as "I want to relax" or "I want to be encouraged." If the user selects "I want to relax," the server will generate a message such as "I'm a little tired today, but it was great chatting!" This process helps users effectively convey their emotions while maintaining natural communication.

[0552] The AI ​​generation model is input with a prompt such as, "If the user's emotional state is fatigued, they have selected 'I want to relax' as the impression they want to give. Based on this, please generate a natural reply message." In this way, the system provides an accurate and flexible conversational experience.

[0553] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0554] Step 1:

[0555] The device collects text messages and voice data in real time from the user's communication tools. Furthermore, it measures the user's heart rate and voice tone through biometric sensors. This data is acquired as input necessary for recognizing emotional states.

[0556] Step 2:

[0557] Inside the device, an emotion engine analyzes the user's emotions based on biosignal data. This analysis uses a machine learning model, with heart rate and voice tone fluctuations serving as foundational data for inferring the user's emotions. The analysis results are output as the user's emotional state and sent to the server.

[0558] Step 3:

[0559] The server receives text messages sent from terminals and analyzes their content using natural language processing (NLP) techniques. Specifically, it performs processes such as text tokenization, sentiment analysis, and keyword extraction to identify the recipient's emotional state. This analysis result serves as foundational data for evaluating the impression a user has of another person.

[0560] Step 4:

[0561] The server generates impression options to present to the user based on the analyzed emotional states of the user and the other party. In this process, optimal options are provided based on past data and statistical models, taking into account the user's communication style and situation. The generated options are then sent from the server to the user.

[0562] Step 5:

[0563] The user reviews the impression options provided by the server and selects the impression they wish to convey. The user's selection is intuitively made through a GUI, and this selection significantly influences the generation of the next response message.

[0564] Step 6:

[0565] The server generates a natural response message using a generative AI model based on the impression selected by the user. In this generation process, the prompt "If the user's emotional state is fatigued, they have selected 'I want to relax' as the impression they want to give to the other party. Based on this, please generate a natural reply message." is used as input data, and the AI ​​model outputs optimized text.

[0566] Step 7:

[0567] The terminal displays the response message received from the server to the user, providing an interface that allows the user to customize the content as needed. The user then makes a final confirmation of the message.

[0568] Step 8:

[0569] The device sends a response message to the other party, acknowledging the user's input. This completes the communication and effectively conveys the user's emotions.

[0570] (Application Example 2)

[0571] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0572] In online communication, it is necessary to appropriately understand users' emotions and provide support for making optimal responses and product suggestions based on those emotions. Conventional systems lack the ability to provide individualized support that takes into account the user's emotional state, resulting in a decline in the quality of dialogue. Furthermore, the lack of systems that offer product suggestions tailored to the user's emotions means that appropriate purchasing support is not being provided.

[0573] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0574] In this invention, the server includes means for collecting communication information, means for analyzing the collected communication information and identifying the emotional state of the other party, and means for generating selectable impression options and product suggestions to present to the user based on the identified emotional state. This enables communication tailored to the user's emotions and optimal product suggestions.

[0575] "Communication information" refers to information such as messages and voice data exchanged between a user and another party.

[0576] "Emotional state" refers to the state of mind and mood of the user or the other party, and is an indicator used to consider the impact on communication by understanding the changes in this state.

[0577] An "impression option" is a set of choices generated to allow a user to select what kind of impression they want to give to the other person.

[0578] A "reply message" is a response that a user sends to another person, optimized based on their emotional state and impression choices.

[0579] A "user interface" refers to the screens and means through which a user operates the system, and to review and customize generated reply messages and product recommendations.

[0580] "Product recommendation" refers to the presentation of product information by suggesting the most suitable product based on the user's emotions and choices.

[0581] "Natural language processing technology" is a general term for technologies used to enable computers to understand and process human language.

[0582] The system for carrying out this invention includes a terminal used by the user and a server that processes information. First, the terminal collects the user's voice data and text messages. This information is analyzed by an emotion engine to identify the user's emotional state. The emotion engine uses biosignals such as the user's heart rate and voice tone as input data to accurately recognize emotions.

[0583] The server receives communication information sent from the terminal and analyzes the information using natural language processing technology. As a result of the analysis, the server identifies the other party's emotional state and understands the impression the user has. Based on this information, the server generates selectable impression options and product recommendations. When the user selects an impression or product from the provided options, the server uses a generative AI model to generate a natural response message corresponding to the selected impression.

[0584] Users can review the generated messages and product suggestions displayed on their device and customize them as needed. The device sends a reply message to the recipient after receiving user approval. One possible application of this system is user support during product purchases in a virtual store. For example, if a user expresses a desire to "relax," the server might suggest a product such as, "This candle promotes relaxation."

[0585] Through this system, users can receive optimal communication and product suggestions that take their emotions into consideration, improving the quality of the interaction. An example of a prompt message is, "Generate a message that provides optimal product suggestions based on the emotion analysis results."

[0586] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0587] Step 1:

[0588] The device collects user voice data and text messages. This data is sent to an emotion engine, which analyzes it along with biosignals (e.g., heart rate, voice tone) necessary to recognize the user's emotional state. The input is the user's raw communication data, and the output is the user's emotional state as recognized by the emotion engine. Machine learning algorithms are applied based on the data to recognize the emotional state.

[0589] Step 2:

[0590] The server receives the user's emotional state and communication information transmitted from the terminal. Based on the received data, natural language processing techniques are used to analyze the other party's emotional state and identify the emotional relationship between the two parties. The input for this step is the user's emotional state and communication information, and the output is the other party's emotional state and what they understood it to be. Text mining and sentiment analysis techniques are used for the analysis.

[0591] Step 3:

[0592] The server generates selectable impression options and product suggestions based on the analysis results. This is done using an AI model that generates options that correspond to the user's emotional state and the other party's emotions. The input is the analysis results from the previous stage, and the output is the impression options and product suggestions. Text generation techniques, including topic modeling, are used to generate the impression options and product suggestions.

[0593] Step 4:

[0594] The user reviews impression options and product suggestions via their device and makes a selection. The user's selection is sent to the server, which uses a generative AI model to generate a natural-sounding reply message based on this selection. The input is the user's selection, and the output is a natural-sounding reply message. Advanced natural language generation techniques, such as transformer models, are used for generation.

[0595] Step 5:

[0596] The terminal displays the generated message to the user, and after the user's final confirmation or customization, sends the approved reply message to the recipient. At this stage, the input consists of the reply message and the user's customizations, while the output is the final response message sent. The sending function utilizes standard communication protocols for data transmission.

[0597] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0598] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0599] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0600] [Fourth Embodiment]

[0601] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0602] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0603] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0604] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0605] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0606] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0607] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0608] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0609] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0610] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0611] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0612] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0613] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0614] This invention is implemented as a system that analyzes the emotions and impressions of others through communication tools used by users on a daily basis and generates the optimal response. The system primarily functions between a terminal and a server.

[0615] The device collects data from various communication methods used by the user, such as LINE, social media, email, and voice calls. This device monitors the content of messages and calls sent and received in real time or at regular intervals, encrypts the collected communication information for privacy protection, and sends it to a server. During this process, voice data is also converted to text.

[0616] The server receives data sent from the terminal and analyzes it. This analysis utilizes emotion analysis AI and natural language processing technology. This allows the server to identify the other party's emotional state and understand the impression gained from the conversation. Based on the analysis results, the server has the function of presenting the user with multiple impression options to choose from.

[0617] Next, the user selects the impression they wish to convey from the presented impression options. Once the user has made their selection, the server uses a generative AI model to generate a natural-sounding reply message based on the selected impression. In this process, contextual language generation is performed to ensure the message is natural and appropriate.

[0618] Finally, the device presents the generated reply message to the user, obtains the user's confirmation, and then sends it to the recipient. This entire process allows the user to efficiently convey the desired impression to the recipient, resulting in better communication.

[0619] For example, consider a scenario where a user receives a message from a friend asking, "How are you doing lately?" In this case, the device collects the message, and the server analyzes it. If the analysis detects that the friend is "concerned," the user is presented with options to respond, such as "reassure" or "cheer up." If the user selects "reassure," a corresponding reply such as "I'm fine, how about you?" is generated and sent. In this way, the system of the present invention helps users communicate effectively.

[0620] The following describes the processing flow.

[0621] Step 1:

[0622] The device monitors messages and voice calls sent and received from the communication tools used by the user, and collects necessary communication information in real time. For voice calls, it also performs the process of converting speech to text.

[0623] Step 2:

[0624] The terminal encrypts the collected communication information and implements security measures to prevent data leakage before sending the data to the server.

[0625] Step 3:

[0626] The server analyzes the communication information received from the terminal. Using sentiment analysis AI, it identifies the recipient's emotional state from the message, and uses natural language processing technology to understand the impression the recipient has of the user.

[0627] Step 4:

[0628] The server presents the user with multiple impression options based on the analysis results. These options include choices related to the other person's feelings (e.g., reassurance, encouragement).

[0629] Step 5:

[0630] The user selects the impression they want to give to the other person from a list of impression options presented by the server.

[0631] Step 6:

[0632] The server generates the optimal reply message using a generative AI model based on the user's selection. This reply message is then adjusted to best match the selected impression.

[0633] Step 7:

[0634] The device presents the generated reply message to the user and obtains the user's confirmation and final approval.

[0635] Step 8:

[0636] After the user reviews and approves the reply message, the device sends the message to the recipient.

[0637] Through this series of steps, users can effectively communicate with others and leave the desired impression.

[0638] (Example 1)

[0639] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0640] In modern digital communication, users often struggle to properly understand the other person's emotions and formulate appropriate responses. This problem is particularly pronounced in text-only communication, where the emotional nuances of the other person cannot be accurately captured. As a result, misunderstandings and inefficient communication can occur, negatively impacting the quality of relationships.

[0641] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0642] In this invention, the server includes a device for collecting communication data, a device for analyzing the collected communication data and identifying the emotional state of another person, and a device for creating selectable impression options to display to the user based on the identified emotional state. This enables the user to efficiently understand the other person's emotional state and generate a response message for optimal communication.

[0643] "Communication data" refers to all digital data exchanged between people, such as messages and voice information, that are sent and received through electronic means.

[0644] "Device" refers to hardware, software, or a combination thereof designed to perform a specific function.

[0645] "Analysis" refers to the process of extracting specific information from collected data and using that information to gain new insights.

[0646] "Emotional state" refers to the type and intensity of emotions that the communication partner is presumed to be experiencing, and is often revealed through text analysis.

[0647] "Users" refer to people who use this system to communicate.

[0648] "Selectable impression options" refers to multiple choices presented to the user to select the impression they want to convey to the other person.

[0649] A "response message" refers to a reply generated based on the user's intent, and its content is appropriate to the recipient's intentions and feelings.

[0650] This invention is a system aimed at enabling users to understand the emotions of others through digital communication and to create appropriate response messages. This system primarily operates between a device held by the user and a server located in the cloud.

[0651] The device collects communication data from communication tools that users use daily, such as messaging applications, social networking services, email, and voice calls. The collected data is protected using standard encryption technologies such as AES and RSA to ensure security and privacy. Voice data is converted to text using real-time speech recognition software.

[0652] When the server receives encrypted communication data sent from a terminal, it decrypts it and analyzes the data using sentiment analysis AI and natural language processing techniques. In this analysis process, machine learning libraries (e.g., TensorFlow and PyTorch) are used to identify the other party's emotional state and tone of communication.

[0653] Based on the analysis results, the server presents the user with multiple selectable impression options. These options are created based on predefined templates and AI-generated suggestions, and are important for improving the user experience of the product or service.

[0654] The generative AI model used in the analysis process receives prompts based on the user's selection and generates natural, contextually appropriate responses. For example, if a user receives a message from a friend saying "How are you?", they might instruct the model with the prompt, "My friend messaged me 'How are you?' Please generate a reassuring response." In this way, the generated response messages help users communicate quickly and accurately with others.

[0655] Finally, the generated message is presented to the user on their device, where it can be modified and reviewed as needed before being sent to the recipient. This entire process allows users to achieve smoother and more effective communication.

[0656] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0657] Step 1:

[0658] The device collects communication data from various communication methods used by the user. Inputs include messages and voice data. Specifically, the device periodically acquires text messages and voice communication data through APIs and interfaces, and converts voice data to text using speech recognition software. The output is encrypted text data.

[0659] Step 2:

[0660] The terminal securely encrypts the collected communication data and sends it to the server. The input is the text data obtained in step 1. Specifically, the data is protected using encryption technologies such as AES and RSA and uploaded to the server via the internet. The output is the encrypted communication data received by the server.

[0661] Step 3:

[0662] The server decrypts encrypted communication data sent from the terminal and analyzes it using sentiment analysis AI and natural language processing technology. The input is encrypted communication data. Specifically, machine learning algorithms are used to analyze the emotional state and important keywords in the text and identify the emotional state of others. The output is the analyzed emotional state and communication tone information.

[0663] Step 4:

[0664] The server generates and provides the user with selectable impression options based on the analysis results. The input is the emotional state information obtained in step 3. Specifically, it generates impression options (e.g., "reassure," "encourage," etc.) using AI or a predefined rule set and sends them to the terminal. The output is the impression options presented to the user.

[0665] Step 5:

[0666] The user selects their desired impression from the impression options presented by the server. The input is the impression options received from the server. Specifically, the user makes a selection by tapping or clicking on an option, and this selection information is sent to the server. The output is the selected impression information.

[0667] Step 6:

[0668] The server utilizes a generative AI model based on the user's selected impression to generate an appropriate response message. The input is the user's selected impression information. Specifically, the generative AI model is given a prompt, and a process is executed to generate a natural response that takes the context into account. The output is the generated response message.

[0669] Step 7:

[0670] The terminal receives a response message generated from the server and presents it to the user. The input is the generated response message. Specifically, it displays a UI that allows the user to review the response and make corrections if necessary. After obtaining user approval, it sends the response message to the other party. The output is the final response message that is sent.

[0671] (Application Example 1)

[0672] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0673] In the security field in particular, immediate responses tailored to the situation on-site are often required. However, conventional systems have a problem in that communication after anomaly detection relies on manual processes, making it difficult to respond quickly and accurately.

[0674] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0675] In this invention, the server includes means for analyzing communication information and identifying the emotional state of the other party, means for detecting anomalies from the surrounding environment and communicating with people on site in response to those anomalies, and means for analyzing people's emotions in real time and generating appropriate instruction messages. This enables emotionally appropriate communication and rapid response tailored to the situation on site.

[0676] "Communication information" refers to digital data that is sent and received between a user and others in the form of voice data, text messages, image data, etc.

[0677] "Emotional state" refers to the psychological reactions and emotional shifts of the other person obtained through communication, and is usually expressed in categories such as positive, negative, and neutral.

[0678] "Impression options" are choices that allow users to select what kind of impression they want to give to the other person, and are presented to the user based on an analysis of their emotional state.

[0679] A "reply message" is a message generated based on the user's selections, and its purpose is to create the best possible impression on the recipient.

[0680] "Detecting anomalies from the surrounding environment" means discovering unusual situations or events by analyzing environmental data such as sound, video, and vibration.

[0681] "Communicating with people on-site" means communicating with people at the scene via voice or text when an anomaly is detected, and sharing necessary information.

[0682] "Analyzing people's emotions in real time" is a process of instantly processing communication data and understanding their emotional state as a result.

[0683] An "appropriate instruction message" is an instruction or guidance message generated based on an analyzed emotional state, and its content is designed to provide appropriate instructions according to the situation on site.

[0684] This invention relates to a system applied in the security field and provides technology for effective communication with people on site in response to abnormal situations within a building.

[0685] The server utilizes a speech recognition module to convert audio data collected from the field into text data. Specifically, it uses the Google Cloud Speech-to-Text API to convert speech to text in real time. The transcribed information is then analyzed using natural language processing techniques to identify emotional states.

[0686] The server analyzes emotions from video data using Amazon Rekognition and other tools via an emotion analysis module. This allows for the evaluation of the psychological state of people on-site and their individual reactions to abnormal situations. Based on the results of the real-time emotion analysis, prompt sentences are supplied to an AI model to obtain the optimal response.

[0687] The AI ​​model generates appropriate instructional messages based on the user's selected impressions, responding to any unusual situation. These messages are transmitted to people on-site via speakers or monitors. For example, if a suspicious sound is detected during a night patrol and workers are on alert, the server will generate and send a message stating, "Security check in progress. Please remain calm."

[0688] As an example of a prompt, we will use a prompt in the format of, "Create a reassuring message for employees who are feeling anxious." Based on this prompt, the generative AI model will provide flexible instructions that are relevant to the context.

[0689] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0690] Step 1:

[0691] The device collects audio and video data from within the building. Audio data is acquired via a microphone, and video data is acquired via a camera. This raw data is transmitted to the server in real time.

[0692] Step 2:

[0693] The server converts the received audio data into text data using the Google Cloud Speech-to-Text API. The input is audio data, and the output is the corresponding text data. This conversion makes the audio information processable as text.

[0694] Step 3:

[0695] The server analyzes the converted text data using natural language processing techniques to identify the emotional state in the communication. The input is text data, and the output is a category of emotional state (e.g., positive, negative, neutral). This enables data processing that understands the emotions behind the text.

[0696] Step 4:

[0697] The server uses Amazon Rekognition to analyze emotional states from video data. The input is video data, and the output is an evaluation of emotional states based on the video. This process makes it possible to understand people's emotions through visual data.

[0698] Step 5:

[0699] Based on these sentiment analysis results, the server generates prompt sentences to give to the generating AI model. These prompt sentences are guiding sentences designed to elicit an appropriate response corresponding to the analyzed sentiment. The input is the evaluation result of the emotional state, and the output is the prompt sentence.

[0700] Step 6:

[0701] The server uses a generative AI model to generate appropriate instruction messages based on the prompt. The input is the prompt, and the output is the generated instruction message. This message is contextualized by the generative AI model.

[0702] Step 7:

[0703] The terminal transmits the generated instruction message to people on-site as actual voice or text via a speaker or monitor. This ensures that messages that help ensure safety and promote reassurance are delivered. Input is the generated instruction message, and output is the transmitted audio or visual information.

[0704] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0705] This invention is implemented as a system that combines the analysis of communication information with an emotion engine that recognizes the user's own emotions, in order to facilitate smooth communication between the user and the other party. The system functions in cooperation with a terminal and a server to provide an optimal dialogue environment for the user.

[0706] The device is responsible for collecting message and voice data sent and received from the user's communication tools. Furthermore, it analyzes biosignals (e.g., heart rate, voice tone) using an emotion engine to read the user's emotional state. Based on this input data, the emotion engine recognizes the user's current emotions and sends that information to the server.

[0707] The server first receives communication information sent from the terminal and analyzes this information using natural language processing technology. During the analysis, it identifies the other party's emotional state during the communication and understands the impression the user has of that party. The server also takes into account the user's emotional state, which is also transmitted from the terminal. Based on this information, impression choices to be presented to the user are generated, and these choices are optimized to reflect the user's emotions.

[0708] The user selects the impression they want to convey to the other person from a set of impression options provided by the server. This selection process is designed to reflect the user's own emotional state. Once the user completes their selection, the server uses a generative AI model to generate a natural-sounding reply message that takes the selected impression into account. The generated message is displayed to the user and can be customized.

[0709] Ultimately, the device sends a reply message to the recipient that has been approved by the user. This system helps users effectively convey their emotions while leaving a desirable impression on the recipient.

[0710] As a concrete example, consider a scenario where a user is asked by their partner, "How was your day?" If the data collected by the device indicates that the user is feeling "tired," the server provides the user with options such as "I want to relax" or "I want to be encouraged." If the user selects "I want to relax," the server generates and sends a message such as, "I was a little tired today, but it was great chatting!" In this way, the system of the present invention takes the user's own emotions into consideration and provides the necessary support in communication.

[0711] The following describes the processing flow.

[0712] Step 1:

[0713] The device collects text messages and voice data from the communication tools used by the user. It also measures biometric signals such as the user's voice tone and heart rate using an emotion engine, preparing to estimate the user's emotional state.

[0714] Step 2:

[0715] The device encrypts the collected communication information and the user's estimated emotional state when sending them to the server to protect the confidentiality of the data.

[0716] Step 3:

[0717] The server analyzes the communication information received from the terminal and uses natural language processing technology to identify the other party's emotional state and impression. At the same time, it receives the user's emotional state transmitted from the terminal and integrates the information from both.

[0718] Step 4:

[0719] The server generates a selection of impression options to present to the user, based on the user's emotions and the identified emotional state of the other party. These options are optimized to be consistent with the emotions the user is currently feeling.

[0720] Step 5:

[0721] The user selects the impression they want to convey to the other person from a set of impression options presented by the server. This selection corresponds to the user's own emotional state and facilitates appropriate dialogue.

[0722] Step 6:

[0723] The server uses a generative AI model to generate the optimal reply message based on the user's selected impression. This message is adjusted based on the user's emotional state and selected impression.

[0724] Step 7:

[0725] The terminal displays the generated reply message to the user, allowing the user to review the content and customize it as needed.

[0726] Step 8:

[0727] Once the user approves the message, the device sends the message to the recipient, completing the communication.

[0728] This process allows users to recognize their own emotional state and effectively communicate the impression they want to make to others.

[0729] (Example 2)

[0730] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0731] Conventional communication systems struggle to appropriately reflect user emotions and generate messages that create a desirable impression on the recipient. This problem hinders efficient and natural dialogue, especially in communication where emotions are complexly intertwined. Therefore, there is a need for technology that simultaneously considers the user's emotions and the recipient's state to generate the optimal response.

[0732] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0733] In this invention, the server includes means for collecting communication data, means for analyzing the communication data to identify the emotional state of the other party, and means for analyzing the user's biosignals to recognize emotions. This makes it possible to generate a response message that gives the most appropriate impression to the other party while taking the user's emotions into consideration.

[0734] "Communication data" refers to all information transmitted and received through electronic means, and specifically includes text messages and voice data.

[0735] "The other party's emotional state" refers to a description of the other party's emotions and mental state, analyzed from communication data.

[0736] "Biosignals" refer to data that indicates the user's physical state, including physical measurements such as heart rate and voice tone.

[0737] "Impression options" refer to multiple choices presented to a user to select the impression they want to convey to the other party.

[0738] A "response message" refers to a message that facilitates natural conversation, generated with consideration for the emotional state of both the user and the other party.

[0739] A "user interface" refers to an interactive mechanism that allows a user to interact with a computer system and perform various operations.

[0740] This invention is a system designed to facilitate communication between a user and another party. It utilizes communication data and biosignals to analyze emotions and generate optimal response messages. The roles of each component of this system are described below.

[0741] The device serves as a communication tool used by users on a daily basis and plays a role in collecting text messages and voice data. This device, such as a smartphone or PC, acquires this data in real time through dedicated applications and APIs, and also collects biometric signals such as heart rate and voice tone. Wearable devices and dedicated sensors are used to collect these biometric signals.

[0742] The collected data is analyzed by a server. The server possesses powerful processing capabilities and utilizes natural language processing techniques and machine learning algorithms to analyze the communication data. In the future, models such as BERT (Bidirectional Encoder Representations from Transformers) and LSTM (Long Short-Term Memory) may be used. Based on the emotional state of the other party and the user's emotions obtained from the emotion analysis, the server generates impression options for the user.

[0743] The user can review the impression options suggested by the server and select the one they deem most effective. Based on this selection, the server uses a generative AI model to create a natural and harmonious response message. This response message is presented to the user through the user interface and can be customized as needed.

[0744] For example, if a user is asked "How was it?" by their partner, and the device recognizes the user's emotion as "tired," the server will present impressions such as "I want to relax" or "I want to be encouraged." If the user selects "I want to relax," the server will generate a message such as "I'm a little tired today, but it was great chatting!" This process helps users effectively convey their emotions while maintaining natural communication.

[0745] The AI ​​generation model is input with a prompt such as, "If the user's emotional state is fatigued, they have selected 'I want to relax' as the impression they want to give. Based on this, please generate a natural reply message." In this way, the system provides an accurate and flexible conversational experience.

[0746] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0747] Step 1:

[0748] The device collects text messages and voice data in real time from the user's communication tools. Furthermore, it measures the user's heart rate and voice tone through biometric sensors. This data is acquired as input necessary for recognizing emotional states.

[0749] Step 2:

[0750] Inside the device, an emotion engine analyzes the user's emotions based on biosignal data. This analysis uses a machine learning model, with heart rate and voice tone fluctuations serving as foundational data for inferring the user's emotions. The analysis results are output as the user's emotional state and sent to the server.

[0751] Step 3:

[0752] The server receives text messages sent from terminals and analyzes their content using natural language processing (NLP) techniques. Specifically, it performs processes such as text tokenization, sentiment analysis, and keyword extraction to identify the recipient's emotional state. This analysis result serves as foundational data for evaluating the impression a user has of another person.

[0753] Step 4:

[0754] The server generates impression options to present to the user based on the analyzed emotional states of the user and the other party. In this process, optimal options are provided based on past data and statistical models, taking into account the user's communication style and situation. The generated options are then sent from the server to the user.

[0755] Step 5:

[0756] The user reviews the impression options provided by the server and selects the impression they wish to convey. The user's selection is intuitively made through a GUI, and this selection significantly influences the generation of the next response message.

[0757] Step 6:

[0758] The server generates a natural response message using a generative AI model based on the impression selected by the user. In this generation process, the prompt "If the user's emotional state is fatigued, they have selected 'I want to relax' as the impression they want to give to the other party. Based on this, please generate a natural reply message." is used as input data, and the AI ​​model outputs optimized text.

[0759] Step 7:

[0760] The terminal displays the response message received from the server to the user, providing an interface that allows the user to customize the content as needed. The user then makes a final confirmation of the message.

[0761] Step 8:

[0762] The device sends a response message to the other party, acknowledging the user's input. This completes the communication and effectively conveys the user's emotions.

[0763] (Application Example 2)

[0764] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0765] In online communication, it is necessary to appropriately understand users' emotions and provide support for making optimal responses and product suggestions based on those emotions. Conventional systems lack the ability to provide individualized support that takes into account the user's emotional state, resulting in a decline in the quality of dialogue. Furthermore, the lack of systems that offer product suggestions tailored to the user's emotions means that appropriate purchasing support is not being provided.

[0766] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0767] In this invention, the server includes means for collecting communication information, means for analyzing the collected communication information and identifying the emotional state of the other party, and means for generating selectable impression options and product suggestions to present to the user based on the identified emotional state. This enables communication tailored to the user's emotions and optimal product suggestions.

[0768] "Communication information" refers to information such as messages and voice data exchanged between a user and another party.

[0769] "Emotional state" refers to the state of mind and mood of the user or the other party, and is an indicator used to consider the impact on communication by understanding the changes in this state.

[0770] An "impression option" is a set of choices generated to allow a user to select what kind of impression they want to give to the other person.

[0771] A "reply message" is a response that a user sends to another person, optimized based on their emotional state and impression choices.

[0772] A "user interface" refers to the screens and means through which a user operates the system, and to review and customize generated reply messages and product recommendations.

[0773] "Product recommendation" refers to the presentation of product information by suggesting the most suitable product based on the user's emotions and choices.

[0774] "Natural language processing technology" is a general term for technologies used to enable computers to understand and process human language.

[0775] The system for carrying out this invention includes a terminal used by the user and a server that processes information. First, the terminal collects the user's voice data and text messages. This information is analyzed by an emotion engine to identify the user's emotional state. The emotion engine uses biosignals such as the user's heart rate and voice tone as input data to accurately recognize emotions.

[0776] The server receives communication information sent from the terminal and analyzes the information using natural language processing technology. As a result of the analysis, the server identifies the other party's emotional state and understands the impression the user has. Based on this information, the server generates selectable impression options and product recommendations. When the user selects an impression or product from the provided options, the server uses a generative AI model to generate a natural response message corresponding to the selected impression.

[0777] Users can review the generated messages and product suggestions displayed on their device and customize them as needed. The device sends a reply message to the recipient after receiving user approval. One possible application of this system is user support during product purchases in a virtual store. For example, if a user expresses a desire to "relax," the server might suggest a product such as, "This candle promotes relaxation."

[0778] Through this system, users can receive optimal communication and product suggestions that take their emotions into consideration, improving the quality of the interaction. An example of a prompt message is, "Generate a message that provides optimal product suggestions based on the emotion analysis results."

[0779] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0780] Step 1:

[0781] The device collects user voice data and text messages. This data is sent to an emotion engine, which analyzes it along with biosignals (e.g., heart rate, voice tone) necessary to recognize the user's emotional state. The input is the user's raw communication data, and the output is the user's emotional state as recognized by the emotion engine. Machine learning algorithms are applied based on the data to recognize the emotional state.

[0782] Step 2:

[0783] The server receives the user's emotional state and communication information transmitted from the terminal. Based on the received data, natural language processing techniques are used to analyze the other party's emotional state and identify the emotional relationship between the two parties. The input for this step is the user's emotional state and communication information, and the output is the other party's emotional state and what they understood it to be. Text mining and sentiment analysis techniques are used for the analysis.

[0784] Step 3:

[0785] The server generates selectable impression options and product suggestions based on the analysis results. This is done using an AI model that generates options that correspond to the user's emotional state and the other party's emotions. The input is the analysis results from the previous stage, and the output is the impression options and product suggestions. Text generation techniques, including topic modeling, are used to generate the impression options and product suggestions.

[0786] Step 4:

[0787] The user reviews impression options and product suggestions via their device and makes a selection. The user's selection is sent to the server, which uses a generative AI model to generate a natural-sounding reply message based on this selection. The input is the user's selection, and the output is a natural-sounding reply message. Advanced natural language generation techniques, such as transformer models, are used for generation.

[0788] Step 5:

[0789] The terminal displays the generated message to the user, and after the user's final confirmation or customization, sends the approved reply message to the recipient. At this stage, the input consists of the reply message and the user's customizations, while the output is the final response message sent. The sending function utilizes standard communication protocols for data transmission.

[0790] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0791] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0792] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0793] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0794] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0795] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0796] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0797] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0798] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0799] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0800] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0801] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0802] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0803] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0804] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0805] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0806] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0807] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0808] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0809] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0810] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0811] The following is further disclosed regarding the embodiments described above.

[0812] (Claim 1)

[0813] Means for collecting communication information,

[0814] A means of analyzing collected communication information to identify the emotional state of the other party,

[0815] A means for generating selectable impression options to present to the user based on an identified emotional state,

[0816] A means of generating the optimal reply message based on the selected impression,

[0817] The means of sending a reply message to the other party,

[0818] A system that includes this.

[0819] (Claim 2)

[0820] The system according to claim 1, which uses natural language processing technology for analyzing communication information.

[0821] (Claim 3)

[0822] The system according to claim 1, further comprising a user interface that allows the user to customize the generated reply message.

[0823] "Example 1"

[0824] (Claim 1)

[0825] A device for collecting communication data,

[0826] A device that analyzes collected communication data to identify the emotional state of others,

[0827] A device that creates selectable impression options to display to the user based on an identified emotional state,

[0828] A device that generates an appropriate response message according to the selected impression,

[0829] A device that sends a response message to another party,

[0830] A system that includes this.

[0831] (Claim 2)

[0832] The system according to claim 1, which utilizes natural language processing technology for analyzing communication data.

[0833] (Claim 3)

[0834] The system according to claim 1, further comprising a human-machine interface that allows the user to modify the generated response message.

[0835] "Application Example 1"

[0836] (Claim 1)

[0837] Means for collecting communication information,

[0838] A means of analyzing collected communication information to identify the emotional state of the other party,

[0839] A means for generating selectable impression options to present to the user based on an identified emotional state,

[0840] A means of generating the optimal reply message based on the selected impression,

[0841] A means of detecting anomalies from the surrounding environment and communicating with people on site in response to those anomalies,

[0842] A means of analyzing people's emotions in real time and generating appropriate instruction messages,

[0843] The means of sending a reply message to the other party,

[0844] A system that includes this.

[0845] (Claim 2)

[0846] The system according to claim 1, which uses natural language processing technology for analyzing communication information.

[0847] (Claim 3)

[0848] The system according to claim 1, further comprising a user interface that allows the user to customize the generated reply message.

[0849] "Example 2 of combining an emotion engine"

[0850] (Claim 1)

[0851] Means for collecting communication data,

[0852] A means of analyzing collected communication data to identify the emotional state of the other party,

[0853] A means of recognizing emotions by analyzing the user's biosignals,

[0854] A means for generating selectable impression options based on identified emotional states and the user's emotions,

[0855] A means of generating the optimal response message based on the selected impression,

[0856] A means of sending a response message to the other party,

[0857] A system that includes this.

[0858] (Claim 2)

[0859] The system according to claim 1, which uses natural language and machine learning techniques for analyzing communication data and biosignals.

[0860] (Claim 3)

[0861] The system according to claim 1, further comprising a user interface that allows the user to modify the generated response message.

[0862] "Application example 2 when combining with an emotional engine"

[0863] (Claim 1)

[0864] Means for collecting communication information,

[0865] A means of analyzing collected communication information to identify the emotional state of the other party,

[0866] A means for generating selectable impression options to present to the user based on an identified emotional state,

[0867] A means of generating the optimal reply message based on the selected impression,

[0868] A device equipped with a user interface for displaying generated reply messages and making product recommendations,

[0869] A means of suggesting the optimal product based on emotion analysis,

[0870] The means of sending a reply message to the other party,

[0871] A system that includes this.

[0872] (Claim 2)

[0873] The system according to claim 1, which uses natural language processing technology for analyzing communication information.

[0874] (Claim 3)

[0875] The system according to claim 1, comprising a user interface that allows the user to customize the generated reply message and product suggestion. [Explanation of symbols]

[0876] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Means for collecting communication information, A means of analyzing collected communication information to identify the emotional state of the other party, A means for generating selectable impression options to present to the user based on an identified emotional state, A means of generating the optimal reply message based on the selected impression, The means of sending a reply message to the other party, A system that includes this.

2. The system according to claim 1, which uses natural language processing technology for analyzing communication information.

3. The system according to claim 1, further comprising a user interface that allows the user to customize the generated reply message.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A