System
A system with a user interface, natural language processing, and notification mechanism addresses the challenge of employee mental state expression, allowing supervisors to provide effective care by analyzing and summarizing employee dialogue.
Patent Information
- Application Number
- JP2024119088
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
AI Technical Summary
The increase in remote work has made it difficult for employees to express their stress and mental state due to psychological barriers, and traditional methods require self-reporting, which hinders effective understanding and care by superiors.
A system with a user interface for interaction, a natural language processing mechanism to analyze dialogue, a filtering mechanism to summarize emotional content, and a notification mechanism to provide feedback to supervisors, ensuring a positive dialogue environment and accurate mental care.
Facilitates comfortable expression of stress and mental state, enabling supervisors to provide effective care by accurately understanding employee emotions and stress levels.
Smart Images

Figure 2026018027000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] With the increase in remote work and fewer opportunities for face-to-face communication, it has become difficult to understand and care for employees' stress and mental state. Traditional methods require employees to report stress and problems to their superiors themselves, which can make it difficult for them to express their true feelings due to psychological barriers. It is also difficult for superiors to properly understand the mental state of each employee and take the necessary measures quickly. There is a need for a system that can solve these issues and enable smooth communication and mental care. [Means for solving the problem]
[0005] The present invention is a system that includes a means having a user interface for employees to interact with a dialogue agent (an all-affirmative bot), a natural language processing means that analyzes the content of the dialogue between the dialogue agent and the employee to determine the employee's emotional state and level of stress, a filtering means that filters and summarizes the analyzed content of the dialogue, and a notification means that notifies a supervisor of the information processed by the filtering means. The dialogue agent always returns a positive response, providing an environment in which employees can easily release stress. Furthermore, the notification means can provide feedback to the supervisor regarding the employee's mental care and encourage appropriate responses. In this way, an environment is created in which employees can comfortably open up, and supervisors can obtain information to provide effective mental care.
[0006] "User" refers to an employee or user who interacts with the system.
[0007] A "conversational agent" refers to artificial intelligence-based software that can converse with a user and provide positive responses.
[0008] "User interface" refers to the screen and operation means through which a user interacts with a dialogue agent.
[0009] "Natural language processing means" refers to technology for analyzing the content of the dialogue between the user and the dialogue agent and determining the emotional state and level of stress.
[0010] "Filtering means" refers to a mechanism for summarizing the content of the dialogue analyzed by the natural language processing means and organizing it to eliminate excessive social influence.
[0011] "Notification means" refers to a reporting technique that transfers information processed by the filtering means to a superior and prompts them to take appropriate action.
[0012] "Superior" refers to the user's direct manager or instructor. [Brief explanation of the drawings]
[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram illustrating a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0021] [First embodiment]
[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0034] This invention is a dialogue agent system that uses an all-affirmation bot to provide mental care and communication support for employees. This system is implemented in the following steps.
[0035] System Configuration
[0036] This system includes users (employees), devices (such as PCs and smartphones), a server, and a conversational agent (an all-affirmation bot). These components work together to understand the user's mental state and provide appropriate feedback to their superiors.
[0037] User Registration and Login
[0038] The user first launches the application and registers by entering the required information such as name, email address, and employee ID on the user registration screen. The device sends the entered information to the server, which stores it in a database. After registration is complete, the user logs in by entering their email address and password on the login screen. The server authenticates the user, and if authentication is successful, displays the user's dashboard.
[0039] Conversation with the all-affirmation bot
[0040] The user clicks the "Start Dialogue" button on the dashboard screen to begin a dialogue with the conversational agent (all-affirmative bot). The device starts voice recognition, and the user begins speaking. The all-affirmative bot always returns a positive response to the user's statements. For example, if the user says, "I'm very tired today," the all-affirmative bot will reply, "That must have been tough. Good job!" The device sends this dialogue content to the server in real time.
[0041] Call analysis
[0042] When the conversation ends, the device notifies the server that the call has ended. The server then sends the saved conversation log to a natural language processing model and begins analyzing the text content. The generative AI model extracts important keywords and emotional trends from the conversation content and creates a summary. The server then applies a social filter to the generated summary to organize the content that should be recognized. For example, a user's statement that "I have too many meetings this week and my work isn't progressing" is summarized as "Work delays due to being busy are causing me stress."
[0043] Information Sharing and Notification
[0044] The server prepares the filtered summary to be sent to the supervisor. The supervisor logs in to a dedicated dashboard from their own device and checks the feedback on the employee's mental state. The server responds to the supervisor's actions by displaying a concise summary of the important parts and providing hints on what points to pay attention to. If necessary, the supervisor can provide individual feedback or send a message to offer support.
[0045] Specific examples
[0046] Example 1: Everyday conversation
[0047] The user schedules a conversation with the All-Affirmation Bot every day at 2:00 PM. Today, the user says, "I have too many meetings this week, so I can't get my work done." The All-Affirmation Bot responds, "That must have been tough, you've worked hard!" The server records the statement, "I have too many meetings this week, so I can't get my work done," and the generative AI model summarizes it as, "The delays in work caused by too many meetings are a stressful factor." The server applies a social filter to the summary and shares it with the user's superiors, concluding, "The user is stressed by delays in work caused by being busy."
[0048] Example 2: Important Feedback
[0049] In conversations with the All-Affirmation Bot, the user repeatedly states, "The project deadline is too tight." The All-Affirmation Bot responds, "I understand the pressure. You're doing a great job!" The server summarizes this as "The tight project deadline is the main cause of stress," and immediately notifies the supervisor. The supervisor can then take action based on the employee's feedback, such as "reevaluating the project deadline."
[0050] This allows the user to more easily reduce stress, and allows the supervisor to obtain information for effective mental care.
[0051] The processing flow will be explained below.
[0052] Step 1:
[0053] The user launches the application, enters the necessary information such as "name," "email address," and "employee ID" on the user registration screen, and presses the "Register" button.
[0054] Step 2:
[0055] The terminal temporarily stores the input user information and transmits the data to the server.
[0056] Step 3:
[0057] The server stores the received user information in a database and returns a response indicating successful registration to the terminal.
[0058] Step 4:
[0059] The terminal displays a "Registration Complete" message to the user.
[0060] Step 5:
[0061] The user enters their "email address" and "password" on the login screen and presses the "Login" button.
[0062] Step 6:
[0063] The terminal transmits the entered login information to the server.
[0064] Step 7:
[0065] The server authenticates the user using a database, and if authentication is successful, returns the user's dashboard information to the terminal.
[0066] Step 8:
[0067] The terminal displays a dashboard screen to the user.
[0068] Step 9:
[0069] The user clicks the "Start conversation" button on the dashboard screen.
[0070] Step 10:
[0071] The terminal starts voice recognition and the user begins speaking.
[0072] Step 11:
[0073] The conversational agent (all-affirmative bot) responds affirmatively to user comments. For example, if the user says, "I'm very tired today," the agent responds, "That must have been hard! You did a great job!"
[0074] Step 12:
[0075] The terminal transmits the content of the conversation between the all-affirming bot and the user to the server in real time.
[0076] Step 13:
[0077] The server stores the content of the conversation as a log and performs processing such as sentiment analysis and keyword extraction in real time as needed.
[0078] Step 14:
[0079] When the interaction ends, the terminal notifies the server of the interaction end.
[0080] Step 15:
[0081] The server sends the saved dialogue log to a natural language processing model to begin analyzing the text content.
[0082] Step 16:
[0083] The generative AI model extracts important keywords and sentiment trends from the conversation and creates a summary, such as, "The user has been busy this week, and is feeling particularly stressed by the large number of meetings."
[0084] Step 17:
[0085] The server applies a social filter to the generated summary to organize the content that should be recognized, for example, summarizing it as "Work delays due to busyness are a cause of stress."
[0086] Step 18:
[0087] The server prepares to send the filtered summary to the superior.
[0088] Step 19:
[0089] Supervisors can log into a dedicated dashboard from their own devices and check feedback on employees' mental health.
[0090] Step 20:
[0091] The server displays a concise summary of important points in response to the supervisor's operation, and also provides hints on what points to pay attention to.
[0092] Step 21:
[0093] If needed, a manager can send a message to provide individual feedback or assistance.
[0094] The above are the specific processing steps for a mental care and communication assistance application using a fully affirmative bot.
[0095] Example 1
[0096] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0097] In modern corporate environments, employee mental health is an important issue, but many systems lack the means to properly understand users' emotional states and stress levels, or to provide appropriate feedback to superiors. Therefore, a method for effectively managing employee mental health is needed. Furthermore, conventional conversational agents may respond negatively to user comments, which may increase employee stress.
[0098] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0099] In this invention, the server includes: a user interface that provides a method for a user to interact with a dialogue agent; a user interface that allows the user to initiate a dialogue with the dialogue agent and record the dialogue content; a generative AI model that allows the dialogue agent to always generate positive responses; a natural language processing unit that analyzes the dialogue content between the dialogue agent and the user to determine the user's emotional state and level of stress; a summarizing unit that summarizes the dialogue content analyzed by the natural language processing unit and extracts important keywords and emotional trends; a notification unit that notifies the supervisor of the information processed by the filtering unit; and a notification unit that provides the supervisor with feedback regarding the user's mental care and encourages the supervisor to take appropriate measures. This makes it possible to accurately grasp the employee's emotional state and level of stress and provide appropriate feedback to the supervisor.
[0100] The "user interface" refers to the operation screen and input means by which the user operates the system and starts a dialogue with the dialogue agent.
[0101] A "generative AI model" is an algorithm or program that generates positive responses in natural language based on what a user says.
[0102] "Natural language processing" is a technology that analyzes text data and extracts important keywords and emotional trends from its content.
[0103] A "dialogue agent" is a program or system that interacts with a user and generates responses to their utterances.
[0104] "Notification means" refers to the communication means or process for sending analyzed information and feedback to superiors.
[0105] The "filtering means" is a processing means for organizing the dialogue content analyzed by natural language processing and generating an appropriate summary.
[0106] "Emotional state" refers to the user's psychological state, and indicates stress, satisfaction, fatigue, etc.
[0107] The "dashboard" is an operation screen accessed by users who log in to the system, and has a central function for starting a dialogue and checking feedback.
[0108] This invention is a dialogue agent system for the purpose of providing mental care and communication support for employees, and includes a user (employee), a terminal (such as a PC or smartphone), a server, and a dialogue agent (an all-affirmation bot). These components work together to grasp the user's mental state and provide appropriate feedback to superiors.
[0109] System Configuration
[0110] The system consists of the following main components:
[0111] User Interface
[0112] The user uses this interface to start a dialogue with the dialogue agent, inputting and operating the agent according to the situation. The user interface is provided as a software application that runs on a device such as a PC or smartphone.
[0113] Generative AI Models
[0114] The conversational agent uses a generative AI model (e.g., ChatGPT) that generates positive responses based on user utterances. This model uses natural language processing techniques to understand what the user is saying and always generates a positive response.
[0115] Natural Language Processing
[0116] On the server side, natural language processing technology (e.g., BERT) is used to analyze the content of the dialogue between the conversation agent and the user, extracting important keywords and emotional trends from the user's speech and determining the user's emotional state and level of stress.
[0117] Notification means
[0118] The information processed by the filtering means is notified to the superior by the server. The notification means ensures that the data is correctly filtered and that the superior is provided with important information along with a concise summary.
[0119] User Registration and Login
[0120] The user starts the application and enters the required information such as name, email address, employee ID, etc. on the user registration screen. The device sends this information to the server, which stores it in a database. When the user enters their email address and password on the login screen, the server authenticates the user and, if successful, displays the dashboard.
[0121] Conversation with the all-affirmation bot
[0122] When the user clicks the "Start Dialogue" button on the dashboard, a dialogue with the conversational agent (all-positive bot) begins. The device starts voice recognition and captures the user's voice input. The server converts the received voice data into text and sends it to the all-positive bot. The all-positive bot uses a generative AI model to generate a positive response, and the device presents this response to the user.
[0123] Call analysis and summarization
[0124] When the dialogue ends, the device notifies the server. The server then sends the saved dialogue log to a natural language processing model, which begins analyzing the text content. The generative AI model extracts important keywords and emotional trends from the dialogue and generates a summary. For example, if a user says, "I have too many meetings this week, so I'm not making much progress on my work," the summary might read, "Work delays due to being too busy are causing me stress."
[0125] Information Sharing and Notification
[0126] The server applies a social filter to the generated summary, organizes important information, and notifies the supervisor. The supervisor logs in to a dedicated dashboard and checks the feedback. This allows the supervisor to take appropriate action based on the employee's mental state.
[0127] Specific prompt examples
[0128] Below are some examples of prompt sentences.
[0129] "Users interact with the all-affirmation bot and send what they say to your server. Your natural language processing model analyzes the text and creates a summary."
[0130] This allows the system to effectively support employees' mental health care and provide supervisors with the information they need to provide appropriate feedback.
[0131] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0132] Step 1: User Registration
[0133] A user starts an application and enters the required information such as name, email address, and employee ID on the user registration screen. The entered information (name, email address, employee ID) is input data. The device sends this input data to the server, which receives it and stores it in a database. Specifically, the server receives an HTTP POST request and registers it as a new record in the database (e.g., MySQL). The output is that the new user has been registered and the information is stored in the database.
[0134] Step 2: User Login
[0135] A user attempts to log in by entering an email address and password. The device sends this authentication information (email address, password) to the server. The server searches for the corresponding user information in the database and compares it with the input data. If authentication is successful, the server returns an authentication success message and the user's dashboard information to the device as output. The device receives this information and displays the user's dashboard.
[0136] Step 3: Initiating a conversation
[0137] The user clicks the "Start conversation" button on the dashboard. This action is registered as input data. The device starts voice input and records what the user says. The recorded voice data is input and the device sends it to the server. The server receives the voice data and converts it into text data using voice recognition software (e.g., Google Cloud Speech-to-Text API). The converted text data is output.
[0138] Step 4: Generate a response
[0139] The server inputs the received text data into a generative AI model (e.g., ChatGPT) to generate a positive response. The generative AI model performs natural language processing based on the input data (user utterances) to generate a positive response. This response text becomes the output. The server sends this response text to the device, which then presents the response to the user in voice or text.
[0140] Step 5: Record the conversation
[0141] The device sends the content of the conversation (user's comments and all affirmative bot's responses) to the server in real time. The server receives this and stores it in a database. The saved conversation log is the output. This log is used for later analysis.
[0142] Step 6: End of conversation
[0143] When the conversation ends, the device notifies the server. This notification becomes the input data. The server acquires the conversation log and sends it to a natural language processing model (e.g., BERT). This starts the analysis of the text content.
[0144] Step 7: Text analysis and summary generation
[0145] The server uses a natural language processing model to extract important keywords and emotional trends from the dialogue. This analysis is a data calculation based on the input data (dialogue log). Summary text is generated as the analysis result. This summary text is the output data. For example, a statement such as "There are too many meetings this week, so work isn't progressing" is summarized as "Work delays due to being busy are causing stress."
[0146] Step 8: Filtering and Notifications
[0147] The server applies a social filter to the generated summary text and organizes important information. The filtered text is output. This text is sent to the superior via a notification method. The superior receives the notification and checks the feedback on a dedicated dashboard. For example, the superior may receive specific feedback such as, "Work delays due to busyness are the cause of stress."
[0148] (Application example 1)
[0149] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0150] In conventional factory work, operators and engineers often suffer from excessive stress and fatigue, which has a negative impact on productivity and quality. In particular, a lack of psychological care can lead to problems such as a decrease in motivation and an increase in turnover. Therefore, a new system is needed to efficiently provide mental care for operators and engineers working in factories and improve the quality of their work.
[0151] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0152] In this invention, the server includes: means for having a user interface that provides a method for a user to interact with a dialogue agent; natural language processing means for analyzing the dialogue content between the dialogue agent and the user to determine the user's emotional state and level of stress; filtering means for filtering and summarizing the dialogue content analyzed by the natural language processing means; notification means for notifying a supervisor of the information processed by the filtering means; means for using a generative AI model that recognizes the user's everyday conversation and generates a positive response; speech recognition means for converting speech input from the user into text in real time; means for executing a program that controls the notification means to summarize and notify the dialogue content; and feedback means for the supervisor to take an appropriate action based on the notified information. This makes it possible to reduce the psychological stress of operators and engineers working in a factory and provide effective mental care.
[0153] A "user interface" is the means of interaction that allows a user to interact with a system or device.
[0154] A "dialogue agent" is a software agent that obtains information through dialogue with a user and returns an appropriate response.
[0155] "Natural language processing means" refers to algorithms or systems that analyze natural language text spoken by a user and determine their emotional state and level of stress.
[0156] The "filtering means" is a process for organizing the dialogue content analyzed by the natural language processing means, and extracting and summarizing important information.
[0157] "Notification means" refers to a function or system for notifying the processed filtering results to superiors or managers.
[0158] A "generative AI model" is an artificial intelligence model that generates appropriate responses to user input.
[0159] "Speech recognition means" refers to technology or devices that convert a user's speech into text in real time.
[0160] The "means for executing a program" refers to the software and hardware configuration for collecting, analyzing, and filtering the contents of the conversation and controlling the notification means.
[0161] "Feedback means" refers to a specific method or system that allows a superior to take appropriate action based on the notified information.
[0162] This invention is a dialogue agent system for providing mental care to operators and engineers working in factories. This system includes users (operators and engineers), terminals (smartphones, tablets, etc.), a server, and a dialogue agent (an all-affirmative bot).
[0163] System Configuration
[0164] The system works by coordinating the following components:
[0165] 1. User Interface: Provides an interface for the user to interact with the conversational agent, which can be operated by a touch screen or voice input.
[0166] 2. Conversational Agent: Engages in natural dialogue with the user. Conversational agents use generative AI models to always generate positive responses to user utterances.
[0167] 3. Speech recognition: Converts voice input into text in real time. This function uses the Google Speech-to-Text API, for example.
[0168] 4. Natural language processing: Analyzes the dialogue and determines the user's emotional state and stress level. This process uses natural language processing technologies such as GPT-3.5.
[0169] 5. Filtering means: Filters the dialogue content analyzed by the natural language processing means and summarizes important information.
[0170] 6. Notification: The information summarized by the filtering method is notified to superiors and managers via email or dashboard.
[0171] 7. Feedback measures: Provide feedback to supervisors and managers to take appropriate action based on the information provided.
[0172] Program processing
[0173] The server receives the dialogue between the conversational agent and the user in real time and generates a positive response using a generative AI model. The use of GPT-3.5 as the generative AI model allows for natural dialogue. All dialogue is saved in text format and later analyzed using natural language processing.
[0174] The device receives voice input from the user and converts it into text using a speech recognition method (e.g., Google Speech-to-Text API), which then transmits the user's speech as text to the server in real time.
[0175] The filtering means processes the acquired text data and extracts important keywords and phrases, and based on this information, summarizes the conversation and notifies superiors or managers.
[0176] The notification method sends filtered summary information to superiors and managers via email or a dedicated dashboard, allowing them to quickly take the necessary measures to care for the user's mental health.
[0177] Specific examples
[0178] For example, if a factory operator uses the system and says, "I'm so busy today, I'm tired," the all-affirmation bot will respond, "That must have been tough, good job!" After the conversation ends, the system summarizes the content of the conversation and provides feedback to the supervisor or manager that "the operator's fatigue due to being too busy is the cause."
[0179] Prompt Sentence Examples
[0180] Input prompt: "I'm too busy and tired today."
[0181] Model response: "That was hard work, good job!"
[0182] By implementing this system, it is possible to improve the working environment within the factory, maintain the motivation of operators and engineers, and reduce stress.
[0183] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0184] Step 1:
[0185] The user performs voice input. Specifically, the user speaks into the device's microphone to start a conversation. The input is what the user says, and is captured by the device as voice data.
[0186] Step 2:
[0187] A speech recognition means is activated on the terminal. The speech recognition means converts the user's voice data into text data. This process uses speech recognition technology such as the Google Speech-to-Text API. The output is data that converts the user's speech into text.
[0188] Step 3:
[0189] The device sends text data to the server, which then passes it to the generative AI model. The input is the converted text data, which is then sent to the server.
[0190] Step 4:
[0191] The conversational agent in the server uses a generative AI model to generate a positive response to the user's utterance. Specifically, a GPT-3.5 model is used to generate a prompt sentence based on the input text. The output is a positive response text.
[0192] Step 5:
[0193] The generated positive response text is returned from the server to the terminal. The terminal reproduces this text data as voice through the voice output means. The input is the generated response text, and the output is a voice response to the user.
[0194] Step 6:
[0195] The server stores the dialogue content as a log, and natural language processing means analyzes the text data. In this analysis process, the user's emotional state and stress level are determined. The input is the user's entire dialogue text, and the output is the evaluation results of the emotional state and stress.
[0196] Step 7:
[0197] The filtering means extracts important keywords and phrases from the dialogue text and generates a summary. In this step, points of particular interest are organized from the dialogue content. The input is the analysis result by the natural language processing means, and the output is the summary text.
[0198] Step 8:
[0199] The summarized information is sent to a supervisor or manager through a notification mechanism, which can be email or a dashboard. The input is the summary text and the output is the notification message provided to the supervisor or manager.
[0200] Step 9:
[0201] The supervisor or manager takes appropriate action based on the notified information. This involves implementing specific measures and countermeasures using feedback methods. The input is the notification message, and the output is specific feedback and countermeasures.
[0202] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0203] The present invention is a dialogue agent system that utilizes an all-affirmation bot, with the aim of providing mental care and communication assistance to users. By incorporating an emotion engine into this system, it is possible to more accurately recognize the user's emotional state and provide appropriate feedback. Specific embodiments of this system are described below.
[0204] System Configuration
[0205] This system includes users (employees), devices (PCs, smartphones, etc.), a server, a dialogue agent (a completely affirmative bot), and an emotion engine. These components work together to understand the user's mental state and provide appropriate feedback to superiors.
[0206] User Registration and Login
[0207] The user first launches the application and registers by entering the required information such as name, email address, and employee ID on the user registration screen. The device sends the entered information to the server, which stores it in a database. After registration is complete, the user logs in by entering their email address and password on the login screen. The server authenticates the user, and if authentication is successful, displays the user's dashboard.
[0208] Conversation with an all-affirmation bot and emotion recognition
[0209] The user clicks the "Start conversation" button on the dashboard screen to begin a conversation with the conversational agent (all-affirmative bot). The device starts voice recognition, and the user begins speaking. The all-affirmative bot always returns a positive response to the user's statements. For example, if the user says, "I'm very tired today," the bot replies, "That must have been tough! You've done well!"
[0210] The emotion engine analyzes the user's voice tone, facial expressions, and text content to recognize their emotions. The recognized emotional information is reflected in the conversational agent's responses. For example, if the user is determined to be very tired, the all-affirmation bot will adjust its responses to be more warmhearted.
[0211] Analysis of call content and emotional information
[0212] When the conversation ends, the device notifies the server that the call has ended. The server then sends the saved conversation log and emotion recognition results to a natural language processing model, which begins analyzing the text content and emotional information. The generative AI model extracts important keywords and emotional trends from the conversation content and creates a summary. The server then applies a social filter to the generated summary to organize the content that should be recognized. For example, information such as "There are too many meetings this week, so work is not progressing" and "The user is feeling tired" is summarized as "Work delays due to busyness are the cause of stress."
[0213] Information Sharing and Notification
[0214] The server prepares the filtered summary to be sent to the supervisor. The supervisor logs in to a dedicated dashboard from their own device and checks the feedback on the employee's mental state. The server responds to the supervisor's actions by displaying a concise summary of the important parts and providing hints on what points to pay attention to. If necessary, the supervisor can provide individual feedback or send a message to offer support.
[0215] Specific examples
[0216] Example 1: Everyday dialogue and emotion recognition
[0217] The user schedules a conversation with the All-Affirmation Bot every day at 2:00 PM. Today, the user says, "There are too many meetings this week, and I'm not making progress on my work." The All-Affirmation Bot replies, "That must have been tough, you've worked hard!" At this point, the emotion engine recognizes a strong sense of fatigue from the user's tone of voice. The server records the statement, "There are too many meetings this week, and I'm not making progress on my work," along with the "strong sense of fatigue," and the generative AI model summarizes it as, "The delays in work due to the many meetings are a stressful factor and cause fatigue." The server applies a sociality filter to the above summary and shares it with the supervisor, concluding, "The user is experiencing stress due to work delays caused by being busy, and is feeling very fatigued."
[0218] Example 2: Critical feedback and emotion recognition
[0219] In conversation with the All-Affirmation Bot, the user repeatedly states, "The project deadline is too tight." The All-Affirmation Bot responds, "I understand the pressure. You're doing a great job!" At this point, the emotion engine recognizes the user's strong stress from their facial expressions and voice. The server summarizes this as "The tight project deadline is the main cause of stress," and immediately notifies this information to the supervisor. The supervisor can then take action based on the employee's feedback, such as "reevaluating the project deadline."
[0220] This makes it easier for users to reduce stress and provides supervisors with information to provide effective mental care.The introduction of the emotion engine allows for a more accurate understanding of the user's emotional state, improving the quality of responses and feedback.
[0221] The processing flow will be explained below.
[0222] Step 1:
[0223] The user launches the application, enters the necessary information such as "name," "email address," and "employee ID" on the user registration screen, and presses the "Register" button.
[0224] Step 2:
[0225] The terminal temporarily stores the input user information and transmits the data to the server.
[0226] Step 3:
[0227] The server stores the received user information in a database and returns a response indicating successful registration to the terminal.
[0228] Step 4:
[0229] The terminal displays a "Registration Complete" message to the user.
[0230] Step 5:
[0231] The user enters their "email address" and "password" on the login screen and presses the "Login" button.
[0232] Step 6:
[0233] The terminal transmits the entered login information to the server.
[0234] Step 7:
[0235] The server authenticates the user using a database, and if authentication is successful, returns the user's dashboard information to the terminal.
[0236] Step 8:
[0237] The terminal displays a dashboard screen to the user.
[0238] Step 9:
[0239] The user clicks the "Start conversation" button on the dashboard screen.
[0240] Step 10:
[0241] The terminal starts voice recognition and the user begins speaking.
[0242] Step 11:
[0243] The conversational agent (all-affirmative bot) responds affirmatively to user comments. For example, if the user says, "I'm very tired today," the agent responds, "That must have been hard! You did a great job!"
[0244] Step 12:
[0245] The emotion engine recognizes the user's emotions by analyzing their voice tone, facial expressions, and text content, for example, reading tiredness from their voice and sadness from their facial expressions.
[0246] Step 13:
[0247] The terminal transmits the content of the conversation between the all-affirmation bot and the user, as well as the emotional data obtained from the emotion engine, to the server in real time.
[0248] Step 14:
[0249] The server stores the dialogue content and emotional data as logs, and performs real-time processing such as emotion analysis and keyword extraction as needed.
[0250] Step 15:
[0251] When the interaction ends, the terminal notifies the server of the interaction end.
[0252] Step 16:
[0253] The server sends the saved dialogue logs and emotion data to a natural language processing model, which begins analyzing the text content and emotion information.
[0254] Step 17:
[0255] The generative AI model extracts important keywords and emotional trends from the conversation and creates a summary, such as, "The user has been busy this week, and is feeling stressed and tired, especially with so many meetings."
[0256] Step 18:
[0257] The server applies a social filter to the generated summary to organize the content that should be recognized, for example, summarizing it as "Work delays due to busyness are a cause of stress."
[0258] Step 19:
[0259] The server prepares to send the filtered summary to the superior.
[0260] Step 20:
[0261] Supervisors can log into a dedicated dashboard from their own devices and check feedback on employees' mental health.
[0262] Step 21:
[0263] The server displays a concise summary of important points in response to the supervisor's operation, and also provides hints on what points to pay attention to.
[0264] Step 22:
[0265] If needed, a manager can send a message to provide individual feedback or assistance.
[0266] Step 23:
[0267] Users receive feedback from their superiors and implement advice on mental care and work improvement.
[0268] The above are the specific processing steps for a mental care and communication assistance application using a fully affirmative bot combined with an emotion engine.
[0269] Example 2
[0270] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0271] To effectively provide mental care to users, it is necessary to accurately grasp the user's emotional state and stress level and provide appropriate feedback accordingly. However, conventional systems lack the means to accurately recognize the user's emotions, making it difficult to provide appropriate responses and feedback. In addition, sharing information with superiors is cumbersome, making it difficult to provide effective mental care to users.
[0272] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0273] In this invention, the server includes means for having a user interface through which the user interacts with the dialogue agent, natural language processing means for analyzing the dialogue between the dialogue agent and the user to determine the user's emotional state and level of stress, filtering means for filtering and summarizing the dialogue analyzed by the natural language processing means, emotion recognition means for analyzing the user's voice tone, facial expressions, and text content to recognize the user's emotions, adjustment means for reflecting the emotion information obtained by the emotion recognition means in the dialogue agent's response content, and notification means for notifying a superior of the information processed by the filtering means. This makes it possible to accurately grasp the user's emotional state and provide appropriate feedback.
[0274] Below are definitions of important terms included in the rewritten claims:
[0275] "User interface" refers to the screen or means by which a user accesses and operates a system.
[0276] An "interactive agent" is a program or system that interacts with a user and responds based on the interaction.
[0277] "Natural language processing means" refers to a technique or means for analyzing the content of a dialogue between a user and a dialogue agent and determining the user's emotional state and stress level.
[0278] "Filtering means" refers to a technology or means for sorting the dialogue content analyzed by natural language processing means and extracting and summarizing important information.
[0279] The "emotion recognition means" refers to a technology or means for analyzing the user's voice tone, facial expression, and text content to recognize the user's emotions.
[0280] The "adjustment means" refers to a technique or means for appropriately changing or adjusting the response content of the dialogue agent based on the emotional information obtained by the emotion recognition means.
[0281] "Notification means" refers to a technique or means for transmitting the information processed by the filtering means to a superior in an appropriate format.
[0282] This invention is a dialogue agent system that uses an all-affirmation bot and an emotion engine, with the aim of providing mental care and communication assistance to users. Specific embodiments of this system are described below.
[0283] System Configuration
[0284] This system includes a user, a terminal, a server, a dialogue agent (a fully affirmative bot), and an emotion engine. These elements work together to understand the user's mental state and provide appropriate feedback.
[0285] User Registration and Login
[0286] 1. The user launches the application and registers by entering information such as their name, email address, and employee ID.
[0287] 2. The terminal sends the entered information to the server, which stores it in a database.
[0288] 3. After completing registration, the user logs in by entering their email address and password.
[0289] 4. The server authenticates the user and, if authentication is successful, displays the user's dashboard.
[0290] Conversation with the all-affirmation bot
[0291] 1. The user clicks the "Start conversation" button on the dashboard screen to begin a conversation with the all-affirmation bot.
[0292] 2. The device starts voice recognition and sends the user's voice to the all-affirmation bot. For example, if the user says, "I'm tired today," the all-affirmation bot replies, "That must have been hard! Good job!"
[0293] emotion recognition
[0294] 1. The device simultaneously transmits the user's voice tone, facial expression, and text content to the emotion engine.
[0295] 2. The emotion engine analyzes this data and recognizes the user's emotions.
[0296] 3. The recognized emotional information is reflected in the responses of the all-positive bot. For example, if the user is judged to be very tired, the all-positive bot will adjust its responses to be more warmhearted.
[0297] Call analysis
[0298] 1. When the conversation ends, the terminal notifies the server that the call has ended.
[0299] 2. The server sends the dialogue log and emotion recognition results to the natural language processing model.
[0300] 3. The generative AI model extracts important keywords and emotional trends from the conversation and creates a summary. For example, the information that "the user is feeling tired" and the statement that "there are too many meetings this week and work isn't progressing" can be summarized as "work delays due to being busy are the cause of stress."
[0301] Information Sharing and Notification
[0302] 1. The server prepares the filtered summary for notification to the supervisor.
[0303] 2. Supervisors log in to a dedicated dashboard and check feedback on employees' mental health.
[0304] 3. The server responds to the supervisor's actions by displaying a concise summary of the important points and providing hints on what points the supervisor should pay attention to. If necessary, the supervisor can provide feedback or send a support message.
[0305] Specific examples
[0306] Example 1: Everyday dialogue and emotion recognition
[0307] Consider a scenario where a user says, "I have too many meetings this week and I'm not getting any work done."
[0308] The all-affirmation bot responds with, "That must have been tough, good job!" At this point, the emotion engine recognizes the strong sense of fatigue from the user's tone of voice and records it on the server.
[0309] The generative AI model summarizes that "work delays caused by numerous meetings are a source of stress and fatigue."
[0310] The server applies a social filter and shares the information with superiors, stating that "users are stressed by delays in work due to being busy, and feel very tired."
[0311] Example 2: Critical feedback and emotion recognition
[0312] Consider a situation where a user repeatedly says, "The project deadline is too tight."
[0313] The all-affirmation bot replies, "I understand the pressure. You're doing a great job!"
[0314] The emotion engine recognizes strong stress from the user's facial expressions and voice.
[0315] The server summarizes that "tight project deadlines are the main cause of stress" and immediately communicates this information to his superiors.
[0316] Based on the feedback, the supervisor can take action such as "reevaluating the project deadline."
[0317] Example prompts for generative AI models
[0318] (Example prompt 1: Summarizing everyday conversations)
[0319] User says: "I have too many meetings this week and I can't get any work done."
[0320] Recognizing the Emotion Engine: Extreme Fatigue
[0321] Summary generated: Work delays due to too many meetings are a major cause of stress and fatigue
[0322] (Example prompt 2: Summary of important feedback)
[0323] User says: "The project deadline is too tight"
[0324] Recognizing the Emotion Engine: High Stress
[0325] Generate summary: Tight project deadlines are a major source of stress
[0326] This system makes it possible to accurately grasp the user's emotional state and provide appropriate feedback, thereby improving the effectiveness of mental care.
[0327] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0328] Step 1:
[0329] The user launches the application and registers by entering their name, email address, and employee ID.
[0330] Input: Name, Email Address, Employee ID
[0331] How it works: The device sends this information to the server.
[0332] Output: User information is sent to the server and a registration completion message is received.
[0333] Step 2:
[0334] The server stores the received information in a database and sends a notification of registration completion to the terminal.
[0335] Input: User information (name, email address, employee ID)
[0336] What happens: The server creates a new record in the database and saves the information.
[0337] Output: A message that registration is complete is displayed on the terminal.
[0338] Step 3:
[0339] The user enters their email address and password and clicks the login button.
[0340] Input: Email address, Password
[0341] How it works: The device sends this information to the server.
[0342] Output: Login information sent to the server.
[0343] Step 4:
[0344] The server compares the entered information with a database and performs authentication.
[0345] Input: Email address, Password
[0346] How it works: The server checks the information in its database and performs authentication.
[0347] Output: If authentication is successful, information is sent to the device to display the user's dashboard.
[0348] Step 5:
[0349] The user clicks the "Start conversation" button on the dashboard screen.
[0350] Input: Click the "Start conversation" button
[0351] Action: The device will begin voice input and perform the initial setup to connect to the all-affirmation bot.
[0352] Output: Ready to connect with all affirmative bots.
[0353] Step 6:
[0354] When the user starts speaking, the device converts the speech into text and sends it to the All Affirmations Bot.
[0355] Input: User's voice data
[0356] Operation: The device performs voice recognition and sends the recognized text data to all affirmative bots.
[0357] Output: Text data sent to all affirmation bots.
[0358] Step 7:
[0359] An all-affirmation bot always responds affirmatively to what the user says.
[0360] Input: Text data based on speech recognition
[0361] How it works: The all-affirmation bot generates appropriate affirmative responses. For example, if someone says "I'm tired today," it will respond with "That must have been hard, good job!"
[0362] Output: Affirmative response to the user.
[0363] Step 8:
[0364] The device sends the user's voice and text data to the emotion engine.
[0365] Input: User voice and text data
[0366] How it works: The emotion engine analyzes this data and recognizes emotions.
[0367] Output: Recognized emotion information.
[0368] Step 9:
[0369] The emotion engine feeds back the analysis results to the dialogue agent and adjusts the response content.
[0370] Input: Emotion recognition results
[0371] How it works: The bot updates its responses based on the results of the emotion engine. For example, if it determines that you are tired, it will respond with something like, "Maybe it would be good to take a break."
[0372] Output: The adjusted response.
[0373] Step 10:
[0374] When the conversation ends, the terminal notifies the server that the call has ended.
[0375] Input: Call end event
[0376] Action: The device sends a call end notification to the server.
[0377] Output: A conversation termination notification is sent to the server.
[0378] Step 11:
[0379] The server sends the call logs and emotion recognition results to the natural language processing model.
[0380] Input: Call logs, emotion recognition results
[0381] How it works: The server sends this data to the natural language processing model and begins analysis.
[0382] Output: The data required for analysis is sent.
[0383] Step 12:
[0384] The generative AI model extracts important keywords and emotional trends from the dialogue and creates a summary.
[0385] Input: Call logs, emotion recognition results
[0386] How it works: The generative AI model analyzes data, extracts key keywords and sentiment trends, and generates summaries. For example, if a user says, "I'm having too many meetings this week, so I'm not making progress on my work," the model summarizes that "Work delays due to being too busy are causing stress."
[0387] Output: Summarized information.
[0388] Step 13:
[0389] The server applies a sociality filter to the generated summary to organize the content to be recognized.
[0390] Input: Summary information
[0391] How it works: The server applies social filters to determine what should be recognized.
[0392] Output: Filtered summary information.
[0393] Step 14:
[0394] The server prepares to send the filtered summary to the supervisor.
[0395] Input: Filtered summary information
[0396] What it does: The server formats the information for notification.
[0397] Output: Filtered summary for notifying superiors.
[0398] Step 15:
[0399] Managers can log in to a dedicated dashboard and check feedback on employees' mental health.
[0400] Input: Filtered summary information
[0401] What happens: Your manager accesses the dashboard and sees the feedback you provided.
[0402] Output: Your supervisor reviews the feedback and is ready to take action if necessary.
[0403] (Application example 2)
[0404] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0405] Conventional conversational agent systems have difficulty accurately grasping a user's emotional state and stress level, making it difficult to provide appropriate feedback. Furthermore, there is a demand for systems that enable sales staff, especially in brick-and-mortar stores, to grasp and respond to customer emotions in real time.
[0406] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means having a user interface for the user to interact with the dialogue agent, natural language processing means for analyzing the content of the dialogue between the dialogue agent and the user to determine the user's emotional state and stress level, filtering means for filtering and summarizing the content of the dialogue analyzed by the natural language processing means, notification means for notifying an administrator of the information processed by the filtering means, and device linking means for recognizing the user's emotional state in real time using smart glasses and displaying feedback. This enables sales staff to grasp customer emotions in real time and respond optimally.
[0407] A "user interface" is a screen or operating means that allows a user to interact with or operate a system.
[0408] A "dialogue agent" is software or a system that interacts with a user and supports communication using natural language processing.
[0409] "Natural language processing means" is a technology for analyzing the content of conversation between a dialogue agent and a user and understanding the user's emotions and intentions.
[0410] The "filtering means" is a technology for organizing information analyzed by the natural language processing means based on specific criteria and extracting necessary information.
[0411] The "notification means" is a function or device for transmitting the information sorted by the filtering means to the administrator.
[0412] "Smart glasses" are eyeglass-type wearable devices that have a display function and provide information to the user.
[0413] "Device integration means" is a technology that links devices such as smart glasses with systems to display and update information in real time.
[0414] "Administrator" is the person or department responsible for overseeing the status of the system and users and providing necessary support and feedback.
[0415] "Real time" refers to a state in which processing or communication is carried out almost immediately after the data is generated.
[0416] This invention is a system for users to receive mental health care through dialogue with a conversational agent. The system uses smart glasses to recognize the user's emotional state in real time and provide appropriate feedback.
[0417] Hardware and software used
[0418] The hardware used includes:
[0419] Smart glasses (e.g. Google Glass, Vuzix Blade)
[0420] Server (e.g. AWS EC2)
[0421] The software used includes:
[0422] Emotion recognition engine (e.g. Microsoft Azure Emotion API)
[0423] Natural language processing models (e.g., OpenAI GPT-4)
[0424] Front-end applications (e.g. React Native)
[0425] Database (e.g. MySQL)
[0426] Overall processing of the program
[0427] The system analyzes the user's emotional state and dialogue content in real time and provides optimal feedback.
[0428] 1. User Registration and Login
[0429] First, the user starts the application and registers by entering basic information on the user registration screen. The terminal sends the entered information to the server, which stores it in a database. After registration, the user can log in and access the system.
[0430] 2. Interaction with a conversational agent
[0431] The user wears the smart glasses and starts a dialogue with the conversational agent. The smart glasses capture the user's voice and facial expressions and send them to an emotion recognition engine. The emotion recognition engine analyzes the user's emotional state and sends it to the server in real time.
[0432] 3. Emotion Recognition and Feedback
[0433] The server then sends the data received from the emotion recognition engine to a natural language processing model to generate appropriate feedback, which is then displayed in real time on the smart glasses display for the user to review.
[0434] Specific examples
[0435] Example prompt sentence 1:
[0436] Generate a conversational agent response and emotion recognition results when a user says, "Tell me more about this product." The customer's expression is filled with curiosity.
[0437] Example system response:
[0438] Emotion recognition result: "Curiosity"
[0439] All-Affirmative bot replies: "I understand. The special feature of this product is that it is made using the latest technology and is extremely durable."
[0440] Example prompt sentence 2:
[0441] Generate a conversational agent response and emotion recognition results when a user says, "This product was a disappointment." The customer's tone sounds angry.
[0442] Example system response:
[0443] Emotion recognition result: "Anger"
[0444] The all-affirmation bot's response: "I'm sorry you felt that way. I'll get back to you shortly. Can you please provide more details?"
[0445] This system allows users to receive appropriate feedback in real time, allowing managers to accurately grasp the user's emotional state and respond promptly, which is expected to improve customer satisfaction in physical stores.
[0446] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0447] Step 1:
[0448] The user starts the application and registers by entering the required information such as name, email address, employee ID, etc. on the user registration screen. The entered information is sent from the device to the server, which stores it in a database. Once the user has completed registration, a login screen is displayed. Here, the user logs in by entering their email address and password. The server authenticates the user, and if authentication is successful, the user's dashboard is displayed.
[0449] Input: Registration information (name, email address, employee ID), login information (email address, password)
[0450] Output: User's dashboard
[0451] Step 2:
[0452] By clicking the "Start Dialogue" button on the dashboard screen, the user begins a dialogue with the conversational agent (all-affirmative bot). The device starts voice recognition, and the user begins speaking. The voice data is sent to the server, where it is input into the emotion recognition engine.
[0453] Input: User's voice data
[0454] Output: Input to the emotion recognition engine
[0455] Step 3:
[0456] The emotion recognition engine analyzes the user's voice tone, facial expressions, and text content to recognize the user's emotional state. The recognized emotional information is used as data for generating appropriate feedback through a natural language processing model.
[0457] Input: User's voice tone, facial expressions, and text content
[0458] Output: Recognized emotion information
[0459] Step 4:
[0460] The recognized emotional information and dialogue content are sent to a natural language processing model, which then uses this data to generate appropriate feedback for the user, which is then displayed on the smart glasses display in real time.
[0461] Input: Recognized emotion information, dialogue content
[0462] Output: Generated feedback
[0463] Step 5:
[0464] When the conversation ends, the device notifies the server of the end of the call. The server then analyzes the conversation using a natural language processing model based on the stored conversation log and emotion recognition results. The generated summary is then further processed by a filtering means to extract important information.
[0465] Input: Dialogue log, emotion recognition results
[0466] Output: A summary with the necessary information extracted
[0467] Step 6:
[0468] The filtered summary is sent to the administrator via a notification mechanism. The administrator can then log in to a dedicated dashboard and check the user's feedback on mental health care. The notification mechanism displays important summary parts according to the administrator's actions and also provides hints on points that require attention.
[0469] Input: Filtered summary
[0470] Output: Administrator notification, hints on what to note
[0471] Step 7:
[0472] Administrators can provide individual feedback based on the user's feedback or send messages to offer assistance, and this information is reflected in the system and used for future interactions.
[0473] Input: Admin feedback, support message
[0474] Output: Information reflected in subsequent interactions
[0475] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0476] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0477] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0478] [Second embodiment]
[0479] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0480] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0481] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0482] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0483] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0484] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0485] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0486] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0487] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0488] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0489] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0490] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0491] This invention is a dialogue agent system that uses an all-affirmation bot to provide mental care and communication support for employees. This system is implemented in the following steps.
[0492] System Configuration
[0493] This system includes users (employees), devices (such as PCs and smartphones), a server, and a conversational agent (an all-affirmation bot). These components work together to understand the user's mental state and provide appropriate feedback to their superiors.
[0494] User Registration and Login
[0495] The user first launches the application and registers by entering the required information such as name, email address, and employee ID on the user registration screen. The device sends the entered information to the server, which stores it in a database. After registration is complete, the user logs in by entering their email address and password on the login screen. The server authenticates the user, and if authentication is successful, displays the user's dashboard.
[0496] Conversation with the all-affirmation bot
[0497] The user clicks the "Start Dialogue" button on the dashboard screen to begin a dialogue with the conversational agent (all-affirmative bot). The device starts voice recognition, and the user begins speaking. The all-affirmative bot always returns a positive response to the user's statements. For example, if the user says, "I'm very tired today," the all-affirmative bot will reply, "That must have been tough. Good job!" The device sends this dialogue content to the server in real time.
[0498] Call analysis
[0499] When the conversation ends, the device notifies the server that the call has ended. The server then sends the saved conversation log to a natural language processing model and begins analyzing the text content. The generative AI model extracts important keywords and emotional trends from the conversation content and creates a summary. The server then applies a social filter to the generated summary to organize the content that should be recognized. For example, a user's statement that "I have too many meetings this week and my work isn't progressing" is summarized as "Work delays due to being busy are causing me stress."
[0500] Information Sharing and Notification
[0501] The server prepares the filtered summary to be sent to the supervisor. The supervisor logs in to a dedicated dashboard from their own device and checks the feedback on the employee's mental state. The server responds to the supervisor's actions by displaying a concise summary of the important parts and providing hints on what points to pay attention to. If necessary, the supervisor can provide individual feedback or send a message to offer support.
[0502] Specific examples
[0503] Example 1: Everyday conversation
[0504] The user schedules a conversation with the All-Affirmation Bot every day at 2:00 PM. Today, the user says, "I have too many meetings this week, so I can't get my work done." The All-Affirmation Bot responds, "That must have been tough, you've worked hard!" The server records the statement, "I have too many meetings this week, so I can't get my work done," and the generative AI model summarizes it as, "The delays in work caused by too many meetings are a stressful factor." The server applies a social filter to the summary and shares it with the user's superiors, concluding, "The user is stressed by delays in work caused by being busy."
[0505] Example 2: Important Feedback
[0506] In conversations with the All-Affirmation Bot, the user repeatedly states, "The project deadline is too tight." The All-Affirmation Bot responds, "I understand the pressure. You're doing a great job!" The server summarizes this as "The tight project deadline is the main cause of stress," and immediately notifies the supervisor. The supervisor can then take action based on the employee's feedback, such as "reevaluating the project deadline."
[0507] This allows the user to more easily reduce stress, and allows the supervisor to obtain information for effective mental care.
[0508] The processing flow will be explained below.
[0509] Step 1:
[0510] The user launches the application, enters the necessary information such as "name," "email address," and "employee ID" on the user registration screen, and presses the "Register" button.
[0511] Step 2:
[0512] The terminal temporarily stores the input user information and transmits the data to the server.
[0513] Step 3:
[0514] The server stores the received user information in a database and returns a response indicating successful registration to the terminal.
[0515] Step 4:
[0516] The terminal displays a "Registration Complete" message to the user.
[0517] Step 5:
[0518] The user enters their "email address" and "password" on the login screen and presses the "Login" button.
[0519] Step 6:
[0520] The terminal transmits the entered login information to the server.
[0521] Step 7:
[0522] The server authenticates the user using a database, and if authentication is successful, returns the user's dashboard information to the terminal.
[0523] Step 8:
[0524] The terminal displays a dashboard screen to the user.
[0525] Step 9:
[0526] The user clicks the "Start conversation" button on the dashboard screen.
[0527] Step 10:
[0528] The terminal starts voice recognition and the user begins speaking.
[0529] Step 11:
[0530] The conversational agent (all-affirmative bot) responds affirmatively to user comments. For example, if the user says, "I'm very tired today," the agent responds, "That must have been hard! You did a great job!"
[0531] Step 12:
[0532] The terminal transmits the content of the conversation between the all-affirming bot and the user to the server in real time.
[0533] Step 13:
[0534] The server stores the content of the conversation as a log and performs processing such as sentiment analysis and keyword extraction in real time as needed.
[0535] Step 14:
[0536] When the interaction ends, the terminal notifies the server of the interaction end.
[0537] Step 15:
[0538] The server sends the saved dialogue log to a natural language processing model to begin analyzing the text content.
[0539] Step 16:
[0540] The generative AI model extracts important keywords and sentiment trends from the conversation and creates a summary, such as, "The user has been busy this week, and is feeling particularly stressed by the large number of meetings."
[0541] Step 17:
[0542] The server applies a social filter to the generated summary to organize the content that should be recognized, for example, summarizing it as "Work delays due to busyness are a cause of stress."
[0543] Step 18:
[0544] The server prepares to send the filtered summary to the superior.
[0545] Step 19:
[0546] Supervisors can log into a dedicated dashboard from their own devices and check feedback on employees' mental health.
[0547] Step 20:
[0548] The server displays a concise summary of important points in response to the supervisor's operation, and also provides hints on what points to pay attention to.
[0549] Step 21:
[0550] If needed, a manager can send a message to provide individual feedback or assistance.
[0551] The above are the specific processing steps for a mental care and communication assistance application using a fully affirmative bot.
[0552] Example 1
[0553] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0554] In modern corporate environments, employee mental health is an important issue, but many systems lack the means to properly understand users' emotional states and stress levels, or to provide appropriate feedback to superiors. Therefore, a method for effectively managing employee mental health is needed. Furthermore, conventional conversational agents may respond negatively to user comments, which may increase employee stress.
[0555] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0556] In this invention, the server includes: a user interface that provides a method for a user to interact with a dialogue agent; a user interface that allows the user to initiate a dialogue with the dialogue agent and record the dialogue content; a generative AI model that allows the dialogue agent to always generate positive responses; a natural language processing unit that analyzes the dialogue content between the dialogue agent and the user to determine the user's emotional state and level of stress; a summarizing unit that summarizes the dialogue content analyzed by the natural language processing unit and extracts important keywords and emotional trends; a notification unit that notifies the supervisor of the information processed by the filtering unit; and a notification unit that provides the supervisor with feedback regarding the user's mental care and encourages the supervisor to take appropriate measures. This makes it possible to accurately grasp the employee's emotional state and level of stress and provide appropriate feedback to the supervisor.
[0557] The "user interface" refers to the operation screen and input means by which the user operates the system and starts a dialogue with the dialogue agent.
[0558] A "generative AI model" is an algorithm or program that generates positive responses in natural language based on what a user says.
[0559] "Natural language processing" is a technology that analyzes text data and extracts important keywords and emotional trends from its content.
[0560] A "dialogue agent" is a program or system that interacts with a user and generates responses to their utterances.
[0561] "Notification means" refers to the communication means or process for sending analyzed information and feedback to superiors.
[0562] The "filtering means" is a processing means for organizing the dialogue content analyzed by natural language processing and generating an appropriate summary.
[0563] "Emotional state" refers to the user's psychological state, and indicates stress, satisfaction, fatigue, etc.
[0564] The "dashboard" is an operation screen accessed by users who log in to the system, and has a central function for starting a dialogue and checking feedback.
[0565] This invention is a dialogue agent system for the purpose of providing mental care and communication support for employees, and includes a user (employee), a terminal (such as a PC or smartphone), a server, and a dialogue agent (an all-affirmation bot). These components work together to grasp the user's mental state and provide appropriate feedback to superiors.
[0566] System Configuration
[0567] The system consists of the following main components:
[0568] User Interface
[0569] The user uses this interface to start a dialogue with the dialogue agent, inputting and operating the agent according to the situation. The user interface is provided as a software application that runs on a device such as a PC or smartphone.
[0570] Generative AI Models
[0571] The conversational agent uses a generative AI model (e.g., ChatGPT) that generates positive responses based on user utterances. This model uses natural language processing techniques to understand what the user is saying and always generates a positive response.
[0572] Natural Language Processing
[0573] On the server side, natural language processing technology (e.g., BERT) is used to analyze the content of the dialogue between the conversation agent and the user, extracting important keywords and emotional trends from the user's speech and determining the user's emotional state and level of stress.
[0574] Notification means
[0575] The information processed by the filtering means is notified to the superior by the server. The notification means ensures that the data is correctly filtered and that the superior is provided with important information along with a concise summary.
[0576] User Registration and Login
[0577] The user starts the application and enters the required information such as name, email address, employee ID, etc. on the user registration screen. The device sends this information to the server, which stores it in a database. When the user enters their email address and password on the login screen, the server authenticates the user and, if successful, displays the dashboard.
[0578] Conversation with the all-affirmation bot
[0579] When the user clicks the "Start Dialogue" button on the dashboard, a dialogue with the conversational agent (all-positive bot) begins. The device starts voice recognition and captures the user's voice input. The server converts the received voice data into text and sends it to the all-positive bot. The all-positive bot uses a generative AI model to generate a positive response, and the device presents this response to the user.
[0580] Call analysis and summarization
[0581] When the dialogue ends, the device notifies the server. The server then sends the saved dialogue log to a natural language processing model, which begins analyzing the text content. The generative AI model extracts important keywords and emotional trends from the dialogue and generates a summary. For example, if a user says, "I have too many meetings this week, so I'm not making much progress on my work," the summary might read, "Work delays due to being too busy are causing me stress."
[0582] Information Sharing and Notification
[0583] The server applies a social filter to the generated summary, organizes important information, and notifies the supervisor. The supervisor logs in to a dedicated dashboard and checks the feedback. This allows the supervisor to take appropriate action based on the employee's mental state.
[0584] Specific prompt examples
[0585] Below are some examples of prompt sentences.
[0586] "Users interact with the all-affirmation bot and send what they say to your server. Your natural language processing model analyzes the text and creates a summary."
[0587] This allows the system to effectively support employees' mental health care and provide supervisors with the information they need to provide appropriate feedback.
[0588] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0589] Step 1: User Registration
[0590] A user starts an application and enters the required information such as name, email address, and employee ID on the user registration screen. The entered information (name, email address, employee ID) is input data. The device sends this input data to the server, which receives it and stores it in a database. Specifically, the server receives an HTTP POST request and registers it as a new record in the database (e.g., MySQL). The output is that the new user has been registered and the information is stored in the database.
[0591] Step 2: User Login
[0592] A user attempts to log in by entering an email address and password. The device sends this authentication information (email address, password) to the server. The server searches for the corresponding user information in the database and compares it with the input data. If authentication is successful, the server returns an authentication success message and the user's dashboard information to the device as output. The device receives this information and displays the user's dashboard.
[0593] Step 3: Initiating a conversation
[0594] The user clicks the "Start conversation" button on the dashboard. This action is registered as input data. The device starts voice input and records what the user says. The recorded voice data is input and the device sends it to the server. The server receives the voice data and converts it into text data using voice recognition software (e.g., Google Cloud Speech-to-Text API). The converted text data is output.
[0595] Step 4: Generate a response
[0596] The server inputs the received text data into a generative AI model (e.g., ChatGPT) to generate a positive response. The generative AI model performs natural language processing based on the input data (user utterances) to generate a positive response. This response text becomes the output. The server sends this response text to the device, which then presents the response to the user in voice or text.
[0597] Step 5: Record the conversation
[0598] The device sends the content of the conversation (user's comments and all affirmative bot's responses) to the server in real time. The server receives this and stores it in a database. The saved conversation log is the output. This log is used for later analysis.
[0599] Step 6: End of conversation
[0600] When the conversation ends, the device notifies the server. This notification becomes the input data. The server acquires the conversation log and sends it to a natural language processing model (e.g., BERT). This starts the analysis of the text content.
[0601] Step 7: Text analysis and summary generation
[0602] The server uses a natural language processing model to extract important keywords and emotional trends from the dialogue. This analysis is a data calculation based on the input data (dialogue log). Summary text is generated as the analysis result. This summary text is the output data. For example, a statement such as "There are too many meetings this week, so work isn't progressing" is summarized as "Work delays due to being busy are causing stress."
[0603] Step 8: Filtering and Notifications
[0604] The server applies a social filter to the generated summary text and organizes important information. The filtered text is output. This text is sent to the superior via a notification method. The superior receives the notification and checks the feedback on a dedicated dashboard. For example, the superior may receive specific feedback such as, "Work delays due to busyness are the cause of stress."
[0605] (Application example 1)
[0606] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0607] In conventional factory work, operators and engineers often suffer from excessive stress and fatigue, which has a negative impact on productivity and quality. In particular, a lack of psychological care can lead to problems such as a decrease in motivation and an increase in turnover. Therefore, a new system is needed to efficiently provide mental care for operators and engineers working in factories and improve the quality of their work.
[0608] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0609] In this invention, the server includes: means for having a user interface that provides a method for a user to interact with a dialogue agent; natural language processing means for analyzing the dialogue content between the dialogue agent and the user to determine the user's emotional state and level of stress; filtering means for filtering and summarizing the dialogue content analyzed by the natural language processing means; notification means for notifying a supervisor of the information processed by the filtering means; means for using a generative AI model that recognizes the user's everyday conversation and generates a positive response; speech recognition means for converting speech input from the user into text in real time; means for executing a program that controls the notification means to summarize and notify the dialogue content; and feedback means for the supervisor to take an appropriate action based on the notified information. This makes it possible to reduce the psychological stress of operators and engineers working in a factory and provide effective mental care.
[0610] A "user interface" is the means of interaction that allows a user to interact with a system or device.
[0611] A "dialogue agent" is a software agent that obtains information through dialogue with a user and returns an appropriate response.
[0612] "Natural language processing means" refers to algorithms or systems that analyze natural language text spoken by a user and determine their emotional state and level of stress.
[0613] The "filtering means" is a process for organizing the dialogue content analyzed by the natural language processing means, and extracting and summarizing important information.
[0614] "Notification means" refers to a function or system for notifying the processed filtering results to superiors or managers.
[0615] A "generative AI model" is an artificial intelligence model that generates appropriate responses to user input.
[0616] "Speech recognition means" refers to technology or devices that convert a user's speech into text in real time.
[0617] The "means for executing a program" refers to the software and hardware configuration for collecting, analyzing, and filtering the contents of the conversation and controlling the notification means.
[0618] "Feedback means" refers to a specific method or system that allows a superior to take appropriate action based on the notified information.
[0619] This invention is a dialogue agent system for providing mental care to operators and engineers working in factories. This system includes users (operators and engineers), terminals (smartphones, tablets, etc.), a server, and a dialogue agent (an all-affirmative bot).
[0620] System Configuration
[0621] The system works by coordinating the following components:
[0622] 1. User Interface: Provides an interface for the user to interact with the conversational agent, which can be operated by a touch screen or voice input.
[0623] 2. Conversational Agent: Engages in natural dialogue with the user. Conversational agents use generative AI models to always generate positive responses to user utterances.
[0624] 3. Speech recognition: Converts voice input into text in real time. This function uses the Google Speech-to-Text API, for example.
[0625] 4. Natural language processing: Analyzes the dialogue and determines the user's emotional state and stress level. This process uses natural language processing technologies such as GPT-3.5.
[0626] 5. Filtering means: Filters the dialogue content analyzed by the natural language processing means and summarizes important information.
[0627] 6. Notification: The information summarized by the filtering method is notified to superiors and managers via email or dashboard.
[0628] 7. Feedback measures: Provide feedback to supervisors and managers to take appropriate action based on the information provided.
[0629] Program processing
[0630] The server receives the dialogue between the conversational agent and the user in real time and generates a positive response using a generative AI model. The use of GPT-3.5 as the generative AI model allows for natural dialogue. All dialogue is saved in text format and later analyzed using natural language processing.
[0631] The device receives voice input from the user and converts it into text using a speech recognition method (e.g., Google Speech-to-Text API), which then transmits the user's speech as text to the server in real time.
[0632] The filtering means processes the acquired text data and extracts important keywords and phrases, and based on this information, summarizes the conversation and notifies superiors or managers.
[0633] The notification method sends filtered summary information to superiors and managers via email or a dedicated dashboard, allowing them to quickly take the necessary measures to care for the user's mental health.
[0634] Specific examples
[0635] For example, if a factory operator uses the system and says, "I'm so busy today, I'm tired," the all-affirmation bot will respond, "That must have been tough, good job!" After the conversation ends, the system summarizes the content of the conversation and provides feedback to the supervisor or manager that "the operator's fatigue due to being too busy is the cause."
[0636] Prompt Sentence Examples
[0637] Input prompt: "I'm too busy and tired today."
[0638] Model response: "That was hard work, good job!"
[0639] By implementing this system, it is possible to improve the working environment within the factory, maintain the motivation of operators and engineers, and reduce stress.
[0640] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0641] Step 1:
[0642] The user performs voice input. Specifically, the user speaks into the device's microphone to start a conversation. The input is what the user says, and is captured by the device as voice data.
[0643] Step 2:
[0644] A speech recognition means is activated on the terminal. The speech recognition means converts the user's voice data into text data. This process uses speech recognition technology such as the Google Speech-to-Text API. The output is data that converts the user's speech into text.
[0645] Step 3:
[0646] The device sends text data to the server, which then passes it to the generative AI model. The input is the converted text data, which is then sent to the server.
[0647] Step 4:
[0648] The conversational agent in the server uses a generative AI model to generate a positive response to the user's utterance. Specifically, a GPT-3.5 model is used to generate a prompt sentence based on the input text. The output is a positive response text.
[0649] Step 5:
[0650] The generated positive response text is returned from the server to the terminal. The terminal reproduces this text data as voice through the voice output means. The input is the generated response text, and the output is a voice response to the user.
[0651] Step 6:
[0652] The server stores the dialogue content as a log, and natural language processing means analyzes the text data. In this analysis process, the user's emotional state and stress level are determined. The input is the user's entire dialogue text, and the output is the evaluation results of the emotional state and stress.
[0653] Step 7:
[0654] The filtering means extracts important keywords and phrases from the dialogue text and generates a summary. In this step, points of particular interest are organized from the dialogue content. The input is the analysis result by the natural language processing means, and the output is the summary text.
[0655] Step 8:
[0656] The summarized information is sent to a supervisor or manager through a notification mechanism, which can be email or a dashboard. The input is the summary text and the output is the notification message provided to the supervisor or manager.
[0657] Step 9:
[0658] The supervisor or manager takes appropriate action based on the notified information. This involves implementing specific measures and countermeasures using feedback methods. The input is the notification message, and the output is specific feedback and countermeasures.
[0659] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0660] The present invention is a dialogue agent system that utilizes an all-affirmation bot, with the aim of providing mental care and communication assistance to users. By incorporating an emotion engine into this system, it is possible to more accurately recognize the user's emotional state and provide appropriate feedback. Specific embodiments of this system are described below.
[0661] System Configuration
[0662] This system includes users (employees), devices (PCs, smartphones, etc.), a server, a dialogue agent (a completely affirmative bot), and an emotion engine. These components work together to understand the user's mental state and provide appropriate feedback to superiors.
[0663] User Registration and Login
[0664] The user first launches the application and registers by entering the required information such as name, email address, and employee ID on the user registration screen. The device sends the entered information to the server, which stores it in a database. After registration is complete, the user logs in by entering their email address and password on the login screen. The server authenticates the user, and if authentication is successful, displays the user's dashboard.
[0665] Conversation with an all-affirmation bot and emotion recognition
[0666] The user clicks the "Start conversation" button on the dashboard screen to begin a conversation with the conversational agent (all-affirmative bot). The device starts voice recognition, and the user begins speaking. The all-affirmative bot always returns a positive response to the user's statements. For example, if the user says, "I'm very tired today," the bot replies, "That must have been tough! You've done well!"
[0667] The emotion engine analyzes the user's voice tone, facial expressions, and text content to recognize their emotions. The recognized emotional information is reflected in the conversational agent's responses. For example, if the user is determined to be very tired, the all-affirmation bot will adjust its responses to be more warmhearted.
[0668] Analysis of call content and emotional information
[0669] When the conversation ends, the device notifies the server that the call has ended. The server then sends the saved conversation log and emotion recognition results to a natural language processing model, which begins analyzing the text content and emotional information. The generative AI model extracts important keywords and emotional trends from the conversation content and creates a summary. The server then applies a social filter to the generated summary to organize the content that should be recognized. For example, information such as "There are too many meetings this week, so work is not progressing" and "The user is feeling tired" is summarized as "Work delays due to busyness are the cause of stress."
[0670] Information Sharing and Notification
[0671] The server prepares the filtered summary to be sent to the supervisor. The supervisor logs in to a dedicated dashboard from their own device and checks the feedback on the employee's mental state. The server responds to the supervisor's actions by displaying a concise summary of the important parts and providing hints on what points to pay attention to. If necessary, the supervisor can provide individual feedback or send a message to offer support.
[0672] Specific examples
[0673] Example 1: Everyday dialogue and emotion recognition
[0674] The user schedules a conversation with the All-Affirmation Bot every day at 2:00 PM. Today, the user says, "There are too many meetings this week, and I'm not making progress on my work." The All-Affirmation Bot replies, "That must have been tough, you've worked hard!" At this point, the emotion engine recognizes a strong sense of fatigue from the user's tone of voice. The server records the statement, "There are too many meetings this week, and I'm not making progress on my work," along with the "strong sense of fatigue," and the generative AI model summarizes it as, "The delays in work due to the many meetings are a stressful factor and cause fatigue." The server applies a sociality filter to the above summary and shares it with the supervisor, concluding, "The user is experiencing stress due to work delays caused by being busy, and is feeling very fatigued."
[0675] Example 2: Critical feedback and emotion recognition
[0676] In conversation with the All-Affirmation Bot, the user repeatedly states, "The project deadline is too tight." The All-Affirmation Bot responds, "I understand the pressure. You're doing a great job!" At this point, the emotion engine recognizes the user's strong stress from their facial expressions and voice. The server summarizes this as "The tight project deadline is the main cause of stress," and immediately notifies this information to the supervisor. The supervisor can then take action based on the employee's feedback, such as "reevaluating the project deadline."
[0677] This makes it easier for users to reduce stress and provides supervisors with information to provide effective mental care.The introduction of the emotion engine allows for a more accurate understanding of the user's emotional state, improving the quality of responses and feedback.
[0678] The processing flow will be explained below.
[0679] Step 1:
[0680] The user launches the application, enters the necessary information such as "name," "email address," and "employee ID" on the user registration screen, and presses the "Register" button.
[0681] Step 2:
[0682] The terminal temporarily stores the input user information and transmits the data to the server.
[0683] Step 3:
[0684] The server stores the received user information in a database and returns a response indicating successful registration to the terminal.
[0685] Step 4:
[0686] The terminal displays a "Registration Complete" message to the user.
[0687] Step 5:
[0688] The user enters their "email address" and "password" on the login screen and presses the "Login" button.
[0689] Step 6:
[0690] The terminal transmits the entered login information to the server.
[0691] Step 7:
[0692] The server authenticates the user using a database, and if authentication is successful, returns the user's dashboard information to the terminal.
[0693] Step 8:
[0694] The terminal displays a dashboard screen to the user.
[0695] Step 9:
[0696] The user clicks the "Start conversation" button on the dashboard screen.
[0697] Step 10:
[0698] The terminal starts voice recognition and the user begins speaking.
[0699] Step 11:
[0700] The conversational agent (all-affirmative bot) responds affirmatively to user comments. For example, if the user says, "I'm very tired today," the agent responds, "That must have been hard! You did a great job!"
[0701] Step 12:
[0702] The emotion engine recognizes the user's emotions by analyzing their voice tone, facial expressions, and text content, for example, reading tiredness from their voice and sadness from their facial expressions.
[0703] Step 13:
[0704] The terminal transmits the content of the conversation between the all-affirmation bot and the user, as well as the emotional data obtained from the emotion engine, to the server in real time.
[0705] Step 14:
[0706] The server stores the dialogue content and emotional data as logs, and performs real-time processing such as emotion analysis and keyword extraction as needed.
[0707] Step 15:
[0708] When the interaction ends, the terminal notifies the server of the interaction end.
[0709] Step 16:
[0710] The server sends the saved dialogue logs and emotion data to a natural language processing model, which begins analyzing the text content and emotion information.
[0711] Step 17:
[0712] The generative AI model extracts important keywords and emotional trends from the conversation and creates a summary, such as, "The user has been busy this week, and is feeling stressed and tired, especially with so many meetings."
[0713] Step 18:
[0714] The server applies a social filter to the generated summary to organize the content that should be recognized, for example, summarizing it as "Work delays due to busyness are a cause of stress."
[0715] Step 19:
[0716] The server prepares to send the filtered summary to the superior.
[0717] Step 20:
[0718] Supervisors can log into a dedicated dashboard from their own devices and check feedback on employees' mental health.
[0719] Step 21:
[0720] The server displays a concise summary of important points in response to the supervisor's operation, and also provides hints on what points to pay attention to.
[0721] Step 22:
[0722] If needed, a manager can send a message to provide individual feedback or assistance.
[0723] Step 23:
[0724] Users receive feedback from their superiors and implement advice on mental care and work improvement.
[0725] The above are the specific processing steps for a mental care and communication assistance application using a fully affirmative bot combined with an emotion engine.
[0726] Example 2
[0727] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0728] To effectively provide mental care to users, it is necessary to accurately grasp the user's emotional state and stress level and provide appropriate feedback accordingly. However, conventional systems lack the means to accurately recognize the user's emotions, making it difficult to provide appropriate responses and feedback. In addition, sharing information with superiors is cumbersome, making it difficult to provide effective mental care to users.
[0729] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0730] In this invention, the server includes means for having a user interface through which the user interacts with the dialogue agent, natural language processing means for analyzing the dialogue between the dialogue agent and the user to determine the user's emotional state and level of stress, filtering means for filtering and summarizing the dialogue analyzed by the natural language processing means, emotion recognition means for analyzing the user's voice tone, facial expressions, and text content to recognize the user's emotions, adjustment means for reflecting the emotion information obtained by the emotion recognition means in the dialogue agent's response content, and notification means for notifying a superior of the information processed by the filtering means. This makes it possible to accurately grasp the user's emotional state and provide appropriate feedback.
[0731] Below are definitions of important terms included in the rewritten claims:
[0732] "User interface" refers to the screen or means by which a user accesses and operates a system.
[0733] An "interactive agent" is a program or system that interacts with a user and responds based on the interaction.
[0734] "Natural language processing means" refers to a technique or means for analyzing the content of a dialogue between a user and a dialogue agent and determining the user's emotional state and stress level.
[0735] "Filtering means" refers to a technology or means for sorting the dialogue content analyzed by natural language processing means and extracting and summarizing important information.
[0736] The "emotion recognition means" refers to a technology or means for analyzing the user's voice tone, facial expression, and text content to recognize the user's emotions.
[0737] The "adjustment means" refers to a technique or means for appropriately changing or adjusting the response content of the dialogue agent based on the emotional information obtained by the emotion recognition means.
[0738] "Notification means" refers to a technique or means for transmitting the information processed by the filtering means to a superior in an appropriate format.
[0739] This invention is a dialogue agent system that uses an all-affirmation bot and an emotion engine, with the aim of providing mental care and communication assistance to users. Specific embodiments of this system are described below.
[0740] System Configuration
[0741] This system includes a user, a terminal, a server, a dialogue agent (a fully affirmative bot), and an emotion engine. These elements work together to understand the user's mental state and provide appropriate feedback.
[0742] User Registration and Login
[0743] 1. The user launches the application and registers by entering information such as their name, email address, and employee ID.
[0744] 2. The terminal sends the entered information to the server, which stores it in a database.
[0745] 3. After completing registration, the user logs in by entering their email address and password.
[0746] 4. The server authenticates the user and, if authentication is successful, displays the user's dashboard.
[0747] Conversation with the all-affirmation bot
[0748] 1. The user clicks the "Start conversation" button on the dashboard screen to begin a conversation with the all-affirmation bot.
[0749] 2. The device starts voice recognition and sends the user's voice to the all-affirmation bot. For example, if the user says, "I'm tired today," the all-affirmation bot replies, "That must have been hard! Good job!"
[0750] emotion recognition
[0751] 1. The device simultaneously transmits the user's voice tone, facial expression, and text content to the emotion engine.
[0752] 2. The emotion engine analyzes this data and recognizes the user's emotions.
[0753] 3. The recognized emotional information is reflected in the responses of the all-positive bot. For example, if the user is judged to be very tired, the all-positive bot will adjust its responses to be more warmhearted.
[0754] Call analysis
[0755] 1. When the conversation ends, the terminal notifies the server that the call has ended.
[0756] 2. The server sends the dialogue log and emotion recognition results to the natural language processing model.
[0757] 3. The generative AI model extracts important keywords and emotional trends from the conversation and creates a summary. For example, the information that "the user is feeling tired" and the statement that "there are too many meetings this week and work isn't progressing" can be summarized as "work delays due to being busy are the cause of stress."
[0758] Information Sharing and Notification
[0759] 1. The server prepares the filtered summary for notification to the supervisor.
[0760] 2. Supervisors log in to a dedicated dashboard and check feedback on employees' mental health.
[0761] 3. The server responds to the supervisor's actions by displaying a concise summary of the important points and providing hints on what points the supervisor should pay attention to. If necessary, the supervisor can provide feedback or send a support message.
[0762] Specific examples
[0763] Example 1: Everyday dialogue and emotion recognition
[0764] Consider a scenario where a user says, "I have too many meetings this week and I'm not getting any work done."
[0765] The all-affirmation bot responds with, "That must have been tough, good job!" At this point, the emotion engine recognizes the strong sense of fatigue from the user's tone of voice and records it on the server.
[0766] The generative AI model summarizes that "work delays caused by numerous meetings are a source of stress and fatigue."
[0767] The server applies a social filter and shares the information with superiors, stating that "users are stressed by delays in work due to being busy, and feel very tired."
[0768] Example 2: Critical feedback and emotion recognition
[0769] Consider a situation where a user repeatedly says, "The project deadline is too tight."
[0770] The all-affirmation bot replies, "I understand the pressure. You're doing a great job!"
[0771] The emotion engine recognizes strong stress from the user's facial expressions and voice.
[0772] The server summarizes that "tight project deadlines are the main cause of stress" and immediately communicates this information to his superiors.
[0773] Based on the feedback, the supervisor can take action such as "reevaluating the project deadline."
[0774] Example prompts for generative AI models
[0775] (Example prompt 1: Summarizing everyday conversations)
[0776] User says: "I have too many meetings this week and I can't get any work done."
[0777] Recognizing the Emotion Engine: Extreme Fatigue
[0778] Summary generated: Work delays due to too many meetings are a major cause of stress and fatigue
[0779] (Example prompt 2: Summary of important feedback)
[0780] User says: "The project deadline is too tight"
[0781] Recognizing the Emotion Engine: High Stress
[0782] Generate summary: Tight project deadlines are a major source of stress
[0783] This system makes it possible to accurately grasp the user's emotional state and provide appropriate feedback, thereby improving the effectiveness of mental care.
[0784] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0785] Step 1:
[0786] The user launches the application and registers by entering their name, email address, and employee ID.
[0787] Input: Name, Email Address, Employee ID
[0788] How it works: The device sends this information to the server.
[0789] Output: User information is sent to the server and a registration completion message is received.
[0790] Step 2:
[0791] The server stores the received information in a database and sends a notification of registration completion to the terminal.
[0792] Input: User information (name, email address, employee ID)
[0793] What happens: The server creates a new record in the database and saves the information.
[0794] Output: A message that registration is complete is displayed on the terminal.
[0795] Step 3:
[0796] The user enters their email address and password and clicks the login button.
[0797] Input: Email address, Password
[0798] How it works: The device sends this information to the server.
[0799] Output: Login information sent to the server.
[0800] Step 4:
[0801] The server compares the entered information with a database and performs authentication.
[0802] Input: Email address, Password
[0803] How it works: The server checks the information in its database and performs authentication.
[0804] Output: If authentication is successful, information is sent to the device to display the user's dashboard.
[0805] Step 5:
[0806] The user clicks the "Start conversation" button on the dashboard screen.
[0807] Input: Click the "Start conversation" button
[0808] Action: The device will begin voice input and perform the initial setup to connect to the all-affirmation bot.
[0809] Output: Ready to connect with all affirmative bots.
[0810] Step 6:
[0811] When the user starts speaking, the device converts the speech into text and sends it to the All Affirmations Bot.
[0812] Input: User's voice data
[0813] Operation: The device performs voice recognition and sends the recognized text data to all affirmative bots.
[0814] Output: Text data sent to all affirmation bots.
[0815] Step 7:
[0816] An all-affirmation bot always responds affirmatively to what the user says.
[0817] Input: Text data based on speech recognition
[0818] How it works: The all-affirmation bot generates appropriate affirmative responses. For example, if someone says "I'm tired today," it will respond with "That must have been hard, good job!"
[0819] Output: Affirmative response to the user.
[0820] Step 8:
[0821] The device sends the user's voice and text data to the emotion engine.
[0822] Input: User voice and text data
[0823] How it works: The emotion engine analyzes this data and recognizes emotions.
[0824] Output: Recognized emotion information.
[0825] Step 9:
[0826] The emotion engine feeds back the analysis results to the dialogue agent and adjusts the response content.
[0827] Input: Emotion recognition results
[0828] How it works: The bot updates its responses based on the results of the emotion engine. For example, if it determines that you are tired, it will respond with something like, "Maybe it would be good to take a break."
[0829] Output: The adjusted response.
[0830] Step 10:
[0831] When the conversation ends, the terminal notifies the server that the call has ended.
[0832] Input: Call end event
[0833] Action: The device sends a call end notification to the server.
[0834] Output: A conversation termination notification is sent to the server.
[0835] Step 11:
[0836] The server sends the call logs and emotion recognition results to the natural language processing model.
[0837] Input: Call logs, emotion recognition results
[0838] How it works: The server sends this data to the natural language processing model and begins analysis.
[0839] Output: The data required for analysis is sent.
[0840] Step 12:
[0841] The generative AI model extracts important keywords and emotional trends from the dialogue and creates a summary.
[0842] Input: Call logs, emotion recognition results
[0843] How it works: The generative AI model analyzes data, extracts key keywords and sentiment trends, and generates summaries. For example, if a user says, "I'm having too many meetings this week, so I'm not making progress on my work," the model summarizes that "Work delays due to being too busy are causing stress."
[0844] Output: Summarized information.
[0845] Step 13:
[0846] The server applies a sociality filter to the generated summary to organize the content to be recognized.
[0847] Input: Summary information
[0848] How it works: The server applies social filters to determine what should be recognized.
[0849] Output: Filtered summary information.
[0850] Step 14:
[0851] The server prepares to send the filtered summary to the supervisor.
[0852] Input: Filtered summary information
[0853] What it does: The server formats the information for notification.
[0854] Output: Filtered summary for notifying superiors.
[0855] Step 15:
[0856] Managers can log in to a dedicated dashboard and check feedback on employees' mental health.
[0857] Input: Filtered summary information
[0858] What happens: Your manager accesses the dashboard and sees the feedback you provided.
[0859] Output: Your supervisor reviews the feedback and is ready to take action if necessary.
[0860] (Application example 2)
[0861] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0862] Conventional conversational agent systems have difficulty accurately grasping a user's emotional state and stress level, making it difficult to provide appropriate feedback. Furthermore, there is a demand for systems that enable sales staff, especially in brick-and-mortar stores, to grasp and respond to customer emotions in real time.
[0863] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means having a user interface for the user to interact with the dialogue agent, natural language processing means for analyzing the content of the dialogue between the dialogue agent and the user to determine the user's emotional state and stress level, filtering means for filtering and summarizing the content of the dialogue analyzed by the natural language processing means, notification means for notifying an administrator of the information processed by the filtering means, and device linking means for recognizing the user's emotional state in real time using smart glasses and displaying feedback. This enables sales staff to grasp customer emotions in real time and respond optimally.
[0864] A "user interface" is a screen or operating means that allows a user to interact with or operate a system.
[0865] A "dialogue agent" is software or a system that interacts with a user and supports communication using natural language processing.
[0866] "Natural language processing means" is a technology for analyzing the content of conversation between a dialogue agent and a user and understanding the user's emotions and intentions.
[0867] The "filtering means" is a technology for organizing information analyzed by the natural language processing means based on specific criteria and extracting necessary information.
[0868] The "notification means" is a function or device for transmitting the information sorted by the filtering means to the administrator.
[0869] "Smart glasses" are eyeglass-type wearable devices that have a display function and provide information to the user.
[0870] "Device integration means" is a technology that links devices such as smart glasses with systems to display and update information in real time.
[0871] "Administrator" is the person or department responsible for overseeing the status of the system and users and providing necessary support and feedback.
[0872] "Real time" refers to a state in which processing or communication is carried out almost immediately after the data is generated.
[0873] This invention is a system for users to receive mental health care through dialogue with a conversational agent. The system uses smart glasses to recognize the user's emotional state in real time and provide appropriate feedback.
[0874] Hardware and software used
[0875] The hardware used includes:
[0876] Smart glasses (e.g. Google Glass, Vuzix Blade)
[0877] Server (e.g. AWS EC2)
[0878] The software used includes:
[0879] Emotion recognition engine (e.g. Microsoft Azure Emotion API)
[0880] Natural language processing models (e.g., OpenAI GPT-4)
[0881] Front-end applications (e.g. React Native)
[0882] Database (e.g. MySQL)
[0883] Overall processing of the program
[0884] The system analyzes the user's emotional state and dialogue content in real time and provides optimal feedback.
[0885] 1. User Registration and Login
[0886] First, the user starts the application and registers by entering basic information on the user registration screen. The terminal sends the entered information to the server, which stores it in a database. After registration, the user can log in and access the system.
[0887] 2. Interaction with a conversational agent
[0888] The user wears the smart glasses and starts a dialogue with the conversational agent. The smart glasses capture the user's voice and facial expressions and send them to an emotion recognition engine. The emotion recognition engine analyzes the user's emotional state and sends it to the server in real time.
[0889] 3. Emotion Recognition and Feedback
[0890] The server then sends the data received from the emotion recognition engine to a natural language processing model to generate appropriate feedback, which is then displayed in real time on the smart glasses display for the user to review.
[0891] Specific examples
[0892] Example prompt sentence 1:
[0893] Generate a conversational agent response and emotion recognition results when a user says, "Tell me more about this product." The customer's expression is filled with curiosity.
[0894] Example system response:
[0895] Emotion recognition result: "Curiosity"
[0896] All-Affirmative bot replies: "I understand. The special feature of this product is that it is made using the latest technology and is extremely durable."
[0897] Example prompt sentence 2:
[0898] Generate a conversational agent response and emotion recognition results when a user says, "This product was a disappointment." The customer's tone sounds angry.
[0899] Example system response:
[0900] Emotion recognition result: "Anger"
[0901] The all-affirmation bot's response: "I'm sorry you felt that way. I'll get back to you shortly. Can you please provide more details?"
[0902] This system allows users to receive appropriate feedback in real time, allowing managers to accurately grasp the user's emotional state and respond promptly, which is expected to improve customer satisfaction in physical stores.
[0903] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0904] Step 1:
[0905] The user starts the application and registers by entering the required information such as name, email address, employee ID, etc. on the user registration screen. The entered information is sent from the device to the server, which stores it in a database. Once the user has completed registration, a login screen is displayed. Here, the user logs in by entering their email address and password. The server authenticates the user, and if authentication is successful, the user's dashboard is displayed.
[0906] Input: Registration information (name, email address, employee ID), login information (email address, password)
[0907] Output: User's dashboard
[0908] Step 2:
[0909] By clicking the "Start Dialogue" button on the dashboard screen, the user begins a dialogue with the conversational agent (all-affirmative bot). The device starts voice recognition, and the user begins speaking. The voice data is sent to the server, where it is input into the emotion recognition engine.
[0910] Input: User's voice data
[0911] Output: Input to the emotion recognition engine
[0912] Step 3:
[0913] The emotion recognition engine analyzes the user's voice tone, facial expressions, and text content to recognize the user's emotional state. The recognized emotional information is used as data for generating appropriate feedback through a natural language processing model.
[0914] Input: User's voice tone, facial expressions, and text content
[0915] Output: Recognized emotion information
[0916] Step 4:
[0917] The recognized emotional information and dialogue content are sent to a natural language processing model, which then uses this data to generate appropriate feedback for the user, which is then displayed on the smart glasses display in real time.
[0918] Input: Recognized emotion information, dialogue content
[0919] Output: Generated feedback
[0920] Step 5:
[0921] When the conversation ends, the device notifies the server of the end of the call. The server then analyzes the conversation using a natural language processing model based on the stored conversation log and emotion recognition results. The generated summary is then further processed by a filtering means to extract important information.
[0922] Input: Dialogue log, emotion recognition results
[0923] Output: A summary with the necessary information extracted
[0924] Step 6:
[0925] The filtered summary is sent to the administrator via a notification mechanism. The administrator can then log in to a dedicated dashboard and check the user's feedback on mental health care. The notification mechanism displays important summary parts according to the administrator's actions and also provides hints on points that require attention.
[0926] Input: Filtered summary
[0927] Output: Administrator notification, hints on what to note
[0928] Step 7:
[0929] Administrators can provide individual feedback based on the user's feedback or send messages to offer assistance, and this information is reflected in the system and used for future interactions.
[0930] Input: Admin feedback, support message
[0931] Output: Information reflected in subsequent interactions
[0932] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0933] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0934] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0935] [Third embodiment]
[0936] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0937] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0938] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0939] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0940] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0941] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0942] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0943] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0944] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0945] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0946] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0947] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0948] This invention is a dialogue agent system that uses an all-affirmation bot to provide mental care and communication support for employees. This system is implemented in the following steps.
[0949] System Configuration
[0950] This system includes users (employees), devices (such as PCs and smartphones), a server, and a conversational agent (an all-affirmation bot). These components work together to understand the user's mental state and provide appropriate feedback to their superiors.
[0951] User Registration and Login
[0952] The user first launches the application and registers by entering the required information such as name, email address, and employee ID on the user registration screen. The device sends the entered information to the server, which stores it in a database. After registration is complete, the user logs in by entering their email address and password on the login screen. The server authenticates the user, and if authentication is successful, displays the user's dashboard.
[0953] Conversation with the all-affirmation bot
[0954] The user clicks the "Start Dialogue" button on the dashboard screen to begin a dialogue with the conversational agent (all-affirmative bot). The device starts voice recognition, and the user begins speaking. The all-affirmative bot always returns a positive response to the user's statements. For example, if the user says, "I'm very tired today," the all-affirmative bot will reply, "That must have been tough. Good job!" The device sends this dialogue content to the server in real time.
[0955] Call analysis
[0956] When the conversation ends, the device notifies the server that the call has ended. The server then sends the saved conversation log to a natural language processing model and begins analyzing the text content. The generative AI model extracts important keywords and emotional trends from the conversation content and creates a summary. The server then applies a social filter to the generated summary to organize the content that should be recognized. For example, a user's statement that "I have too many meetings this week and my work isn't progressing" is summarized as "Work delays due to being busy are causing me stress."
[0957] Information Sharing and Notification
[0958] The server prepares the filtered summary to be sent to the supervisor. The supervisor logs in to a dedicated dashboard from their own device and checks the feedback on the employee's mental state. The server responds to the supervisor's actions by displaying a concise summary of the important parts and providing hints on what points to pay attention to. If necessary, the supervisor can provide individual feedback or send a message to offer support.
[0959] Specific examples
[0960] Example 1: Everyday conversation
[0961] The user schedules a conversation with the All-Affirmation Bot every day at 2:00 PM. Today, the user says, "I have too many meetings this week, so I can't get my work done." The All-Affirmation Bot responds, "That must have been tough, you've worked hard!" The server records the statement, "I have too many meetings this week, so I can't get my work done," and the generative AI model summarizes it as, "The delays in work caused by too many meetings are a stressful factor." The server applies a social filter to the summary and shares it with the user's superiors, concluding, "The user is stressed by delays in work caused by being busy."
[0962] Example 2: Important Feedback
[0963] In conversations with the All-Affirmation Bot, the user repeatedly states, "The project deadline is too tight." The All-Affirmation Bot responds, "I understand the pressure. You're doing a great job!" The server summarizes this as "The tight project deadline is the main cause of stress," and immediately notifies the supervisor. The supervisor can then take action based on the employee's feedback, such as "reevaluating the project deadline."
[0964] This allows the user to more easily reduce stress, and allows the supervisor to obtain information for effective mental care.
[0965] The processing flow will be explained below.
[0966] Step 1:
[0967] The user launches the application, enters the necessary information such as "name," "email address," and "employee ID" on the user registration screen, and presses the "Register" button.
[0968] Step 2:
[0969] The terminal temporarily stores the input user information and transmits the data to the server.
[0970] Step 3:
[0971] The server stores the received user information in a database and returns a response indicating successful registration to the terminal.
[0972] Step 4:
[0973] The terminal displays a "Registration Complete" message to the user.
[0974] Step 5:
[0975] The user enters their "email address" and "password" on the login screen and presses the "Login" button.
[0976] Step 6:
[0977] The terminal transmits the entered login information to the server.
[0978] Step 7:
[0979] The server authenticates the user using a database, and if authentication is successful, returns the user's dashboard information to the terminal.
[0980] Step 8:
[0981] The terminal displays a dashboard screen to the user.
[0982] Step 9:
[0983] The user clicks the "Start conversation" button on the dashboard screen.
[0984] Step 10:
[0985] The terminal starts voice recognition and the user begins speaking.
[0986] Step 11:
[0987] The conversational agent (all-affirmative bot) responds affirmatively to user comments. For example, if the user says, "I'm very tired today," the agent responds, "That must have been hard! You did a great job!"
[0988] Step 12:
[0989] The terminal transmits the content of the conversation between the all-affirming bot and the user to the server in real time.
[0990] Step 13:
[0991] The server stores the content of the conversation as a log and performs processing such as sentiment analysis and keyword extraction in real time as needed.
[0992] Step 14:
[0993] When the interaction ends, the terminal notifies the server of the interaction end.
[0994] Step 15:
[0995] The server sends the saved dialogue log to a natural language processing model to begin analyzing the text content.
[0996] Step 16:
[0997] The generative AI model extracts important keywords and sentiment trends from the conversation and creates a summary, such as, "The user has been busy this week, and is feeling particularly stressed by the large number of meetings."
[0998] Step 17:
[0999] The server applies a social filter to the generated summary to organize the content that should be recognized, for example, summarizing it as "Work delays due to busyness are a cause of stress."
[1000] Step 18:
[1001] The server prepares to send the filtered summary to the superior.
[1002] Step 19:
[1003] Supervisors can log into a dedicated dashboard from their own devices and check feedback on employees' mental health.
[1004] Step 20:
[1005] The server displays a concise summary of important points in response to the supervisor's operation, and also provides hints on what points to pay attention to.
[1006] Step 21:
[1007] If needed, a manager can send a message to provide individual feedback or assistance.
[1008] The above are the specific processing steps for a mental care and communication assistance application using a fully affirmative bot.
[1009] Example 1
[1010] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1011] In modern corporate environments, employee mental health is an important issue, but many systems lack the means to properly understand users' emotional states and stress levels, or to provide appropriate feedback to superiors. Therefore, a method for effectively managing employee mental health is needed. Furthermore, conventional conversational agents may respond negatively to user comments, which may increase employee stress.
[1012] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1013] In this invention, the server includes: a user interface that provides a method for a user to interact with a dialogue agent; a user interface that allows the user to initiate a dialogue with the dialogue agent and record the dialogue content; a generative AI model that allows the dialogue agent to always generate positive responses; a natural language processing unit that analyzes the dialogue content between the dialogue agent and the user to determine the user's emotional state and level of stress; a summarizing unit that summarizes the dialogue content analyzed by the natural language processing unit and extracts important keywords and emotional trends; a notification unit that notifies the supervisor of the information processed by the filtering unit; and a notification unit that provides the supervisor with feedback regarding the user's mental care and encourages the supervisor to take appropriate measures. This makes it possible to accurately grasp the employee's emotional state and level of stress and provide appropriate feedback to the supervisor.
[1014] The "user interface" refers to the operation screen and input means by which the user operates the system and starts a dialogue with the dialogue agent.
[1015] A "generative AI model" is an algorithm or program that generates positive responses in natural language based on what a user says.
[1016] "Natural language processing" is a technology that analyzes text data and extracts important keywords and emotional trends from its content.
[1017] A "dialogue agent" is a program or system that interacts with a user and generates responses to their utterances.
[1018] "Notification means" refers to the communication means or process for sending analyzed information and feedback to superiors.
[1019] The "filtering means" is a processing means for organizing the dialogue content analyzed by natural language processing and generating an appropriate summary.
[1020] "Emotional state" refers to the user's psychological state, and indicates stress, satisfaction, fatigue, etc.
[1021] The "dashboard" is an operation screen accessed by users who log in to the system, and has a central function for starting a dialogue and checking feedback.
[1022] This invention is a dialogue agent system for the purpose of providing mental care and communication support for employees, and includes a user (employee), a terminal (such as a PC or smartphone), a server, and a dialogue agent (an all-affirmation bot). These components work together to grasp the user's mental state and provide appropriate feedback to superiors.
[1023] System Configuration
[1024] The system consists of the following main components:
[1025] User Interface
[1026] The user uses this interface to start a dialogue with the dialogue agent, inputting and operating the agent according to the situation. The user interface is provided as a software application that runs on a device such as a PC or smartphone.
[1027] Generative AI Models
[1028] The conversational agent uses a generative AI model (e.g., ChatGPT) that generates positive responses based on user utterances. This model uses natural language processing techniques to understand what the user is saying and always generates a positive response.
[1029] Natural Language Processing
[1030] On the server side, natural language processing technology (e.g., BERT) is used to analyze the content of the dialogue between the conversation agent and the user, extracting important keywords and emotional trends from the user's speech and determining the user's emotional state and level of stress.
[1031] Notification means
[1032] The information processed by the filtering means is notified to the superior by the server. The notification means ensures that the data is correctly filtered and that the superior is provided with important information along with a concise summary.
[1033] User Registration and Login
[1034] The user starts the application and enters the required information such as name, email address, employee ID, etc. on the user registration screen. The device sends this information to the server, which stores it in a database. When the user enters their email address and password on the login screen, the server authenticates the user and, if successful, displays the dashboard.
[1035] Conversation with the all-affirmation bot
[1036] When the user clicks the "Start Dialogue" button on the dashboard, a dialogue with the conversational agent (all-positive bot) begins. The device starts voice recognition and captures the user's voice input. The server converts the received voice data into text and sends it to the all-positive bot. The all-positive bot uses a generative AI model to generate a positive response, and the device presents this response to the user.
[1037] Call analysis and summarization
[1038] When the dialogue ends, the device notifies the server. The server then sends the saved dialogue log to a natural language processing model, which begins analyzing the text content. The generative AI model extracts important keywords and emotional trends from the dialogue and generates a summary. For example, if a user says, "I have too many meetings this week, so I'm not making much progress on my work," the summary might read, "Work delays due to being too busy are causing me stress."
[1039] Information Sharing and Notification
[1040] The server applies a social filter to the generated summary, organizes important information, and notifies the supervisor. The supervisor logs in to a dedicated dashboard and checks the feedback. This allows the supervisor to take appropriate action based on the employee's mental state.
[1041] Specific prompt examples
[1042] Below are some examples of prompt sentences.
[1043] "Users interact with the all-affirmation bot and send what they say to your server. Your natural language processing model analyzes the text and creates a summary."
[1044] This allows the system to effectively support employees' mental health care and provide supervisors with the information they need to provide appropriate feedback.
[1045] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1046] Step 1: User Registration
[1047] A user starts an application and enters the required information such as name, email address, and employee ID on the user registration screen. The entered information (name, email address, employee ID) is input data. The device sends this input data to the server, which receives it and stores it in a database. Specifically, the server receives an HTTP POST request and registers it as a new record in the database (e.g., MySQL). The output is that the new user has been registered and the information is stored in the database.
[1048] Step 2: User Login
[1049] A user attempts to log in by entering an email address and password. The device sends this authentication information (email address, password) to the server. The server searches for the corresponding user information in the database and compares it with the input data. If authentication is successful, the server returns an authentication success message and the user's dashboard information to the device as output. The device receives this information and displays the user's dashboard.
[1050] Step 3: Initiating a conversation
[1051] The user clicks the "Start conversation" button on the dashboard. This action is registered as input data. The device starts voice input and records what the user says. The recorded voice data is input and the device sends it to the server. The server receives the voice data and converts it into text data using voice recognition software (e.g., Google Cloud Speech-to-Text API). The converted text data is output.
[1052] Step 4: Generate a response
[1053] The server inputs the received text data into a generative AI model (e.g., ChatGPT) to generate a positive response. The generative AI model performs natural language processing based on the input data (user utterances) to generate a positive response. This response text becomes the output. The server sends this response text to the device, which then presents the response to the user in voice or text.
[1054] Step 5: Record the conversation
[1055] The device sends the content of the conversation (user's comments and all affirmative bot's responses) to the server in real time. The server receives this and stores it in a database. The saved conversation log is the output. This log is used for later analysis.
[1056] Step 6: End of conversation
[1057] When the conversation ends, the device notifies the server. This notification becomes the input data. The server acquires the conversation log and sends it to a natural language processing model (e.g., BERT). This starts the analysis of the text content.
[1058] Step 7: Text analysis and summary generation
[1059] The server uses a natural language processing model to extract important keywords and emotional trends from the dialogue. This analysis is a data calculation based on the input data (dialogue log). Summary text is generated as the analysis result. This summary text is the output data. For example, a statement such as "There are too many meetings this week, so work isn't progressing" is summarized as "Work delays due to being busy are causing stress."
[1060] Step 8: Filtering and Notifications
[1061] The server applies a social filter to the generated summary text and organizes important information. The filtered text is output. This text is sent to the superior via a notification method. The superior receives the notification and checks the feedback on a dedicated dashboard. For example, the superior may receive specific feedback such as, "Work delays due to busyness are the cause of stress."
[1062] (Application example 1)
[1063] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1064] In conventional factory work, operators and engineers often suffer from excessive stress and fatigue, which has a negative impact on productivity and quality. In particular, a lack of psychological care can lead to problems such as a decrease in motivation and an increase in turnover. Therefore, a new system is needed to efficiently provide mental care for operators and engineers working in factories and improve the quality of their work.
[1065] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1066] In this invention, the server includes: means for having a user interface that provides a method for a user to interact with a dialogue agent; natural language processing means for analyzing the dialogue content between the dialogue agent and the user to determine the user's emotional state and level of stress; filtering means for filtering and summarizing the dialogue content analyzed by the natural language processing means; notification means for notifying a supervisor of the information processed by the filtering means; means for using a generative AI model that recognizes the user's everyday conversation and generates a positive response; speech recognition means for converting speech input from the user into text in real time; means for executing a program that controls the notification means to summarize and notify the dialogue content; and feedback means for the supervisor to take an appropriate action based on the notified information. This makes it possible to reduce the psychological stress of operators and engineers working in a factory and provide effective mental care.
[1067] A "user interface" is the means of interaction that allows a user to interact with a system or device.
[1068] A "dialogue agent" is a software agent that obtains information through dialogue with a user and returns an appropriate response.
[1069] "Natural language processing means" refers to algorithms or systems that analyze natural language text spoken by a user and determine their emotional state and level of stress.
[1070] The "filtering means" is a process for organizing the dialogue content analyzed by the natural language processing means, and extracting and summarizing important information.
[1071] "Notification means" refers to a function or system for notifying the processed filtering results to superiors or managers.
[1072] A "generative AI model" is an artificial intelligence model that generates appropriate responses to user input.
[1073] "Speech recognition means" refers to technology or devices that convert a user's speech into text in real time.
[1074] The "means for executing a program" refers to the software and hardware configuration for collecting, analyzing, and filtering the contents of the conversation and controlling the notification means.
[1075] "Feedback means" refers to a specific method or system that allows a superior to take appropriate action based on the notified information.
[1076] This invention is a dialogue agent system for providing mental care to operators and engineers working in factories. This system includes users (operators and engineers), terminals (smartphones, tablets, etc.), a server, and a dialogue agent (an all-affirmative bot).
[1077] System Configuration
[1078] The system works by coordinating the following components:
[1079] 1. User Interface: Provides an interface for the user to interact with the conversational agent, which can be operated by a touch screen or voice input.
[1080] 2. Conversational Agent: Engages in natural dialogue with the user. Conversational agents use generative AI models to always generate positive responses to user utterances.
[1081] 3. Speech recognition: Converts voice input into text in real time. This function uses the Google Speech-to-Text API, for example.
[1082] 4. Natural language processing: Analyzes the dialogue and determines the user's emotional state and stress level. This process uses natural language processing technologies such as GPT-3.5.
[1083] 5. Filtering means: Filters the dialogue content analyzed by the natural language processing means and summarizes important information.
[1084] 6. Notification: The information summarized by the filtering method is notified to superiors and managers via email or dashboard.
[1085] 7. Feedback measures: Provide feedback to supervisors and managers to take appropriate action based on the information provided.
[1086] Program processing
[1087] The server receives the dialogue between the conversational agent and the user in real time and generates a positive response using a generative AI model. The use of GPT-3.5 as the generative AI model allows for natural dialogue. All dialogue is saved in text format and later analyzed using natural language processing.
[1088] The device receives voice input from the user and converts it into text using a speech recognition method (e.g., Google Speech-to-Text API), which then transmits the user's speech as text to the server in real time.
[1089] The filtering means processes the acquired text data and extracts important keywords and phrases, and based on this information, summarizes the conversation and notifies superiors or managers.
[1090] The notification method sends filtered summary information to superiors and managers via email or a dedicated dashboard, allowing them to quickly take the necessary measures to care for the user's mental health.
[1091] Specific examples
[1092] For example, if a factory operator uses the system and says, "I'm so busy today, I'm tired," the all-affirmation bot will respond, "That must have been tough, good job!" After the conversation ends, the system summarizes the content of the conversation and provides feedback to the supervisor or manager that "the operator's fatigue due to being too busy is the cause."
[1093] Prompt Sentence Examples
[1094] Input prompt: "I'm too busy and tired today."
[1095] Model response: "That was hard work, good job!"
[1096] By implementing this system, it is possible to improve the working environment within the factory, maintain the motivation of operators and engineers, and reduce stress.
[1097] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1098] Step 1:
[1099] The user performs voice input. Specifically, the user speaks into the device's microphone to start a conversation. The input is what the user says, and is captured by the device as voice data.
[1100] Step 2:
[1101] A speech recognition means is activated on the terminal. The speech recognition means converts the user's voice data into text data. This process uses speech recognition technology such as the Google Speech-to-Text API. The output is data that converts the user's speech into text.
[1102] Step 3:
[1103] The device sends text data to the server, which then passes it to the generative AI model. The input is the converted text data, which is then sent to the server.
[1104] Step 4:
[1105] The conversational agent in the server uses a generative AI model to generate a positive response to the user's utterance. Specifically, a GPT-3.5 model is used to generate a prompt sentence based on the input text. The output is a positive response text.
[1106] Step 5:
[1107] The generated positive response text is returned from the server to the terminal. The terminal reproduces this text data as voice through the voice output means. The input is the generated response text, and the output is a voice response to the user.
[1108] Step 6:
[1109] The server stores the dialogue content as a log, and natural language processing means analyzes the text data. In this analysis process, the user's emotional state and stress level are determined. The input is the user's entire dialogue text, and the output is the evaluation results of the emotional state and stress.
[1110] Step 7:
[1111] The filtering means extracts important keywords and phrases from the dialogue text and generates a summary. In this step, points of particular interest are organized from the dialogue content. The input is the analysis result by the natural language processing means, and the output is the summary text.
[1112] Step 8:
[1113] The summarized information is sent to a supervisor or manager through a notification mechanism, which can be email or a dashboard. The input is the summary text and the output is the notification message provided to the supervisor or manager.
[1114] Step 9:
[1115] The supervisor or manager takes appropriate action based on the notified information. This involves implementing specific measures and countermeasures using feedback methods. The input is the notification message, and the output is specific feedback and countermeasures.
[1116] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1117] The present invention is a dialogue agent system that utilizes an all-affirmation bot, with the aim of providing mental care and communication assistance to users. By incorporating an emotion engine into this system, it is possible to more accurately recognize the user's emotional state and provide appropriate feedback. Specific embodiments of this system are described below.
[1118] System Configuration
[1119] This system includes users (employees), devices (PCs, smartphones, etc.), a server, a dialogue agent (a completely affirmative bot), and an emotion engine. These components work together to understand the user's mental state and provide appropriate feedback to superiors.
[1120] User Registration and Login
[1121] The user first launches the application and registers by entering the required information such as name, email address, and employee ID on the user registration screen. The device sends the entered information to the server, which stores it in a database. After registration is complete, the user logs in by entering their email address and password on the login screen. The server authenticates the user, and if authentication is successful, displays the user's dashboard.
[1122] Conversation with an all-affirmation bot and emotion recognition
[1123] The user clicks the "Start conversation" button on the dashboard screen to begin a conversation with the conversational agent (all-affirmative bot). The device starts voice recognition, and the user begins speaking. The all-affirmative bot always returns a positive response to the user's statements. For example, if the user says, "I'm very tired today," the bot replies, "That must have been tough! You've done well!"
[1124] The emotion engine analyzes the user's voice tone, facial expressions, and text content to recognize their emotions. The recognized emotional information is reflected in the conversational agent's responses. For example, if the user is determined to be very tired, the all-affirmation bot will adjust its responses to be more warmhearted.
[1125] Analysis of call content and emotional information
[1126] When the conversation ends, the device notifies the server that the call has ended. The server then sends the saved conversation log and emotion recognition results to a natural language processing model, which begins analyzing the text content and emotional information. The generative AI model extracts important keywords and emotional trends from the conversation content and creates a summary. The server then applies a social filter to the generated summary to organize the content that should be recognized. For example, information such as "There are too many meetings this week, so work is not progressing" and "The user is feeling tired" is summarized as "Work delays due to busyness are the cause of stress."
[1127] Information Sharing and Notification
[1128] The server prepares the filtered summary to be sent to the supervisor. The supervisor logs in to a dedicated dashboard from their own device and checks the feedback on the employee's mental state. The server responds to the supervisor's actions by displaying a concise summary of the important parts and providing hints on what points to pay attention to. If necessary, the supervisor can provide individual feedback or send a message to offer support.
[1129] Specific examples
[1130] Example 1: Everyday dialogue and emotion recognition
[1131] The user schedules a conversation with the All-Affirmation Bot every day at 2:00 PM. Today, the user says, "There are too many meetings this week, and I'm not making progress on my work." The All-Affirmation Bot replies, "That must have been tough, you've worked hard!" At this point, the emotion engine recognizes a strong sense of fatigue from the user's tone of voice. The server records the statement, "There are too many meetings this week, and I'm not making progress on my work," along with the "strong sense of fatigue," and the generative AI model summarizes it as, "The delays in work due to the many meetings are a stressful factor and cause fatigue." The server applies a sociality filter to the above summary and shares it with the supervisor, concluding, "The user is experiencing stress due to work delays caused by being busy, and is feeling very fatigued."
[1132] Example 2: Critical feedback and emotion recognition
[1133] In conversation with the All-Affirmation Bot, the user repeatedly states, "The project deadline is too tight." The All-Affirmation Bot responds, "I understand the pressure. You're doing a great job!" At this point, the emotion engine recognizes the user's strong stress from their facial expressions and voice. The server summarizes this as "The tight project deadline is the main cause of stress," and immediately notifies this information to the supervisor. The supervisor can then take action based on the employee's feedback, such as "reevaluating the project deadline."
[1134] This makes it easier for users to reduce stress and provides supervisors with information to provide effective mental care.The introduction of the emotion engine allows for a more accurate understanding of the user's emotional state, improving the quality of responses and feedback.
[1135] The processing flow will be explained below.
[1136] Step 1:
[1137] The user launches the application, enters the necessary information such as "name," "email address," and "employee ID" on the user registration screen, and presses the "Register" button.
[1138] Step 2:
[1139] The terminal temporarily stores the input user information and transmits the data to the server.
[1140] Step 3:
[1141] The server stores the received user information in a database and returns a response indicating successful registration to the terminal.
[1142] Step 4:
[1143] The terminal displays a "Registration Complete" message to the user.
[1144] Step 5:
[1145] The user enters their "email address" and "password" on the login screen and presses the "Login" button.
[1146] Step 6:
[1147] The terminal transmits the entered login information to the server.
[1148] Step 7:
[1149] The server authenticates the user using a database, and if authentication is successful, returns the user's dashboard information to the terminal.
[1150] Step 8:
[1151] The terminal displays a dashboard screen to the user.
[1152] Step 9:
[1153] The user clicks the "Start conversation" button on the dashboard screen.
[1154] Step 10:
[1155] The terminal starts voice recognition and the user begins speaking.
[1156] Step 11:
[1157] The conversational agent (all-affirmative bot) responds affirmatively to user comments. For example, if the user says, "I'm very tired today," the agent responds, "That must have been hard! You did a great job!"
[1158] Step 12:
[1159] The emotion engine recognizes the user's emotions by analyzing their voice tone, facial expressions, and text content, for example, reading tiredness from their voice and sadness from their facial expressions.
[1160] Step 13:
[1161] The terminal transmits the content of the conversation between the all-affirmation bot and the user, as well as the emotional data obtained from the emotion engine, to the server in real time.
[1162] Step 14:
[1163] The server stores the dialogue content and emotional data as logs, and performs real-time processing such as emotion analysis and keyword extraction as needed.
[1164] Step 15:
[1165] When the interaction ends, the terminal notifies the server of the interaction end.
[1166] Step 16:
[1167] The server sends the saved dialogue logs and emotion data to a natural language processing model, which begins analyzing the text content and emotion information.
[1168] Step 17:
[1169] The generative AI model extracts important keywords and emotional trends from the conversation and creates a summary, such as, "The user has been busy this week, and is feeling stressed and tired, especially with so many meetings."
[1170] Step 18:
[1171] The server applies a social filter to the generated summary to organize the content that should be recognized, for example, summarizing it as "Work delays due to busyness are a cause of stress."
[1172] Step 19:
[1173] The server prepares to send the filtered summary to the superior.
[1174] Step 20:
[1175] Supervisors can log into a dedicated dashboard from their own devices and check feedback on employees' mental health.
[1176] Step 21:
[1177] The server displays a concise summary of important points in response to the supervisor's operation, and also provides hints on what points to pay attention to.
[1178] Step 22:
[1179] If needed, a manager can send a message to provide individual feedback or assistance.
[1180] Step 23:
[1181] Users receive feedback from their superiors and implement advice on mental care and work improvement.
[1182] The above are the specific processing steps for a mental care and communication assistance application using a fully affirmative bot combined with an emotion engine.
[1183] Example 2
[1184] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1185] To effectively provide mental care to users, it is necessary to accurately grasp the user's emotional state and stress level and provide appropriate feedback accordingly. However, conventional systems lack the means to accurately recognize the user's emotions, making it difficult to provide appropriate responses and feedback. In addition, sharing information with superiors is cumbersome, making it difficult to provide effective mental care to users.
[1186] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1187] In this invention, the server includes means for having a user interface through which the user interacts with the dialogue agent, natural language processing means for analyzing the dialogue between the dialogue agent and the user to determine the user's emotional state and level of stress, filtering means for filtering and summarizing the dialogue analyzed by the natural language processing means, emotion recognition means for analyzing the user's voice tone, facial expressions, and text content to recognize the user's emotions, adjustment means for reflecting the emotion information obtained by the emotion recognition means in the dialogue agent's response content, and notification means for notifying a superior of the information processed by the filtering means. This makes it possible to accurately grasp the user's emotional state and provide appropriate feedback.
[1188] Below are definitions of important terms included in the rewritten claims:
[1189] "User interface" refers to the screen or means by which a user accesses and operates a system.
[1190] An "interactive agent" is a program or system that interacts with a user and responds based on the interaction.
[1191] "Natural language processing means" refers to a technique or means for analyzing the content of a dialogue between a user and a dialogue agent and determining the user's emotional state and stress level.
[1192] "Filtering means" refers to a technology or means for sorting the dialogue content analyzed by natural language processing means and extracting and summarizing important information.
[1193] The "emotion recognition means" refers to a technology or means for analyzing the user's voice tone, facial expression, and text content to recognize the user's emotions.
[1194] The "adjustment means" refers to a technique or means for appropriately changing or adjusting the response content of the dialogue agent based on the emotional information obtained by the emotion recognition means.
[1195] "Notification means" refers to a technique or means for transmitting the information processed by the filtering means to a superior in an appropriate format.
[1196] This invention is a dialogue agent system that uses an all-affirmation bot and an emotion engine, with the aim of providing mental care and communication assistance to users. Specific embodiments of this system are described below.
[1197] System Configuration
[1198] This system includes a user, a terminal, a server, a dialogue agent (a fully affirmative bot), and an emotion engine. These elements work together to understand the user's mental state and provide appropriate feedback.
[1199] User Registration and Login
[1200] 1. The user launches the application and registers by entering information such as their name, email address, and employee ID.
[1201] 2. The terminal sends the entered information to the server, which stores it in a database.
[1202] 3. After completing registration, the user logs in by entering their email address and password.
[1203] 4. The server authenticates the user and, if authentication is successful, displays the user's dashboard.
[1204] Conversation with the all-affirmation bot
[1205] 1. The user clicks the "Start conversation" button on the dashboard screen to begin a conversation with the all-affirmation bot.
[1206] 2. The device starts voice recognition and sends the user's voice to the all-affirmation bot. For example, if the user says, "I'm tired today," the all-affirmation bot replies, "That must have been hard! Good job!"
[1207] emotion recognition
[1208] 1. The device simultaneously transmits the user's voice tone, facial expression, and text content to the emotion engine.
[1209] 2. The emotion engine analyzes this data and recognizes the user's emotions.
[1210] 3. The recognized emotional information is reflected in the responses of the all-positive bot. For example, if the user is judged to be very tired, the all-positive bot will adjust its responses to be more warmhearted.
[1211] Call analysis
[1212] 1. When the conversation ends, the terminal notifies the server that the call has ended.
[1213] 2. The server sends the dialogue log and emotion recognition results to the natural language processing model.
[1214] 3. The generative AI model extracts important keywords and emotional trends from the conversation and creates a summary. For example, the information that "the user is feeling tired" and the statement that "there are too many meetings this week and work isn't progressing" can be summarized as "work delays due to being busy are the cause of stress."
[1215] Information Sharing and Notification
[1216] 1. The server prepares the filtered summary for notification to the supervisor.
[1217] 2. Supervisors log in to a dedicated dashboard and check feedback on employees' mental health.
[1218] 3. The server responds to the supervisor's actions by displaying a concise summary of the important points and providing hints on what points the supervisor should pay attention to. If necessary, the supervisor can provide feedback or send a support message.
[1219] Specific examples
[1220] Example 1: Everyday dialogue and emotion recognition
[1221] Consider a scenario where a user says, "I have too many meetings this week and I'm not getting any work done."
[1222] The all-affirmation bot responds with, "That must have been tough, good job!" At this point, the emotion engine recognizes the strong sense of fatigue from the user's tone of voice and records it on the server.
[1223] The generative AI model summarizes that "work delays caused by numerous meetings are a source of stress and fatigue."
[1224] The server applies a social filter and shares the information with superiors, stating that "users are stressed by delays in work due to being busy, and feel very tired."
[1225] Example 2: Critical feedback and emotion recognition
[1226] Consider a situation where a user repeatedly says, "The project deadline is too tight."
[1227] The all-affirmation bot replies, "I understand the pressure. You're doing a great job!"
[1228] The emotion engine recognizes strong stress from the user's facial expressions and voice.
[1229] The server summarizes that "tight project deadlines are the main cause of stress" and immediately communicates this information to his superiors.
[1230] Based on the feedback, the supervisor can take action such as "reevaluating the project deadline."
[1231] Example prompts for generative AI models
[1232] (Example prompt 1: Summarizing everyday conversations)
[1233] User says: "I have too many meetings this week and I can't get any work done."
[1234] Recognizing the Emotion Engine: Extreme Fatigue
[1235] Summary generated: Work delays due to too many meetings are a major cause of stress and fatigue
[1236] (Example prompt 2: Summary of important feedback)
[1237] User says: "The project deadline is too tight"
[1238] Recognizing the Emotion Engine: High Stress
[1239] Generate summary: Tight project deadlines are a major source of stress
[1240] This system makes it possible to accurately grasp the user's emotional state and provide appropriate feedback, thereby improving the effectiveness of mental care.
[1241] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1242] Step 1:
[1243] The user launches the application and registers by entering their name, email address, and employee ID.
[1244] Input: Name, Email Address, Employee ID
[1245] How it works: The device sends this information to the server.
[1246] Output: User information is sent to the server and a registration completion message is received.
[1247] Step 2:
[1248] The server stores the received information in a database and sends a notification of registration completion to the terminal.
[1249] Input: User information (name, email address, employee ID)
[1250] What happens: The server creates a new record in the database and saves the information.
[1251] Output: A message that registration is complete is displayed on the terminal.
[1252] Step 3:
[1253] The user enters their email address and password and clicks the login button.
[1254] Input: Email address, Password
[1255] How it works: The device sends this information to the server.
[1256] Output: Login information sent to the server.
[1257] Step 4:
[1258] The server compares the entered information with a database and performs authentication.
[1259] Input: Email address, Password
[1260] How it works: The server checks the information in its database and performs authentication.
[1261] Output: If authentication is successful, information is sent to the device to display the user's dashboard.
[1262] Step 5:
[1263] The user clicks the "Start conversation" button on the dashboard screen.
[1264] Input: Click the "Start conversation" button
[1265] Action: The device will begin voice input and perform the initial setup to connect to the all-affirmation bot.
[1266] Output: Ready to connect with all affirmative bots.
[1267] Step 6:
[1268] When the user starts speaking, the device converts the speech into text and sends it to the All Affirmations Bot.
[1269] Input: User's voice data
[1270] Operation: The device performs voice recognition and sends the recognized text data to all affirmative bots.
[1271] Output: Text data sent to all affirmation bots.
[1272] Step 7:
[1273] An all-affirmation bot always responds affirmatively to what the user says.
[1274] Input: Text data based on speech recognition
[1275] How it works: The all-affirmation bot generates appropriate affirmative responses. For example, if someone says "I'm tired today," it will respond with "That must have been hard, good job!"
[1276] Output: Affirmative response to the user.
[1277] Step 8:
[1278] The device sends the user's voice and text data to the emotion engine.
[1279] Input: User voice and text data
[1280] How it works: The emotion engine analyzes this data and recognizes emotions.
[1281] Output: Recognized emotion information.
[1282] Step 9:
[1283] The emotion engine feeds back the analysis results to the dialogue agent and adjusts the response content.
[1284] Input: Emotion recognition results
[1285] How it works: The bot updates its responses based on the results of the emotion engine. For example, if it determines that you are tired, it will respond with something like, "Maybe it would be good to take a break."
[1286] Output: The adjusted response.
[1287] Step 10:
[1288] When the conversation ends, the terminal notifies the server that the call has ended.
[1289] Input: Call end event
[1290] Action: The device sends a call end notification to the server.
[1291] Output: A conversation termination notification is sent to the server.
[1292] Step 11:
[1293] The server sends the call logs and emotion recognition results to the natural language processing model.
[1294] Input: Call logs, emotion recognition results
[1295] How it works: The server sends this data to the natural language processing model and begins analysis.
[1296] Output: The data required for analysis is sent.
[1297] Step 12:
[1298] The generative AI model extracts important keywords and emotional trends from the dialogue and creates a summary.
[1299] Input: Call logs, emotion recognition results
[1300] How it works: The generative AI model analyzes data, extracts key keywords and sentiment trends, and generates summaries. For example, if a user says, "I'm having too many meetings this week, so I'm not making progress on my work," the model summarizes that "Work delays due to being too busy are causing stress."
[1301] Output: Summarized information.
[1302] Step 13:
[1303] The server applies a sociality filter to the generated summary to organize the content to be recognized.
[1304] Input: Summary information
[1305] How it works: The server applies social filters to determine what should be recognized.
[1306] Output: Filtered summary information.
[1307] Step 14:
[1308] The server prepares to send the filtered summary to the supervisor.
[1309] Input: Filtered summary information
[1310] What it does: The server formats the information for notification.
[1311] Output: Filtered summary for notifying superiors.
[1312] Step 15:
[1313] Managers can log in to a dedicated dashboard and check feedback on employees' mental health.
[1314] Input: Filtered summary information
[1315] What happens: Your manager accesses the dashboard and sees the feedback you provided.
[1316] Output: Your supervisor reviews the feedback and is ready to take action if necessary.
[1317] (Application example 2)
[1318] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1319] Conventional conversational agent systems have difficulty accurately grasping a user's emotional state and stress level, making it difficult to provide appropriate feedback. Furthermore, there is a demand for systems that enable sales staff, especially in brick-and-mortar stores, to grasp and respond to customer emotions in real time.
[1320] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means having a user interface for the user to interact with the dialogue agent, natural language processing means for analyzing the content of the dialogue between the dialogue agent and the user to determine the user's emotional state and stress level, filtering means for filtering and summarizing the content of the dialogue analyzed by the natural language processing means, notification means for notifying an administrator of the information processed by the filtering means, and device linking means for recognizing the user's emotional state in real time using smart glasses and displaying feedback. This enables sales staff to grasp customer emotions in real time and respond optimally.
[1321] A "user interface" is a screen or operating means that allows a user to interact with or operate a system.
[1322] A "dialogue agent" is software or a system that interacts with a user and supports communication using natural language processing.
[1323] "Natural language processing means" is a technology for analyzing the content of conversation between a dialogue agent and a user and understanding the user's emotions and intentions.
[1324] The "filtering means" is a technology for organizing information analyzed by the natural language processing means based on specific criteria and extracting necessary information.
[1325] The "notification means" is a function or device for transmitting the information sorted by the filtering means to the administrator.
[1326] "Smart glasses" are eyeglass-type wearable devices that have a display function and provide information to the user.
[1327] "Device integration means" is a technology that links devices such as smart glasses with systems to display and update information in real time.
[1328] "Administrator" is the person or department responsible for overseeing the status of the system and users and providing necessary support and feedback.
[1329] "Real time" refers to a state in which processing or communication is carried out almost immediately after the data is generated.
[1330] This invention is a system for users to receive mental health care through dialogue with a conversational agent. The system uses smart glasses to recognize the user's emotional state in real time and provide appropriate feedback.
[1331] Hardware and software used
[1332] The hardware used includes:
[1333] Smart glasses (e.g. Google Glass, Vuzix Blade)
[1334] Server (e.g. AWS EC2)
[1335] The software used includes:
[1336] Emotion recognition engine (e.g. Microsoft Azure Emotion API)
[1337] Natural language processing models (e.g., OpenAI GPT-4)
[1338] Front-end applications (e.g. React Native)
[1339] Database (e.g. MySQL)
[1340] Overall processing of the program
[1341] The system analyzes the user's emotional state and dialogue content in real time and provides optimal feedback.
[1342] 1. User Registration and Login
[1343] First, the user starts the application and registers by entering basic information on the user registration screen. The terminal sends the entered information to the server, which stores it in a database. After registration, the user can log in and access the system.
[1344] 2. Interaction with a conversational agent
[1345] The user wears the smart glasses and starts a dialogue with the conversational agent. The smart glasses capture the user's voice and facial expressions and send them to an emotion recognition engine. The emotion recognition engine analyzes the user's emotional state and sends it to the server in real time.
[1346] 3. Emotion Recognition and Feedback
[1347] The server then sends the data received from the emotion recognition engine to a natural language processing model to generate appropriate feedback, which is then displayed in real time on the smart glasses display for the user to review.
[1348] Specific examples
[1349] Example prompt sentence 1:
[1350] Generate a conversational agent response and emotion recognition results when a user says, "Tell me more about this product." The customer's expression is filled with curiosity.
[1351] Example system response:
[1352] Emotion recognition result: "Curiosity"
[1353] All-Affirmative bot replies: "I understand. The special feature of this product is that it is made using the latest technology and is extremely durable."
[1354] Example prompt sentence 2:
[1355] Generate a conversational agent response and emotion recognition results when a user says, "This product was a disappointment." The customer's tone sounds angry.
[1356] Example system response:
[1357] Emotion recognition result: "Anger"
[1358] The all-affirmation bot's response: "I'm sorry you felt that way. I'll get back to you shortly. Can you please provide more details?"
[1359] This system allows users to receive appropriate feedback in real time, allowing managers to accurately grasp the user's emotional state and respond promptly, which is expected to improve customer satisfaction in physical stores.
[1360] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1361] Step 1:
[1362] The user starts the application and registers by entering the required information such as name, email address, employee ID, etc. on the user registration screen. The entered information is sent from the device to the server, which stores it in a database. Once the user has completed registration, a login screen is displayed. Here, the user logs in by entering their email address and password. The server authenticates the user, and if authentication is successful, the user's dashboard is displayed.
[1363] Input: Registration information (name, email address, employee ID), login information (email address, password)
[1364] Output: User's dashboard
[1365] Step 2:
[1366] By clicking the "Start Dialogue" button on the dashboard screen, the user begins a dialogue with the conversational agent (all-affirmative bot). The device starts voice recognition, and the user begins speaking. The voice data is sent to the server, where it is input into the emotion recognition engine.
[1367] Input: User's voice data
[1368] Output: Input to the emotion recognition engine
[1369] Step 3:
[1370] The emotion recognition engine analyzes the user's voice tone, facial expressions, and text content to recognize the user's emotional state. The recognized emotional information is used as data for generating appropriate feedback through a natural language processing model.
[1371] Input: User's voice tone, facial expressions, and text content
[1372] Output: Recognized emotion information
[1373] Step 4:
[1374] The recognized emotional information and dialogue content are sent to a natural language processing model, which then uses this data to generate appropriate feedback for the user, which is then displayed on the smart glasses display in real time.
[1375] Input: Recognized emotion information, dialogue content
[1376] Output: Generated feedback
[1377] Step 5:
[1378] When the conversation ends, the device notifies the server of the end of the call. The server then analyzes the conversation using a natural language processing model based on the stored conversation log and emotion recognition results. The generated summary is then further processed by a filtering means to extract important information.
[1379] Input: Dialogue log, emotion recognition results
[1380] Output: A summary with the necessary information extracted
[1381] Step 6:
[1382] The filtered summary is sent to the administrator via a notification mechanism. The administrator can then log in to a dedicated dashboard and check the user's feedback on mental health care. The notification mechanism displays important summary parts according to the administrator's actions and also provides hints on points that require attention.
[1383] Input: Filtered summary
[1384] Output: Administrator notification, hints on what to note
[1385] Step 7:
[1386] Administrators can provide individual feedback based on the user's feedback or send messages to offer assistance, and this information is reflected in the system and used for future interactions.
[1387] Input: Admin feedback, support message
[1388] Output: Information reflected in subsequent interactions
[1389] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1390] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1391] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1392] [Fourth embodiment]
[1393] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1394] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1395] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1396] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1397] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1398] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1399] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1400] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1401] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1402] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1403] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1404] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1405] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1406] This invention is a dialogue agent system that uses an all-affirmation bot to provide mental care and communication support for employees. This system is implemented in the following steps.
[1407] System Configuration
[1408] This system includes users (employees), devices (such as PCs and smartphones), a server, and a conversational agent (an all-affirmation bot). These components work together to understand the user's mental state and provide appropriate feedback to their superiors.
[1409] User Registration and Login
[1410] The user first launches the application and registers by entering the required information such as name, email address, and employee ID on the user registration screen. The device sends the entered information to the server, which stores it in a database. After registration is complete, the user logs in by entering their email address and password on the login screen. The server authenticates the user, and if authentication is successful, displays the user's dashboard.
[1411] Conversation with the all-affirmation bot
[1412] The user clicks the "Start Dialogue" button on the dashboard screen to begin a dialogue with the conversational agent (all-affirmative bot). The device starts voice recognition, and the user begins speaking. The all-affirmative bot always returns a positive response to the user's statements. For example, if the user says, "I'm very tired today," the all-affirmative bot will reply, "That must have been tough. Good job!" The device sends this dialogue content to the server in real time.
[1413] Call analysis
[1414] When the conversation ends, the device notifies the server that the call has ended. The server then sends the saved conversation log to a natural language processing model and begins analyzing the text content. The generative AI model extracts important keywords and emotional trends from the conversation content and creates a summary. The server then applies a social filter to the generated summary to organize the content that should be recognized. For example, a user's statement that "I have too many meetings this week and my work isn't progressing" is summarized as "Work delays due to being busy are causing me stress."
[1415] Information Sharing and Notification
[1416] The server prepares the filtered summary to be sent to the supervisor. The supervisor logs in to a dedicated dashboard from their own device and checks the feedback on the employee's mental state. The server responds to the supervisor's actions by displaying a concise summary of the important parts and providing hints on what points to pay attention to. If necessary, the supervisor can provide individual feedback or send a message to offer support.
[1417] Specific examples
[1418] Example 1: Everyday conversation
[1419] The user schedules a conversation with the All-Affirmation Bot every day at 2:00 PM. Today, the user says, "I have too many meetings this week, so I can't get my work done." The All-Affirmation Bot responds, "That must have been tough, you've worked hard!" The server records the statement, "I have too many meetings this week, so I can't get my work done," and the generative AI model summarizes it as, "The delays in work caused by too many meetings are a stressful factor." The server applies a social filter to the summary and shares it with the user's superiors, concluding, "The user is stressed by delays in work caused by being busy."
[1420] Example 2: Important Feedback
[1421] In conversations with the All-Affirmation Bot, the user repeatedly states, "The project deadline is too tight." The All-Affirmation Bot responds, "I understand the pressure. You're doing a great job!" The server summarizes this as "The tight project deadline is the main cause of stress," and immediately notifies the supervisor. The supervisor can then take action based on the employee's feedback, such as "reevaluating the project deadline."
[1422] This allows the user to more easily reduce stress, and allows the supervisor to obtain information for effective mental care.
[1423] The processing flow will be explained below.
[1424] Step 1:
[1425] The user launches the application, enters the necessary information such as "name," "email address," and "employee ID" on the user registration screen, and presses the "Register" button.
[1426] Step 2:
[1427] The terminal temporarily stores the input user information and transmits the data to the server.
[1428] Step 3:
[1429] The server stores the received user information in a database and returns a response indicating successful registration to the terminal.
[1430] Step 4:
[1431] The terminal displays a "Registration Complete" message to the user.
[1432] Step 5:
[1433] The user enters their "email address" and "password" on the login screen and presses the "Login" button.
[1434] Step 6:
[1435] The terminal transmits the entered login information to the server.
[1436] Step 7:
[1437] The server authenticates the user using a database, and if authentication is successful, returns the user's dashboard information to the terminal.
[1438] Step 8:
[1439] The terminal displays a dashboard screen to the user.
[1440] Step 9:
[1441] The user clicks the "Start conversation" button on the dashboard screen.
[1442] Step 10:
[1443] The terminal starts voice recognition and the user begins speaking.
[1444] Step 11:
[1445] The conversational agent (all-affirmative bot) responds affirmatively to user comments. For example, if the user says, "I'm very tired today," the agent responds, "That must have been hard! You did a great job!"
[1446] Step 12:
[1447] The terminal transmits the content of the conversation between the all-affirming bot and the user to the server in real time.
[1448] Step 13:
[1449] The server stores the content of the conversation as a log and performs processing such as sentiment analysis and keyword extraction in real time as needed.
[1450] Step 14:
[1451] When the interaction ends, the terminal notifies the server of the interaction end.
[1452] Step 15:
[1453] The server sends the saved dialogue log to a natural language processing model to begin analyzing the text content.
[1454] Step 16:
[1455] The generative AI model extracts important keywords and sentiment trends from the conversation and creates a summary, such as, "The user has been busy this week, and is feeling particularly stressed by the large number of meetings."
[1456] Step 17:
[1457] The server applies a social filter to the generated summary to organize the content that should be recognized, for example, summarizing it as "Work delays due to busyness are a cause of stress."
[1458] Step 18:
[1459] The server prepares to send the filtered summary to the superior.
[1460] Step 19:
[1461] Supervisors can log into a dedicated dashboard from their own devices and check feedback on employees' mental health.
[1462] Step 20:
[1463] The server displays a concise summary of important points in response to the supervisor's operation, and also provides hints on what points to pay attention to.
[1464] Step 21:
[1465] If needed, a manager can send a message to provide individual feedback or assistance.
[1466] The above are the specific processing steps for a mental care and communication assistance application using a fully affirmative bot.
[1467] Example 1
[1468] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1469] In modern corporate environments, employee mental health is an important issue, but many systems lack the means to properly understand users' emotional states and stress levels, or to provide appropriate feedback to superiors. Therefore, a method for effectively managing employee mental health is needed. Furthermore, conventional conversational agents may respond negatively to user comments, which may increase employee stress.
[1470] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1471] In this invention, the server includes: a user interface that provides a method for a user to interact with a dialogue agent; a user interface that allows the user to initiate a dialogue with the dialogue agent and record the dialogue content; a generative AI model that allows the dialogue agent to always generate positive responses; a natural language processing unit that analyzes the dialogue content between the dialogue agent and the user to determine the user's emotional state and level of stress; a summarizing unit that summarizes the dialogue content analyzed by the natural language processing unit and extracts important keywords and emotional trends; a notification unit that notifies the supervisor of the information processed by the filtering unit; and a notification unit that provides the supervisor with feedback regarding the user's mental care and encourages the supervisor to take appropriate measures. This makes it possible to accurately grasp the employee's emotional state and level of stress and provide appropriate feedback to the supervisor.
[1472] The "user interface" refers to the operation screen and input means by which the user operates the system and starts a dialogue with the dialogue agent.
[1473] A "generative AI model" is an algorithm or program that generates positive responses in natural language based on what a user says.
[1474] "Natural language processing" is a technology that analyzes text data and extracts important keywords and emotional trends from its content.
[1475] A "dialogue agent" is a program or system that interacts with a user and generates responses to their utterances.
[1476] "Notification means" refers to the communication means or process for sending analyzed information and feedback to superiors.
[1477] The "filtering means" is a processing means for organizing the dialogue content analyzed by natural language processing and generating an appropriate summary.
[1478] "Emotional state" refers to the user's psychological state, and indicates stress, satisfaction, fatigue, etc.
[1479] The "dashboard" is an operation screen accessed by users who log in to the system, and has a central function for starting a dialogue and checking feedback.
[1480] This invention is a dialogue agent system for the purpose of providing mental care and communication support for employees, and includes a user (employee), a terminal (such as a PC or smartphone), a server, and a dialogue agent (an all-affirmation bot). These components work together to grasp the user's mental state and provide appropriate feedback to superiors.
[1481] System Configuration
[1482] The system consists of the following main components:
[1483] User Interface
[1484] The user uses this interface to start a dialogue with the dialogue agent, inputting and operating the agent according to the situation. The user interface is provided as a software application that runs on a device such as a PC or smartphone.
[1485] Generative AI Models
[1486] The conversational agent uses a generative AI model (e.g., ChatGPT) that generates positive responses based on user utterances. This model uses natural language processing techniques to understand what the user is saying and always generates a positive response.
[1487] Natural Language Processing
[1488] On the server side, natural language processing technology (e.g., BERT) is used to analyze the content of the dialogue between the conversation agent and the user, extracting important keywords and emotional trends from the user's speech and determining the user's emotional state and level of stress.
[1489] Notification means
[1490] The information processed by the filtering means is notified to the superior by the server. The notification means ensures that the data is correctly filtered and that the superior is provided with important information along with a concise summary.
[1491] User Registration and Login
[1492] The user starts the application and enters the required information such as name, email address, employee ID, etc. on the user registration screen. The device sends this information to the server, which stores it in a database. When the user enters their email address and password on the login screen, the server authenticates the user and, if successful, displays the dashboard.
[1493] Conversation with the all-affirmation bot
[1494] When the user clicks the "Start Dialogue" button on the dashboard, a dialogue with the conversational agent (all-positive bot) begins. The device starts voice recognition and captures the user's voice input. The server converts the received voice data into text and sends it to the all-positive bot. The all-positive bot uses a generative AI model to generate a positive response, and the device presents this response to the user.
[1495] Call analysis and summarization
[1496] When the dialogue ends, the device notifies the server. The server then sends the saved dialogue log to a natural language processing model, which begins analyzing the text content. The generative AI model extracts important keywords and emotional trends from the dialogue and generates a summary. For example, if a user says, "I have too many meetings this week, so I'm not making much progress on my work," the summary might read, "Work delays due to being too busy are causing me stress."
[1497] Information Sharing and Notification
[1498] The server applies a social filter to the generated summary, organizes important information, and notifies the supervisor. The supervisor logs in to a dedicated dashboard and checks the feedback. This allows the supervisor to take appropriate action based on the employee's mental state.
[1499] Specific prompt examples
[1500] Below are some examples of prompt sentences.
[1501] "Users interact with the all-affirmation bot and send what they say to your server. Your natural language processing model analyzes the text and creates a summary."
[1502] This allows the system to effectively support employees' mental health care and provide supervisors with the information they need to provide appropriate feedback.
[1503] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1504] Step 1: User Registration
[1505] A user starts an application and enters the required information such as name, email address, and employee ID on the user registration screen. The entered information (name, email address, employee ID) is input data. The device sends this input data to the server, which receives it and stores it in a database. Specifically, the server receives an HTTP POST request and registers it as a new record in the database (e.g., MySQL). The output is that the new user has been registered and the information is stored in the database.
[1506] Step 2: User Login
[1507] A user attempts to log in by entering an email address and password. The device sends this authentication information (email address, password) to the server. The server searches for the corresponding user information in the database and compares it with the input data. If authentication is successful, the server returns an authentication success message and the user's dashboard information to the device as output. The device receives this information and displays the user's dashboard.
[1508] Step 3: Initiating a conversation
[1509] The user clicks the "Start conversation" button on the dashboard. This action is registered as input data. The device starts voice input and records what the user says. The recorded voice data is input and the device sends it to the server. The server receives the voice data and converts it into text data using voice recognition software (e.g., Google Cloud Speech-to-Text API). The converted text data is output.
[1510] Step 4: Generate a response
[1511] The server inputs the received text data into a generative AI model (e.g., ChatGPT) to generate a positive response. The generative AI model performs natural language processing based on the input data (user utterances) to generate a positive response. This response text becomes the output. The server sends this response text to the device, which then presents the response to the user in voice or text.
[1512] Step 5: Record the conversation
[1513] The device sends the content of the conversation (user's comments and all affirmative bot's responses) to the server in real time. The server receives this and stores it in a database. The saved conversation log is the output. This log is used for later analysis.
[1514] Step 6: End of conversation
[1515] When the conversation ends, the device notifies the server. This notification becomes the input data. The server acquires the conversation log and sends it to a natural language processing model (e.g., BERT). This starts the analysis of the text content.
[1516] Step 7: Text analysis and summary generation
[1517] The server uses a natural language processing model to extract important keywords and emotional trends from the dialogue. This analysis is a data calculation based on the input data (dialogue log). Summary text is generated as the analysis result. This summary text is the output data. For example, a statement such as "There are too many meetings this week, so work isn't progressing" is summarized as "Work delays due to being busy are causing stress."
[1518] Step 8: Filtering and Notifications
[1519] The server applies a social filter to the generated summary text and organizes important information. The filtered text is output. This text is sent to the superior via a notification method. The superior receives the notification and checks the feedback on a dedicated dashboard. For example, the superior may receive specific feedback such as, "Work delays due to busyness are the cause of stress."
[1520] (Application example 1)
[1521] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1522] In conventional factory work, operators and engineers often suffer from excessive stress and fatigue, which has a negative impact on productivity and quality. In particular, a lack of psychological care can lead to problems such as a decrease in motivation and an increase in turnover. Therefore, a new system is needed to efficiently provide mental care for operators and engineers working in factories and improve the quality of their work.
[1523] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1524] In this invention, the server includes: means for having a user interface that provides a method for a user to interact with a dialogue agent; natural language processing means for analyzing the dialogue content between the dialogue agent and the user to determine the user's emotional state and level of stress; filtering means for filtering and summarizing the dialogue content analyzed by the natural language processing means; notification means for notifying a supervisor of the information processed by the filtering means; means for using a generative AI model that recognizes the user's everyday conversation and generates a positive response; speech recognition means for converting speech input from the user into text in real time; means for executing a program that controls the notification means to summarize and notify the dialogue content; and feedback means for the supervisor to take an appropriate action based on the notified information. This makes it possible to reduce the psychological stress of operators and engineers working in a factory and provide effective mental care.
[1525] A "user interface" is the means of interaction that allows a user to interact with a system or device.
[1526] A "dialogue agent" is a software agent that obtains information through dialogue with a user and returns an appropriate response.
[1527] "Natural language processing means" refers to algorithms or systems that analyze natural language text spoken by a user and determine their emotional state and level of stress.
[1528] The "filtering means" is a process for organizing the dialogue content analyzed by the natural language processing means, and extracting and summarizing important information.
[1529] "Notification means" refers to a function or system for notifying the processed filtering results to superiors or managers.
[1530] A "generative AI model" is an artificial intelligence model that generates appropriate responses to user input.
[1531] "Speech recognition means" refers to technology or devices that convert a user's speech into text in real time.
[1532] The "means for executing a program" refers to the software and hardware configuration for collecting, analyzing, and filtering the contents of the conversation and controlling the notification means.
[1533] "Feedback means" refers to a specific method or system that allows a superior to take appropriate action based on the notified information.
[1534] This invention is a dialogue agent system for providing mental care to operators and engineers working in factories. This system includes users (operators and engineers), terminals (smartphones, tablets, etc.), a server, and a dialogue agent (an all-affirmative bot).
[1535] System Configuration
[1536] The system works by coordinating the following components:
[1537] 1. User Interface: Provides an interface for the user to interact with the conversational agent, which can be operated by a touch screen or voice input.
[1538] 2. Conversational Agent: Engages in natural dialogue with the user. Conversational agents use generative AI models to always generate positive responses to user utterances.
[1539] 3. Speech recognition: Converts voice input into text in real time. This function uses the Google Speech-to-Text API, for example.
[1540] 4. Natural language processing: Analyzes the dialogue and determines the user's emotional state and stress level. This process uses natural language processing technologies such as GPT-3.5.
[1541] 5. Filtering means: Filters the dialogue content analyzed by the natural language processing means and summarizes important information.
[1542] 6. Notification: The information summarized by the filtering method is notified to superiors and managers via email or dashboard.
[1543] 7. Feedback measures: Provide feedback to supervisors and managers to take appropriate action based on the information provided.
[1544] Program processing
[1545] The server receives the dialogue between the conversational agent and the user in real time and generates a positive response using a generative AI model. The use of GPT-3.5 as the generative AI model allows for natural dialogue. All dialogue is saved in text format and later analyzed using natural language processing.
[1546] The device receives voice input from the user and converts it into text using a speech recognition method (e.g., Google Speech-to-Text API), which then transmits the user's speech as text to the server in real time.
[1547] The filtering means processes the acquired text data and extracts important keywords and phrases, and based on this information, summarizes the conversation and notifies superiors or managers.
[1548] The notification method sends filtered summary information to superiors and managers via email or a dedicated dashboard, allowing them to quickly take the necessary measures to care for the user's mental health.
[1549] Specific examples
[1550] For example, if a factory operator uses the system and says, "I'm so busy today, I'm tired," the all-affirmation bot will respond, "That must have been tough, good job!" After the conversation ends, the system summarizes the content of the conversation and provides feedback to the supervisor or manager that "the operator's fatigue due to being too busy is the cause."
[1551] Prompt Sentence Examples
[1552] Input prompt: "I'm too busy and tired today."
[1553] Model response: "That was hard work, good job!"
[1554] By implementing this system, it is possible to improve the working environment within the factory, maintain the motivation of operators and engineers, and reduce stress.
[1555] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1556] Step 1:
[1557] The user performs voice input. Specifically, the user speaks into the device's microphone to start a conversation. The input is what the user says, and is captured by the device as voice data.
[1558] Step 2:
[1559] A speech recognition means is activated on the terminal. The speech recognition means converts the user's voice data into text data. This process uses speech recognition technology such as the Google Speech-to-Text API. The output is data that converts the user's speech into text.
[1560] Step 3:
[1561] The device sends text data to the server, which then passes it to the generative AI model. The input is the converted text data, which is then sent to the server.
[1562] Step 4:
[1563] The conversational agent in the server uses a generative AI model to generate a positive response to the user's utterance. Specifically, a GPT-3.5 model is used to generate a prompt sentence based on the input text. The output is a positive response text.
[1564] Step 5:
[1565] The generated positive response text is returned from the server to the terminal. The terminal reproduces this text data as voice through the voice output means. The input is the generated response text, and the output is a voice response to the user.
[1566] Step 6:
[1567] The server stores the dialogue content as a log, and natural language processing means analyzes the text data. In this analysis process, the user's emotional state and stress level are determined. The input is the user's entire dialogue text, and the output is the evaluation results of the emotional state and stress.
[1568] Step 7:
[1569] The filtering means extracts important keywords and phrases from the dialogue text and generates a summary. In this step, points of particular interest are organized from the dialogue content. The input is the analysis result by the natural language processing means, and the output is the summary text.
[1570] Step 8:
[1571] The summarized information is sent to a supervisor or manager through a notification mechanism, which can be email or a dashboard. The input is the summary text and the output is the notification message provided to the supervisor or manager.
[1572] Step 9:
[1573] The supervisor or manager takes appropriate action based on the notified information. This involves implementing specific measures and countermeasures using feedback methods. The input is the notification message, and the output is specific feedback and countermeasures.
[1574] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1575] The present invention is a dialogue agent system that utilizes an all-affirmation bot, with the aim of providing mental care and communication assistance to users. By incorporating an emotion engine into this system, it is possible to more accurately recognize the user's emotional state and provide appropriate feedback. Specific embodiments of this system are described below.
[1576] System Configuration
[1577] This system includes users (employees), devices (PCs, smartphones, etc.), a server, a dialogue agent (a completely affirmative bot), and an emotion engine. These components work together to understand the user's mental state and provide appropriate feedback to superiors.
[1578] User Registration and Login
[1579] The user first launches the application and registers by entering the required information such as name, email address, and employee ID on the user registration screen. The device sends the entered information to the server, which stores it in a database. After registration is complete, the user logs in by entering their email address and password on the login screen. The server authenticates the user, and if authentication is successful, displays the user's dashboard.
[1580] Conversation with an all-affirmation bot and emotion recognition
[1581] The user clicks the "Start conversation" button on the dashboard screen to begin a conversation with the conversational agent (all-affirmative bot). The device starts voice recognition, and the user begins speaking. The all-affirmative bot always returns a positive response to the user's statements. For example, if the user says, "I'm very tired today," the bot replies, "That must have been tough! You've done well!"
[1582] The emotion engine analyzes the user's voice tone, facial expressions, and text content to recognize their emotions. The recognized emotional information is reflected in the conversational agent's responses. For example, if the user is determined to be very tired, the all-affirmation bot will adjust its responses to be more warmhearted.
[1583] Analysis of call content and emotional information
[1584] When the conversation ends, the device notifies the server that the call has ended. The server then sends the saved conversation log and emotion recognition results to a natural language processing model, which begins analyzing the text content and emotional information. The generative AI model extracts important keywords and emotional trends from the conversation content and creates a summary. The server then applies a social filter to the generated summary to organize the content that should be recognized. For example, information such as "There are too many meetings this week, so work is not progressing" and "The user is feeling tired" is summarized as "Work delays due to busyness are the cause of stress."
[1585] Information Sharing and Notification
[1586] The server prepares the filtered summary to be sent to the supervisor. The supervisor logs in to a dedicated dashboard from their own device and checks the feedback on the employee's mental state. The server responds to the supervisor's actions by displaying a concise summary of the important parts and providing hints on what points to pay attention to. If necessary, the supervisor can provide individual feedback or send a message to offer support.
[1587] Specific examples
[1588] Example 1: Everyday dialogue and emotion recognition
[1589] The user schedules a conversation with the All-Affirmation Bot every day at 2:00 PM. Today, the user says, "There are too many meetings this week, and I'm not making progress on my work." The All-Affirmation Bot replies, "That must have been tough, you've worked hard!" At this point, the emotion engine recognizes a strong sense of fatigue from the user's tone of voice. The server records the statement, "There are too many meetings this week, and I'm not making progress on my work," along with the "strong sense of fatigue," and the generative AI model summarizes it as, "The delays in work due to the many meetings are a stressful factor and cause fatigue." The server applies a sociality filter to the above summary and shares it with the supervisor, concluding, "The user is experiencing stress due to work delays caused by being busy, and is feeling very fatigued."
[1590] Example 2: Critical feedback and emotion recognition
[1591] In conversation with the All-Affirmation Bot, the user repeatedly states, "The project deadline is too tight." The All-Affirmation Bot responds, "I understand the pressure. You're doing a great job!" At this point, the emotion engine recognizes the user's strong stress from their facial expressions and voice. The server summarizes this as "The tight project deadline is the main cause of stress," and immediately notifies this information to the supervisor. The supervisor can then take action based on the employee's feedback, such as "reevaluating the project deadline."
[1592] This makes it easier for users to reduce stress and provides supervisors with information to provide effective mental care.The introduction of the emotion engine allows for a more accurate understanding of the user's emotional state, improving the quality of responses and feedback.
[1593] The processing flow will be explained below.
[1594] Step 1:
[1595] The user launches the application, enters the necessary information such as "name," "email address," and "employee ID" on the user registration screen, and presses the "Register" button.
[1596] Step 2:
[1597] The terminal temporarily stores the input user information and transmits the data to the server.
[1598] Step 3:
[1599] The server stores the received user information in a database and returns a response indicating successful registration to the terminal.
[1600] Step 4:
[1601] The terminal displays a "Registration Complete" message to the user.
[1602] Step 5:
[1603] The user enters their "email address" and "password" on the login screen and presses the "Login" button.
[1604] Step 6:
[1605] The terminal transmits the entered login information to the server.
[1606] Step 7:
[1607] The server authenticates the user using a database, and if authentication is successful, returns the user's dashboard information to the terminal.
[1608] Step 8:
[1609] The terminal displays a dashboard screen to the user.
[1610] Step 9:
[1611] The user clicks the "Start conversation" button on the dashboard screen.
[1612] Step 10:
[1613] The terminal starts voice recognition and the user begins speaking.
[1614] Step 11:
[1615] The conversational agent (all-affirmative bot) responds affirmatively to user comments. For example, if the user says, "I'm very tired today," the agent responds, "That must have been hard! You did a great job!"
[1616] Step 12:
[1617] The emotion engine recognizes the user's emotions by analyzing their voice tone, facial expressions, and text content, for example, reading tiredness from their voice and sadness from their facial expressions.
[1618] Step 13:
[1619] The terminal transmits the content of the conversation between the all-affirmation bot and the user, as well as the emotional data obtained from the emotion engine, to the server in real time.
[1620] Step 14:
[1621] The server stores the dialogue content and emotional data as logs, and performs real-time processing such as emotion analysis and keyword extraction as needed.
[1622] Step 15:
[1623] When the interaction ends, the terminal notifies the server of the interaction end.
[1624] Step 16:
[1625] The server sends the saved dialogue logs and emotion data to a natural language processing model, which begins analyzing the text content and emotion information.
[1626] Step 17:
[1627] The generative AI model extracts important keywords and emotional trends from the conversation and creates a summary, such as, "The user has been busy this week, and is feeling stressed and tired, especially with so many meetings."
[1628] Step 18:
[1629] The server applies a social filter to the generated summary to organize the content that should be recognized, for example, summarizing it as "Work delays due to busyness are a cause of stress."
[1630] Step 19:
[1631] The server prepares to send the filtered summary to the superior.
[1632] Step 20:
[1633] Supervisors can log into a dedicated dashboard from their own devices and check feedback on employees' mental health.
[1634] Step 21:
[1635] The server displays a concise summary of important points in response to the supervisor's operation, and also provides hints on what points to pay attention to.
[1636] Step 22:
[1637] If needed, a manager can send a message to provide individual feedback or assistance.
[1638] Step 23:
[1639] Users receive feedback from their superiors and implement advice on mental care and work improvement.
[1640] The above are the specific processing steps for a mental care and communication assistance application using a fully affirmative bot combined with an emotion engine.
[1641] Example 2
[1642] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1643] To effectively provide mental care to users, it is necessary to accurately grasp the user's emotional state and stress level and provide appropriate feedback accordingly. However, conventional systems lack the means to accurately recognize the user's emotions, making it difficult to provide appropriate responses and feedback. In addition, sharing information with superiors is cumbersome, making it difficult to provide effective mental care to users.
[1644] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1645] In this invention, the server includes means for having a user interface through which the user interacts with the dialogue agent, natural language processing means for analyzing the dialogue between the dialogue agent and the user to determine the user's emotional state and level of stress, filtering means for filtering and summarizing the dialogue analyzed by the natural language processing means, emotion recognition means for analyzing the user's voice tone, facial expressions, and text content to recognize the user's emotions, adjustment means for reflecting the emotion information obtained by the emotion recognition means in the dialogue agent's response content, and notification means for notifying a superior of the information processed by the filtering means. This makes it possible to accurately grasp the user's emotional state and provide appropriate feedback.
[1646] Below are definitions of important terms included in the rewritten claims:
[1647] "User interface" refers to the screen or means by which a user accesses and operates a system.
[1648] An "interactive agent" is a program or system that interacts with a user and responds based on the interaction.
[1649] "Natural language processing means" refers to a technique or means for analyzing the content of a dialogue between a user and a dialogue agent and determining the user's emotional state and stress level.
[1650] "Filtering means" refers to a technology or means for sorting the dialogue content analyzed by natural language processing means and extracting and summarizing important information.
[1651] The "emotion recognition means" refers to a technology or means for analyzing the user's voice tone, facial expression, and text content to recognize the user's emotions.
[1652] The "adjustment means" refers to a technique or means for appropriately changing or adjusting the response content of the dialogue agent based on the emotional information obtained by the emotion recognition means.
[1653] "Notification means" refers to a technique or means for transmitting the information processed by the filtering means to a superior in an appropriate format.
[1654] This invention is a dialogue agent system that uses an all-affirmation bot and an emotion engine, with the aim of providing mental care and communication assistance to users. Specific embodiments of this system are described below.
[1655] System Configuration
[1656] This system includes a user, a terminal, a server, a dialogue agent (a fully affirmative bot), and an emotion engine. These elements work together to understand the user's mental state and provide appropriate feedback.
[1657] User Registration and Login
[1658] 1. The user launches the application and registers by entering information such as their name, email address, and employee ID.
[1659] 2. The terminal sends the entered information to the server, which stores it in a database.
[1660] 3. After completing registration, the user logs in by entering their email address and password.
[1661] 4. The server authenticates the user and, if authentication is successful, displays the user's dashboard.
[1662] Conversation with the all-affirmation bot
[1663] 1. The user clicks the "Start conversation" button on the dashboard screen to begin a conversation with the all-affirmation bot.
[1664] 2. The device starts voice recognition and sends the user's voice to the all-affirmation bot. For example, if the user says, "I'm tired today," the all-affirmation bot replies, "That must have been hard! Good job!"
[1665] emotion recognition
[1666] 1. The device simultaneously transmits the user's voice tone, facial expression, and text content to the emotion engine.
[1667] 2. The emotion engine analyzes this data and recognizes the user's emotions.
[1668] 3. The recognized emotional information is reflected in the responses of the all-positive bot. For example, if the user is judged to be very tired, the all-positive bot will adjust its responses to be more warmhearted.
[1669] Call analysis
[1670] 1. When the conversation ends, the terminal notifies the server that the call has ended.
[1671] 2. The server sends the dialogue log and emotion recognition results to the natural language processing model.
[1672] 3. The generative AI model extracts important keywords and emotional trends from the conversation and creates a summary. For example, the information that "the user is feeling tired" and the statement that "there are too many meetings this week and work isn't progressing" can be summarized as "work delays due to being busy are the cause of stress."
[1673] Information Sharing and Notification
[1674] 1. The server prepares the filtered summary for notification to the supervisor.
[1675] 2. Supervisors log in to a dedicated dashboard and check feedback on employees' mental health.
[1676] 3. The server responds to the supervisor's actions by displaying a concise summary of the important points and providing hints on what points the supervisor should pay attention to. If necessary, the supervisor can provide feedback or send a support message.
[1677] Specific examples
[1678] Example 1: Everyday dialogue and emotion recognition
[1679] Consider a scenario where a user says, "I have too many meetings this week and I'm not getting any work done."
[1680] The all-affirmation bot responds with, "That must have been tough, good job!" At this point, the emotion engine recognizes the strong sense of fatigue from the user's tone of voice and records it on the server.
[1681] The generative AI model summarizes that "work delays caused by numerous meetings are a source of stress and fatigue."
[1682] The server applies a social filter and shares the information with superiors, stating that "users are stressed by delays in work due to being busy, and feel very tired."
[1683] Example 2: Critical feedback and emotion recognition
[1684] Consider a situation where a user repeatedly says, "The project deadline is too tight."
[1685] The all-affirmation bot replies, "I understand the pressure. You're doing a great job!"
[1686] The emotion engine recognizes strong stress from the user's facial expressions and voice.
[1687] The server summarizes that "tight project deadlines are the main cause of stress" and immediately communicates this information to his superiors.
[1688] Based on the feedback, the supervisor can take action such as "reevaluating the project deadline."
[1689] Example prompts for generative AI models
[1690] (Example prompt 1: Summarizing everyday conversations)
[1691] User says: "I have too many meetings this week and I can't get any work done."
[1692] Recognizing the Emotion Engine: Extreme Fatigue
[1693] Summary generated: Work delays due to too many meetings are a major cause of stress and fatigue
[1694] (Example prompt 2: Summary of important feedback)
[1695] User says: "The project deadline is too tight"
[1696] Recognizing the Emotion Engine: High Stress
[1697] Generate summary: Tight project deadlines are a major source of stress
[1698] This system makes it possible to accurately grasp the user's emotional state and provide appropriate feedback, thereby improving the effectiveness of mental care.
[1699] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1700] Step 1:
[1701] The user launches the application and registers by entering their name, email address, and employee ID.
[1702] Input: Name, Email Address, Employee ID
[1703] How it works: The device sends this information to the server.
[1704] Output: User information is sent to the server and a registration completion message is received.
[1705] Step 2:
[1706] The server stores the received information in a database and sends a notification of registration completion to the terminal.
[1707] Input: User information (name, email address, employee ID)
[1708] What happens: The server creates a new record in the database and saves the information.
[1709] Output: A message that registration is complete is displayed on the terminal.
[1710] Step 3:
[1711] The user enters their email address and password and clicks the login button.
[1712] Input: Email address, Password
[1713] How it works: The device sends this information to the server.
[1714] Output: Login information sent to the server.
[1715] Step 4:
[1716] The server compares the entered information with a database and performs authentication.
[1717] Input: Email address, Password
[1718] How it works: The server checks the information in its database and performs authentication.
[1719] Output: If authentication is successful, information is sent to the device to display the user's dashboard.
[1720] Step 5:
[1721] The user clicks the "Start conversation" button on the dashboard screen.
[1722] Input: Click the "Start conversation" button
[1723] Action: The device will begin voice input and perform the initial setup to connect to the all-affirmation bot.
[1724] Output: Ready to connect with all affirmative bots.
[1725] Step 6:
[1726] When the user starts speaking, the device converts the speech into text and sends it to the All Affirmations Bot.
[1727] Input: User's voice data
[1728] Operation: The device performs voice recognition and sends the recognized text data to all affirmative bots.
[1729] Output: Text data sent to all affirmation bots.
[1730] Step 7:
[1731] An all-affirmation bot always responds affirmatively to what the user says.
[1732] Input: Text data based on speech recognition
[1733] How it works: The all-affirmation bot generates appropriate affirmative responses. For example, if someone says "I'm tired today," it will respond with "That must have been hard, good job!"
[1734] Output: Affirmative response to the user.
[1735] Step 8:
[1736] The device sends the user's voice and text data to the emotion engine.
[1737] Input: User voice and text data
[1738] How it works: The emotion engine analyzes this data and recognizes emotions.
[1739] Output: Recognized emotion information.
[1740] Step 9:
[1741] The emotion engine feeds back the analysis results to the dialogue agent and adjusts the response content.
[1742] Input: Emotion recognition results
[1743] How it works: The bot updates its responses based on the results of the emotion engine. For example, if it determines that you are tired, it will respond with something like, "Maybe it would be good to take a break."
[1744] Output: The adjusted response.
[1745] Step 10:
[1746] When the conversation ends, the terminal notifies the server that the call has ended.
[1747] Input: Call end event
[1748] Action: The device sends a call end notification to the server.
[1749] Output: A conversation termination notification is sent to the server.
[1750] Step 11:
[1751] The server sends the call logs and emotion recognition results to the natural language processing model.
[1752] Input: Call logs, emotion recognition results
[1753] How it works: The server sends this data to the natural language processing model and begins analysis.
[1754] Output: The data required for analysis is sent.
[1755] Step 12:
[1756] The generative AI model extracts important keywords and emotional trends from the dialogue and creates a summary.
[1757] Input: Call logs, emotion recognition results
[1758] How it works: The generative AI model analyzes data, extracts key keywords and sentiment trends, and generates summaries. For example, if a user says, "I'm having too many meetings this week, so I'm not making progress on my work," the model summarizes that "Work delays due to being too busy are causing stress."
[1759] Output: Summarized information.
[1760] Step 13:
[1761] The server applies a sociality filter to the generated summary to organize the content to be recognized.
[1762] Input: Summary information
[1763] How it works: The server applies social filters to determine what should be recognized.
[1764] Output: Filtered summary information.
[1765] Step 14:
[1766] The server prepares to send the filtered summary to the supervisor.
[1767] Input: Filtered summary information
[1768] What it does: The server formats the information for notification.
[1769] Output: Filtered summary for notifying superiors.
[1770] Step 15:
[1771] Managers can log in to a dedicated dashboard and check feedback on employees' mental health.
[1772] Input: Filtered summary information
[1773] What happens: Your manager accesses the dashboard and sees the feedback you provided.
[1774] Output: Your supervisor reviews the feedback and is ready to take action if necessary.
[1775] (Application example 2)
[1776] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1777] Conventional conversational agent systems have difficulty accurately grasping a user's emotional state and stress level, making it difficult to provide appropriate feedback. Furthermore, there is a demand for systems that enable sales staff, especially in brick-and-mortar stores, to grasp and respond to customer emotions in real time.
[1778] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means having a user interface for the user to interact with the dialogue agent, natural language processing means for analyzing the content of the dialogue between the dialogue agent and the user to determine the user's emotional state and stress level, filtering means for filtering and summarizing the content of the dialogue analyzed by the natural language processing means, notification means for notifying an administrator of the information processed by the filtering means, and device linking means for recognizing the user's emotional state in real time using smart glasses and displaying feedback. This enables sales staff to grasp customer emotions in real time and respond optimally.
[1779] A "user interface" is a screen or operating means that allows a user to interact with or operate a system.
[1780] A "dialogue agent" is software or a system that interacts with a user and supports communication using natural language processing.
[1781] "Natural language processing means" is a technology for analyzing the content of conversation between a dialogue agent and a user and understanding the user's emotions and intentions.
[1782] The "filtering means" is a technology for organizing information analyzed by the natural language processing means based on specific criteria and extracting necessary information.
[1783] The "notification means" is a function or device for transmitting the information sorted by the filtering means to the administrator.
[1784] "Smart glasses" are eyeglass-type wearable devices that have a display function and provide information to the user.
[1785] "Device integration means" is a technology that links devices such as smart glasses with systems to display and update information in real time.
[1786] "Administrator" is the person or department responsible for overseeing the status of the system and users and providing necessary support and feedback.
[1787] "Real time" refers to a state in which processing or communication is carried out almost immediately after the data is generated.
[1788] This invention is a system for users to receive mental health care through dialogue with a conversational agent. The system uses smart glasses to recognize the user's emotional state in real time and provide appropriate feedback.
[1789] Hardware and software used
[1790] The hardware used includes:
[1791] Smart glasses (e.g. Google Glass, Vuzix Blade)
[1792] Server (e.g. AWS EC2)
[1793] The software used includes:
[1794] Emotion recognition engine (e.g. Microsoft Azure Emotion API)
[1795] Natural language processing models (e.g., OpenAI GPT-4)
[1796] Front-end applications (e.g. React Native)
[1797] Database (e.g. MySQL)
[1798] Overall processing of the program
[1799] The system analyzes the user's emotional state and dialogue content in real time and provides optimal feedback.
[1800] 1. User Registration and Login
[1801] First, the user starts the application and registers by entering basic information on the user registration screen. The terminal sends the entered information to the server, which stores it in a database. After registration, the user can log in and access the system.
[1802] 2. Interaction with a conversational agent
[1803] The user wears the smart glasses and starts a dialogue with the conversational agent. The smart glasses capture the user's voice and facial expressions and send them to an emotion recognition engine. The emotion recognition engine analyzes the user's emotional state and sends it to the server in real time.
[1804] 3. Emotion Recognition and Feedback
[1805] The server then sends the data received from the emotion recognition engine to a natural language processing model to generate appropriate feedback, which is then displayed in real time on the smart glasses display for the user to review.
[1806] Specific examples
[1807] Example prompt sentence 1:
[1808] Generate a conversational agent response and emotion recognition results when a user says, "Tell me more about this product." The customer's expression is filled with curiosity.
[1809] Example system response:
[1810] Emotion recognition result: "Curiosity"
[1811] All-Affirmative bot replies: "I understand. The special feature of this product is that it is made using the latest technology and is extremely durable."
[1812] Example prompt sentence 2:
[1813] Generate a conversational agent response and emotion recognition results when a user says, "This product was a disappointment." The customer's tone sounds angry.
[1814] Example system response:
[1815] Emotion recognition result: "Anger"
[1816] The all-affirmation bot's response: "I'm sorry you felt that way. I'll get back to you shortly. Can you please provide more details?"
[1817] This system allows users to receive appropriate feedback in real time, allowing managers to accurately grasp the user's emotional state and respond promptly, which is expected to improve customer satisfaction in physical stores.
[1818] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1819] Step 1:
[1820] The user starts the application and registers by entering the required information such as name, email address, employee ID, etc. on the user registration screen. The entered information is sent from the device to the server, which stores it in a database. Once the user has completed registration, a login screen is displayed. Here, the user logs in by entering their email address and password. The server authenticates the user, and if authentication is successful, the user's dashboard is displayed.
[1821] Input: Registration information (name, email address, employee ID), login information (email address, password)
[1822] Output: User's dashboard
[1823] Step 2:
[1824] By clicking the "Start Dialogue" button on the dashboard screen, the user begins a dialogue with the conversational agent (all-affirmative bot). The device starts voice recognition, and the user begins speaking. The voice data is sent to the server, where it is input into the emotion recognition engine.
[1825] Input: User's voice data
[1826] Output: Input to the emotion recognition engine
[1827] Step 3:
[1828] The emotion recognition engine analyzes the user's voice tone, facial expressions, and text content to recognize the user's emotional state. The recognized emotional information is used as data for generating appropriate feedback through a natural language processing model.
[1829] Input: User's voice tone, facial expressions, and text content
[1830] Output: Recognized emotion information
[1831] Step 4:
[1832] The recognized emotional information and dialogue content are sent to a natural language processing model, which then uses this data to generate appropriate feedback for the user, which is then displayed on the smart glasses display in real time.
[1833] Input: Recognized emotion information, dialogue content
[1834] Output: Generated feedback
[1835] Step 5:
[1836] When the conversation ends, the device notifies the server of the end of the call. The server then analyzes the conversation using a natural language processing model based on the stored conversation log and emotion recognition results. The generated summary is then further processed by a filtering means to extract important information.
[1837] Input: Dialogue log, emotion recognition results
[1838] Output: A summary with the necessary information extracted
[1839] Step 6:
[1840] The filtered summary is sent to the administrator via a notification mechanism. The administrator can then log in to a dedicated dashboard and check the user's feedback on mental health care. The notification mechanism displays important summary parts according to the administrator's actions and also provides hints on points that require attention.
[1841] Input: Filtered summary
[1842] Output: Administrator notification, hints on what to note
[1843] Step 7:
[1844] Administrators can provide individual feedback based on the user's feedback or send messages to offer assistance, and this information is reflected in the system and used for future interactions.
[1845] Input: Admin feedback, support message
[1846] Output: Information reflected in subsequent interactions
[1847] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1848] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1849] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1850] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1851] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1852] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1853] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1854] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1855] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1856] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1857] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1858] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1859] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1860] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1861] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1862] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1863] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1864] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1865] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1866] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1867] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1868] The following is further disclosed regarding the above embodiment.
[1869] (Claim 1)
[1870] means having a user interface that provides a way for a user to interact with the interactive agent;
[1871] natural language processing means for determining the emotional state and stress level of a user by analyzing the content of a dialogue between a dialogue agent and the user;
[1872] a filtering means for filtering and summarizing the dialogue content analyzed by the natural language processing means;
[1873] a notification means for notifying a superior of the information processed by the filtering means;
[1874] A system including:
[1875] (Claim 2)
[1876] 2. The system of claim 1, wherein the dialogue agent always responds with a positive response.
[1877] (Claim 3)
[1878] 2. The system according to claim 1, wherein the notification means provides feedback regarding the user's mental care to the superior, and encourages the superior to take appropriate action.
[1879] "Example 1"
[1880] (Claim 1)
[1881] means having a user interface that provides a way for a user to interact with the interactive agent;
[1882] a means for a user to initiate a dialogue with a dialogue agent through the user interface and record the dialogue content;
[1883] a means for the dialogue agent to have a generative AI model that always generates a positive response;
[1884] natural language processing means for determining the emotional state and stress level of a user by analyzing the content of a dialogue between a dialogue agent and the user;
[1885] a means for summarizing the dialogue content analyzed by the natural language processing means and extracting important keywords and emotional tendencies;
[1886] a notification means for notifying a superior of the information processed by the filtering means;
[1887] a means for providing feedback regarding the user's mental care to a superior by the notification means and encouraging the superior to take appropriate action;
[1888] A system including:
[1889] (Claim 2)
[1890] 2. The system of claim 1, wherein the generative AI model always returns a positive response.
[1891] (Claim 3)
[1892] 2. The system according to claim 1, wherein the notification means provides feedback to the supervisor regarding the user's mental care and presents hints regarding important keywords and emotional tendencies.
[1893] "Application Example 1"
[1894] (Claim 1)
[1895] means having a user interface that provides a way for a user to interact with the interactive agent;
[1896] natural language processing means for determining the emotional state and stress level of a user by analyzing the content of a dialogue between a dialogue agent and the user;
[1897] a filtering means for filtering and summarizing the dialogue content analyzed by the natural language processing means;
[1898] a notification means for notifying a superior of the information processed by the filtering means;
[1899] A means for using a generative AI model to recognize a user's everyday speech and generate a positive response;
[1900] a speech recognition means for converting speech input from a user into text in real time;
[1901] means for executing a program that controls a notification means for summarizing and notifying the contents of the dialogue;
[1902] A feedback method for superiors to take appropriate action based on the notified information;
[1903] A system including:
[1904] (Claim 2)
[1905] 10. The system of claim 1, wherein the dialogue agent uses a generative AI model that always returns a positive response.
[1906] (Claim 3)
[1907] 2. The system according to claim 1, wherein the notification means provides a supervisor with summary information regarding the user's mental care, and encourages the supervisor to take appropriate action.
[1908] "Example 2: Combining Emotion Engines"
[1909] (Claim 1)
[1910] means having a user interface for a user to interact with the dialogue agent;
[1911] natural language processing means for determining the emotional state and stress level of a user by analyzing the content of a dialogue between a dialogue agent and the user;
[1912] a filtering means for filtering and summarizing the dialogue content analyzed by the natural language processing means;
[1913] emotion recognition means for recognizing the user's emotion by analyzing the user's voice tone, facial expression, and text content;
[1914] an adjustment means for reflecting the emotional information obtained by the emotion recognition means in the response content of the dialogue agent;
[1915] a notification means for notifying a superior of the information processed by the filtering means;
[1916] A system including:
[1917] (Claim 2)
[1918] 2. The system of claim 1, wherein the dialogue agent always responds with a positive response.
[1919] (Claim 3)
[1920] 2. The system according to claim 1, wherein the notification means provides feedback regarding the user's mental care to the superior, and encourages appropriate action.
[1921] "Application example 2 when combining emotion engines"
[1922] (Claim 1)
[1923] means having a user interface for a user to interact with the dialogue agent;
[1924] natural language processing means for determining the emotional state and stress level of a user by analyzing the content of a dialogue between a dialogue agent and the user;
[1925] a filtering means for filtering and summarizing the dialogue content analyzed by the natural language processing means;
[1926] a notification means for notifying an administrator of the information processed by the filtering means;
[1927] A device cooperation means for recognizing a user's emotional state in real time using smart glasses and displaying feedback;
[1928] A system including:
[1929] (Claim 2)
[1930] 2. The system of claim 1, wherein the dialogue agent always responds with a positive response.
[1931] (Claim 3)
[1932] 2. The system according to claim 1, wherein the notification means provides feedback to the administrator regarding the user's mental care and urges the administrator to take appropriate action. [Explanation of symbols]
[1933] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means having a user interface that provides a way for a user to interact with the interactive agent; natural language processing means for determining the emotional state and stress level of a user by analyzing the content of a dialogue between a dialogue agent and the user; a filtering means for filtering and summarizing the dialogue content analyzed by the natural language processing means; a notification means for notifying a superior of the information processed by the filtering means; A system including:
2. 2. The system of claim 1, wherein the dialogue agent always responds with a positive response.
3. 2. The system according to claim 1, wherein the notification means provides feedback regarding the user's mental care to the superior, and encourages the superior to take appropriate action.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A