System

The system addresses the lack of comprehensive emotional health management by collecting and analyzing voice and message data to provide personalized improvement suggestions, enhancing emotional health management through accurate emotional state assessment and visualization.

JP2026014900APending Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024116374
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Conventional health management systems primarily focus on physical data and lack comprehensive emotional health management, making it difficult to accurately grasp emotional fluctuations and provide personalized improvement suggestions.

Method used

A system that collects voice and message data, converts it into text, integrates and analyzes sentiment, and generates improvement suggestions based on emotional analysis, with timestamp organization and visualization for intuitive understanding.

Benefits of technology

Enables accurate emotional state assessment and provides specific improvement suggestions, facilitating better emotional health management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026014900000001_ABST
    Figure 2026014900000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for collecting voice data; means for converting the collected voice data into text data; means for collecting message data; means for integrating the voice data and the message data; means for performing sentiment analysis on the integrated data; and means for generating an improvement suggestion based on a sentiment analysis result.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In modern society, people's physical and mental health is increasingly important, but managing their emotional health in particular is extremely difficult. Conventional health management systems are primarily based on physical data, and few comprehensively manage emotional health. Therefore, there is a demand for a system that can accurately grasp the emotional fluctuations in a user's daily life and provide personalized improvement suggestions based on their emotions, but there is a lack of concrete means to achieve this. [Means for solving the problem]

[0005] To solve the above problems, the present invention has the following features. By providing a system including a means for collecting voice data, a means for converting the collected voice data into text data, a means for collecting message data, a means for integrating the voice data and the message data, a means for sentiment analysis of the integrated data, and a means for generating improvement suggestions based on the sentiment analysis results, it becomes possible to accurately grasp a user's daily emotional fluctuations and make appropriate improvement suggestions. Furthermore, by further including a means for assigning a timestamp to each voice data and message data and organizing the sentiment analysis results in chronological order, it becomes possible to visualize emotional fluctuations by daily activity and provide more specific data to the user. Furthermore, by including a means for converting the sentiment analysis results into a visual format such as a graph or chart, it becomes easier for the user to intuitively understand their own emotional state.

[0006] "Audio data" refers to digital data that records a user's speech or conversation as sound.

[0007] "Text data" refers to data containing character information obtained by converting voice data into text format.

[0008] "Message data" refers to data that includes information about text messages sent and received by users.

[0009] "Means of collection" refers to hardware or software mechanisms for obtaining voice data or message data.

[0010] The "means of conversion" refers to the voice recognition technology or algorithm that automatically converts voice data into text data.

[0011] The "integrating means" is a mechanism for combining voice data and message data obtained from different data sources into a single data set.

[0012] "Means for emotion analysis" refers to algorithms and technologies for analyzing a user's emotional state (happiness, sadness, anger, etc.) from text data.

[0013] The "means for generating improvement suggestions" is a mechanism that automatically creates appropriate suggestions and advice for users based on the results of sentiment analysis.

[0014] A "timestamp assignment means" is a mechanism for adding the time of generation and other time information to each piece of data.

[0015] "Visualization means" refers to a mechanism that converts the results of sentiment analysis into a visual format such as a graph or chart, and displays it in a way that allows users to intuitively understand it. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] The present invention is a system that collects voice data and text data, integrates the data, performs sentiment analysis, and generates improvement suggestions for users based on emotional fluctuations. To realize this system, hardware and software with the following functions are required.

[0038] System configuration

[0039] 1. Device for collecting voice data

[0040] The user's voice is recorded by using a meeting application or the like.

[0041] 2. Device for collecting message data

[0042] Get text messages sent and received from LINE and other messaging applications.

[0043] 3. Data analysis server

[0044] A speech recognition engine for converting collected voice data into text data.

[0045] An analysis engine that integrates message data and text data converted from voice to perform sentiment analysis.

[0046] An engine for generating specific improvement suggestions for users based on the results of sentiment analysis.

[0047] System operation procedure

[0048] Data collection

[0049] Audio data collection

[0050] A user speaks during a conference using a meeting application.

[0051] The device records what is said, generates an audio file, and sends it to the server.

[0052] Message Data Collection

[0053] A user sends a message via LINE.

[0054] The device receives the LINE text message and sends it to the server.

[0055] Data Preprocessing

[0056] Converting voice data to text

[0057] The server inputs the received voice file into a voice recognition engine, which converts the voice into text data.

[0058] Text data integration

[0059] The server integrates the converted voice data with the message data obtained from LINE.

[0060] The integrated data is cleansed to remove unnecessary symbols and spaces.

[0061] sentiment analysis

[0062] Sentiment analysis of text data

[0063] The server inputs the cleansed text data into a sentiment analysis engine.

[0064] A sentiment analysis engine calculates an emotion score (e.g., happy, sad, anger) for each text block.

[0065] Data integration

[0066] Timestamp of emotion data

[0067] The server assigns a timestamp to each piece of data and organizes it in chronological order.

[0068] Visualizing Emotion Data

[0069] The server converts the emotional data into a visual format such as graphs and charts, and displays it so that users can intuitively understand the fluctuations in their emotions throughout the day.

[0070] Proposal Generation

[0071] Generate improvement suggestions

[0072] The server analyzes the user's emotional tendencies based on the emotion analysis results.

[0073] The server generates appropriate improvement suggestions for the user based on the analysis results (e.g., "Try taking deep relaxation breaths before a meeting").

[0074] Specific examples

[0075] Example 1:

[0076] A user says during a meeting, "What are your tasks for today?"

[0077] The device records what is said and sends the audio file to the server.

[0078] The server converts the audio file into text data: "What is your task today?"

[0079] A user sends a message on LINE asking, "How did the afternoon meeting go?"

[0080] The terminal sends the message to the server.

[0081] The server combines both sets of text data and feeds it into a sentiment analysis engine.

[0082] The emotion analysis detects "impatience" and generates a suggestion based on that, such as "We recommend you relax before your afternoon meeting."

[0083] In this way, the system of the present invention mainly operates in cooperation with the server, the terminal, and the user, and constitutes a specific means for supporting the user's emotional health management.

[0084] The processing flow will be explained below.

[0085] Step 1:

[0086] A user speaks up during a meeting, for example, "What are our goals for next week?"

[0087] Step 2:

[0088] The device records what the user says and generates an audio file that is temporarily stored on the device.

[0089] Step 3:

[0090] The terminal transmits the recorded audio file to the server via the meeting application.

[0091] Step 4:

[0092] The server inputs the received voice file into a voice recognition engine, which converts the voice data into text data, generating a string of characters such as "What are your goals for next week?"

[0093] Step 5:

[0094] A user sends a message on LINE saying, "What day is the meeting next week?"

[0095] Step 6:

[0096] The device retrieves message data from the LINE application and sends it to the server.

[0097] Step 7:

[0098] The server combines the received LINE message data with the text data converted from the voice. At this time, a cleansing process is performed to unify the data and remove unnecessary symbols and spaces.

[0099] Step 8:

[0100] The server inputs the formatted text data into a sentiment analysis engine, which calculates sentiment scores for the texts "What are your goals for next week?" and "What day of the week is the meeting next week?"

[0101] Step 9:

[0102] The server assigns a timestamp to each block of text based on the sentiment score obtained from the sentiment analysis engine, for example, recording that "What are your goals for next week?" was spoken at 10:00 AM and "What day is our meeting next week?" was spoken at 10:05 AM.

[0103] Step 10:

[0104] The server chronologically organizes the time-stamped emotion data and summarizes the emotional fluctuations by daily activity. The emotion data is then converted into a visual format such as a graph or chart.

[0105] Step 11:

[0106] The server analyzes the user's emotional tendencies based on the emotional data and timestamp data, and obtains analysis results such as "people often feel anxious during morning meetings."

[0107] Step 12:

[0108] Based on the analysis results, the server generates improvement suggestions for the user, such as "I suggest you meditate for five minutes before your morning meeting."

[0109] Step 13:

[0110] The user checks the suggestions from the server on a device such as a smartphone and adjusts their behavior based on the advice.

[0111] Example 1

[0112] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0113] In today's communication environment, there is a lack of methods to properly understand users' emotional fluctuations and stress levels and support their emotional health management. In particular, there is a need for a method that can more accurately assess users' emotional state and provide effective improvement suggestions by integrating and analyzing voice data and text messages.

[0114] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0115] In this invention, the server includes means for collecting voice data, means for converting the collected voice data into text data, means for collecting text messages, means for integrating the voice data and the text messages, means for preprocessing the integrated data, means for sentiment analysis of the preprocessed data, and means for generating improvement suggestions based on the sentiment analysis results. This makes it possible to comprehensively analyze a user's various communication data, appropriately evaluate the emotional state, and provide effective improvement suggestions.

[0116] "Audio data" refers to digital audio files that record the content of user speech occurring during meetings, conversations, and the like.

[0117] "Text data" refers to character information converted from voice data or character information sent and received by a message application.

[0118] A "text message" is text information that a user sends and receives using a text application (e.g., a messaging app).

[0119] "Preprocessing" is the procedure of removing unnecessary information and extracting necessary parts in order to prepare data in a format that can be analyzed.

[0120] "Sentiment analysis" is the process of analyzing a user's emotional state (e.g., joy, sadness, anger) from text data and classifying it into scores or categories.

[0121] "Improvement suggestions" are specific advice on improving behavior or status that is provided to the user based on the results of sentiment analysis.

[0122] "Fusion" refers to combining multiple data sources (e.g., voice data and text messages) into a single dataset.

[0123] A "timestamp" is a means of clarifying time-series information by adding information about the date and time when data was generated or collected.

[0124] "Visualization" is the process of transforming data into a visually understandable format such as a graph or chart.

[0125] The present invention provides a system for collecting voice data and text messages, integrating the collected data to perform sentiment analysis, and generating improvement suggestions for users based on emotional fluctuations. Specific embodiments are described below.

[0126] Hardware and Software Configuration

[0127] 1. Device for collecting audio data:

[0128] A user speaks during a meeting using a meeting application (e.g., an online conference system).

[0129] The terminal records the user's speech during the conference, generates an audio file (e.g., a WAV file), and sends it to the server.

[0130] 2. Devices for Text Message Data Collection:

[0131] A user has a conversation using a messaging application (e.g., a messaging app).

[0132] The terminal retrieves the text message from the message application and sends it to the server as a text file.

[0133] 3. Data analysis server:

[0134] The server uses speech recognition and sentiment analysis engines such as Google Cloud Speech-to-Text and IBM Watson Natural Language Understanding to convert the voice data into text data and perform further sentiment analysis.

[0135] The server includes an engine for generating specific improvement suggestions for the user based on the sentiment analysis results.

[0136] Specific examples

[0137] Example 1:

[0138] A user says, "What are your tasks for today?" during an online meeting.

[0139] The device records what is said and sends the audio file to the server.

[0140] The server inputs the audio file into the Google Cloud Speech-to-Text engine and converts it into text data: "What is your task today?"

[0141] A user sends a message in a messaging app asking, "How did the meeting this afternoon go?"

[0142] The terminal receives the message and sends it to the server.

[0143] The server integrates both sets of text data and performs a cleansing process.

[0144] The server inputs the cleansed text data into IBM Watson Natural Language Understanding for sentiment analysis.

[0145] As a result of the emotion analysis, "impatience" is detected, and based on that, an improvement suggestion is generated, such as "recommending relaxation before the afternoon meeting."

[0146] Examples of prompt statements

[0147] Example prompt sentence:

[0148] "A user might say, 'What are your tasks for today?' in an online meeting, and then later send a message in a messaging app asking, 'How did your afternoon meeting go?' Combine these data points, perform sentiment analysis, and generate appropriate improvement suggestions."

[0149] In this way, the system of the present invention constitutes a concrete means for supporting the emotional health management of users, with the server, terminal, and user working together.

[0150] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0151] Step 1:

[0152] Collection and transmission of voice data

[0153] A user speaks during a conference using an online conference system.

[0154] The terminal records the user's speech and generates an audio file (e.g., WAV format).

[0155] The terminal transmits the generated audio file to the server.

[0156] Input: User's speech in an online conference system

[0157] Data processing: Recording audio data and generating WAV format audio files

[0158] Output: Audio file sent to the server

[0159] Specific operation:

[0160] Execute a script to record speech on the online conference system, generate an audio file, and upload the generated audio file to the server using FTP or HTTP protocol.

[0161] Step 2:

[0162] Collection and transmission of message data

[0163] A user sends a text message using a messaging application.

[0164] The terminal acquires text messages sent and received from the message application and saves them as text files.

[0165] The terminal transmits the obtained text file to the server.

[0166] Input: The user's text message in the Messages application

[0167] Data processing: Obtaining text messages and generating text files in TXT format

[0168] Output: A text file that is sent to the server.

[0169] Specific operation:

[0170] When a new message is detected, a script running in the background of the Messages app extracts its contents, saves the extracted text message in a text file, and uploads it to a server.

[0171] Step 3:

[0172] Speech-to-text

[0173] The server inputs the received audio file into a speech recognition engine (e.g., Google Cloud Speech-to-Text).

[0174] The server converts the audio file into text data using a speech recognition engine.

[0175] The server stores the converted text data in temporary storage.

[0176] Input: Audio file sent to the server

[0177] Data processing: Converting audio files into text using a speech recognition engine

[0178] Output: Text data saved in temporary storage

[0179] Specific operation:

[0180] A Python script is executed on the server to send the audio file to the Google Cloud Speech-to-Text API, and the text data returned by the API is retrieved in JSON format and saved in a database on the server.

[0181] Step 4:

[0182] Text data integration and preprocessing

[0183] The server integrates the text data converted from the voice data with the text file obtained from the messaging app.

[0184] The server cleanses the integrated text data, removing unnecessary symbols and spaces.

[0185] Input: Text data converted from voice, text data obtained from messaging apps

[0186] Data processing: text data integration and cleansing

[0187] Output: Integrated text data after cleansing

[0188] Specific operation:

[0189] The integration process runs an SQL query to combine the voice text and message text, and uses regular expressions (RegEx) on the combined data to remove unnecessary symbols and spaces.

[0190] Step 5:

[0191] Conducting sentiment analysis

[0192] The server inputs the cleansed text data into a sentiment analysis engine (e.g., IBM Watson Natural Language Understanding).

[0193] The server obtains the sentiment score (e.g., happy, sad, anger) for each text block returned by the sentiment analysis engine.

[0194] Input: Text data after cleansing

[0195] Data calculation: Calculating sentiment scores using a sentiment analysis engine

[0196] Output: Text data with sentiment scores

[0197] Specific operation:

[0198] The server makes an API request to send the cleansed text data to the sentiment analysis engine, which analyzes the resulting sentiment scores and stores them in a database.

[0199] Step 6:

[0200] Timestamp and organize emotion data

[0201] The server assigns a timestamp to each piece of emotion data and organizes it in chronological order.

[0202] Input: Text data with sentiment scores

[0203] Data processing: Adding timestamps and organizing data in chronological order

[0204] Output: Emotion data with timestamps and organized in chronological order

[0205] Specific operation:

[0206] The Python Pandas library is used to assign timestamps, and a sorting algorithm is applied to organize the data in chronological order, before storing the results in a database.

[0207] Step 7:

[0208] Visualizing Emotion Data

[0209] The server converts the emotional data into a visual format such as a graph or chart, and displays it so that the user can intuitively understand the fluctuations in their emotions throughout the day.

[0210] Input: Organized emotion data

[0211] Data processing: Converting data into graphs, charts, etc.

[0212] Output: Emotion data displayed in a visually understandable format

[0213] Specific operation:

[0214] We create timeline graphs using libraries such as Matplotlib and Plotly, and display the generated graphs on a web dashboard to help users intuitively understand the fluctuations in sentiment throughout the day.

[0215] Step 8:

[0216] Generate improvement suggestions

[0217] The server analyzes the user's emotional tendency based on the emotion analysis result.

[0218] The server generates specific improvement suggestions for the user from the analysis results.

[0219] Input: Sentiment analysis results

[0220] Data calculation: analyzing emotional trends and generating improvement suggestions

[0221] Output: Improvement suggestions

[0222] Specific operation:

[0223] It uses machine learning models to learn patterns and trends from past data and generate suggestions for future improvements, saving the suggestions in a text file in natural language format and sending them to a module that notifies the user.

[0224] (Application example 1)

[0225] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0226] Conventional emotion analysis systems only collect and analyze voice and text data, making it difficult to utilize the resulting emotion information to detect risks or propose countermeasures. In particular, to improve corporate security and safety, it is necessary to quickly grasp changes in users' emotions and propose appropriate countermeasures. To solve this problem, a system is needed that can detect risks from emotion analysis results and propose countermeasures in a timely manner.

[0227] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0228] In this invention, the server includes means for collecting voice data, means for converting the collected voice data into text data, means for collecting message data, means for integrating the voice data and the message data, means for sentiment analysis of the integrated data, means for generating improvement proposals based on the sentiment analysis results, and means for detecting risks from the collected data and proposing appropriate responses. This makes it possible to grasp the emotional state of the user in real time, detect potential risks early, and propose appropriate measures.

[0229] A "means for collecting voice data" is a device or software that records a user's voice and stores it as digital data.

[0230] The "means for converting collected voice data into text data" refers to a device or software that converts voice data into text using voice recognition technology.

[0231] A "means for collecting message data" is a device or software that obtains text data sent and received from a messaging application.

[0232] "Means for integrating voice data and message data" refers to a device or software that unifies data collected in different formats into one format.

[0233] A "means for sentiment analysis of integrated data" is a device or software that analyzes and evaluates sentiment from collected text data.

[0234] The "means for generating improvement suggestions based on the results of sentiment analysis" is a device or software that provides a specific action plan or advice to the user based on the results of sentiment analysis.

[0235] "Means for detecting risks from collected data and proposing appropriate responses" refers to devices or software that analyze emotional data, identify potential risks, and suggest necessary measures.

[0236] The "means for assigning timestamps and organizing the results of sentiment analysis in chronological order" refers to a device or software that assigns time information to each piece of data and arranges it in chronological order.

[0237] "Means for converting the results of sentiment analysis into a visual format such as a graph or chart" refers to a device or software that visually displays the analysis results so that they can be intuitively understood.

[0238] The "means for warning the user of the occurrence of a risk" refers to a device or software that notifies the user when a risk increases based on the analysis results.

[0239] To implement this invention, a system with the following functions is required. First, as a means for collecting voice data, a microphone on a smartphone or smart glasses is used to record the user's voice in real time. As a means for converting collected voice data into text data, voice recognition technology is used. Google Speech-to-Text is a suitable software.

[0240] Additionally, the "means of collecting message data" involves using APIs to obtain text data from messaging applications used within the company (e.g., Slack, Microsoft Teams). The terminals collecting this data must be appropriate devices connected to the server.

[0241] Next, the different formats of data are unified into one format using a "means for integrating voice data and message data." This allows the data to be analyzed using a "means for sentiment analysis of the integrated data" to calculate an emotional score. For this part, a sentiment analysis engine such as Amazon Comprehend is used.

[0242] Furthermore, the "means for generating improvement proposals based on the results of sentiment analysis" proposes specific measures to users based on the obtained sentiment data. This system is capable of detecting risks from the collected data and proposing appropriate responses in a timely manner.

[0243] The "means of assigning timestamps and organizing the sentiment analysis results in chronological order" involves assigning time information to each piece of data and organizing it in chronological order. This allows for an intuitive understanding of sentiment fluctuations. The "means of converting the sentiment analysis results into visual formats such as graphs and charts and alerting users to emerging risks" involves visually displaying them using D3.js and Grafana.

[0244] As a specific example, audio spoken by a user during a meeting is collected using a smartphone microphone and sent to a server. The server then uses Google Speech-to-Text to convert the audio into text data. In parallel, text data sent and received by the user via a messaging application is also sent to the server, and both sets of data are integrated. Sentiment analysis is then performed using Amazon Comprehend, and risks are detected based on the results. For example, specific improvement suggestions are displayed, such as, "A high stress level has been detected from comments made during the meeting. We recommend that you take a break to relax."

[0245] An example prompt is, "Generate an application that uses Google Speech-to-Text to convert employees' real-time voice data into text data, perform sentiment analysis, detect security risks early, and generate improvement suggestions."

[0246] In this way, the entire system works together to grasp the user's emotional state in real time, detect potential risks early, and propose appropriate countermeasures.

[0247] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0248] Step 1:

[0249] Audio data collection

[0250] Users communicate using the microphone on their smartphones or smart glasses. The device collects voice data in real time through the microphone and stores it in a digital file format. The input is the user's voice, and the output is a voice data file.

[0251] Step 2:

[0252] Converting audio data to text

[0253] The device sends the collected voice data to the server, which then uses a voice recognition engine (e.g., Google Speech-to-Text) to convert the voice data into text data. The input is a voice data file, and the output is text data.

[0254] Step 3:

[0255] Message Data Collection

[0256] The device acquires text data from messaging applications used within the company (e.g., Slack, Microsoft Teams), and uses an API to send the message data to the server. The input is the text message from the messaging application, and the output is the collected message data.

[0257] Step 4:

[0258] Data integration

[0259] The server integrates the text data converted from the voice with the message data obtained from the messaging application. Data integration is a process for unifying data of different formats into a single format. The input is the text data from the voice and the message data, and the output is the integrated text data.

[0260] Step 5:

[0261] sentiment analysis

[0262] The server inputs the integrated text data into a sentiment analysis engine (e.g., Amazon Comprehend). The sentiment analysis engine calculates an emotional score (e.g., happy, sad, or angry) from the text data. The input is the integrated text data, and the output is the emotional score.

[0263] Step 6:

[0264] Risk detection and response proposals

[0265] The server runs an algorithm to detect potential risks from the sentiment analysis results. If a risk is detected, the server generates a response suggestion (e.g., "High stress levels detected. We recommend taking a break to relax."). The input is the sentiment score, and the output is the risk detection result and a response suggestion.

[0266] Step 7:

[0267] Organizing data chronologically

[0268] The server assigns a timestamp to each voice and message data and organizes the emotion analysis results in chronological order, making it easier to understand emotion fluctuations throughout the day. The input is the emotion score and the original data, and the output is the time-stamped data.

[0269] Step 8:

[0270] Visualization of sentiment analysis results

[0271] The server converts the sentiment analysis results into visual formats such as graphs and charts and displays them to the user. This process uses visualization tools such as D3.js and Grafana. The input is time-stamped sentiment data, and the output is visualized graphs and charts.

[0272] Step 9:

[0273] Risk warning notification

[0274] If a risk is detected, the server sends a notification to the user to warn them of the risk. This notification can be in the form of a push notification to a smartphone or other device. The input is the risk detection result, and the output is a warning notification to the user.

[0275] By following these steps, users can understand their own emotional state in real time, detect potential risks early, and take appropriate measures.

[0276] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0277] This invention combines a system that collects voice data and message data, integrates the data, performs emotion analysis, and generates improvement proposals, with an emotion engine that recognizes the user's emotions. To realize this system, hardware and software with the following functions are required.

[0278] System configuration

[0279] 1. Device for collecting voice data

[0280] A terminal that uses a meeting application or the like to record user speech.

[0281] 2. Device for collecting message data

[0282] A device that retrieves text messages from a messaging application (e.g., LINE).

[0283] 3. Data analysis server

[0284] A speech recognition engine for converting collected voice data into text data.

[0285] An analysis engine that integrates message data and text data converted from voice to perform sentiment analysis.

[0286] Processing to recognize user emotions using an emotion engine.

[0287] An engine that generates specific improvement suggestions for users based on the results of sentiment analysis.

[0288] System operation procedure

[0289] Data collection

[0290] Audio data collection

[0291] A user speaks during a meeting using a meeting application, for example, "How are we progressing with this week's tasks?"

[0292] The device records the user's speech and generates an audio file.

[0293] Message Data Collection

[0294] A user sends a message on LINE asking, "When is the next meeting?"

[0295] The device receives the LINE text message.

[0296] Data Preprocessing

[0297] Converting voice data to text

[0298] The server receives the audio file and inputs it into a speech recognition engine, which converts the audio data into text data. For example, it generates text data such as "How will you proceed with this week's tasks?"

[0299] Text data integration

[0300] The server combines the text data obtained from LINE with the text data converted from the voice data. The data is then formatted into a unified format and a cleansing process is performed to remove unnecessary symbols and spaces.

[0301] sentiment analysis

[0302] Sentiment analysis of text data

[0303] The server inputs the formatted text data into the emotion engine, which directly recognizes the user's emotions from the text "How will you proceed with your tasks this week?" and "When is the next meeting?" and calculates an emotion score (e.g., joy, sadness, impatience, etc.).

[0304] Data integration

[0305] Timestamp of emotion data

[0306] The server assigns a timestamp to each piece of data and organizes it in chronological order. For example, it records that "How are we progressing with our tasks this week?" was said at 9:00 AM and "When is our next meeting?" was said at 9:10 AM.

[0307] Visualizing Emotion Data

[0308] The server converts the emotional data into a visual format such as graphs and charts, and displays it so that users can intuitively understand the fluctuations in their emotions throughout the day.

[0309] Proposal Generation

[0310] Generate improvement suggestions

[0311] The server analyzes the user's emotional tendencies based on the results of the emotion analysis, and obtains an analysis result such as "people often feel anxious during morning meetings."

[0312] The server generates appropriate improvement suggestions for the user based on the analysis results (e.g., "I suggest you meditate for five minutes before your morning meeting").

[0313] Specific examples

[0314] Example 1:

[0315] A user says during a meeting, "How are we progressing with our tasks this week?"

[0316] The device records what is said and sends the audio file to the server.

[0317] The server converts the audio file into text data such as "How will you proceed with your tasks this week?"

[0318] A user sends a message on LINE asking, "When is the next meeting?"

[0319] The terminal sends the message to the server.

[0320] The server combines both sets of text data and inputs them into the emotion engine.

[0321] The emotion engine recognizes the emotion "impatient" for "How are you progressing with your tasks this week?" and "When is your next meeting?"

[0322] Based on the sentiment score, the server generates a suggestion such as "Try taking some deep breaths to relax before your morning meeting."

[0323] In this way, the system of the present invention operates in cooperation with the server, the terminal, and the user, and in particular, combines the emotion engine to form a specific means for supporting the user's emotional health management.

[0324] The processing flow will be explained below.

[0325] Step 1:

[0326] A user says during a meeting, "What are our goals for next week?"

[0327] Step 2:

[0328] The device records what the user says and generates an audio file that is temporarily stored on the device.

[0329] Step 3:

[0330] The terminal transmits the recorded audio file to the server via the meeting application.

[0331] Step 4:

[0332] The server inputs the received voice file into a speech recognition engine, which converts the voice data into text data and generates the string "What are your goals for next week?"

[0333] Step 5:

[0334] A user sends a message on LINE saying, "What day is the meeting next week?"

[0335] Step 6:

[0336] The device retrieves message data from the LINE application and sends it to the server.

[0337] Step 7:

[0338] The server combines the received LINE message data with the text data converted from the voice. At this time, a cleansing process is performed to unify the data and remove unnecessary symbols and spaces.

[0339] Step 8:

[0340] The server inputs the formatted text data into the emotion engine, which calculates emotion scores for the texts "What are your goals for next week?" and "What day of the week is the meeting next week?"

[0341] Step 9:

[0342] The server recognizes the emotion of each text based on the emotion score obtained from the emotion engine. For example, "What are your goals for next week?" will get an emotion score of "Interest," and "What day of the week is the meeting next week?" will get an emotion score of "Impatience."

[0343] Step 10:

[0344] The server assigns a timestamp to the sentiment score and organizes the data chronologically, for example, recording that "What are your goals for next week?" was said at 10:00 AM and "What day is the meeting next week?" at 10:05 AM.

[0345] Step 11:

[0346] The server converts the time-stamped emotion data into visual formats such as graphs and charts, allowing users to intuitively understand the fluctuations of emotions throughout the day.

[0347] Step 12:

[0348] The server analyzes the user's emotional tendencies based on the emotional data and timestamp data. For example, it determines whether the user frequently feels impatient during certain times of the day.

[0349] Step 13:

[0350] The server generates improvement suggestions for the user based on the analysis results, such as "Try taking deep breaths to relax before your morning meeting."

[0351] Step 14:

[0352] The user checks the suggestions from the server on a device such as a smartphone and adjusts their behavior based on the advice.

[0353] Example 2

[0354] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0355] Conventional systems collect and analyze voice data and message data separately, making it difficult to comprehensively and accurately recognize user emotions. Furthermore, the collected data is often insufficiently cleansed, hindering accurate emotion analysis. Furthermore, it is difficult to provide specific improvement suggestions to users based on the results of emotion analysis, and there is a lack of a way to visually understand daily emotional fluctuations.

[0356] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0357] In this invention, the server includes means for collecting voice data, means for converting the collected voice data into text data, means for collecting message data, means for integrating the voice data and the message data, means for sentiment analysis of the integrated data, means for generating improvement suggestions based on the sentiment analysis results, means for data cleansing to remove unnecessary symbols and spaces from each data, means for generating suggestions to support the user's emotional health management based on the sentiment analysis results, and means for organizing the emotional data in timestamp order. This enables the integrated collection and analysis of voice and message data to accurately recognize the user's emotions. Furthermore, the cleansing process improves data accuracy, and appropriate improvement suggestions can be provided to support the user's emotional health. Organizing the emotional data in timestamp order allows for a visual understanding of emotional fluctuations.

[0358] A "means for collecting voice data" is a device or application that records a user's speech or voice and stores it as a digital audio file.

[0359] The "means for converting collected voice data into text data" refers to a voice recognition engine or software for analyzing voice files and converting their contents into text data in sentence format.

[0360] A "means for collecting message data" is a device or application that captures and stores text messages or chat messages.

[0361] A "means for integrating voice and message data" is software or a process for consolidating data collected from different formats and sources into a single format and managing it in a unified manner.

[0362] "Means for sentiment analysis of the integrated data" refers to an emotion engine or software for analyzing the integrated text data and determining a user's emotion based on the content of the text.

[0363] The "means for generating improvement suggestions based on the results of sentiment analysis" is software or an engine for generating feedback and advice for users based on the results of sentiment analysis.

[0364] A "data cleansing method that removes unnecessary symbols and spaces from each piece of data" is a process or software that automatically removes unnecessary symbols and spaces from collected data to improve the quality of the data.

[0365] The "means for generating suggestions to support the user's emotional health management based on the results of sentiment analysis" is software or an engine for utilizing the results of sentiment analysis to present specific advice and schedules for improving the user's emotional health.

[0366] "Means for organizing emotion data in timestamp order" refers to software or a process for adding time information to collected emotion data and organizing and storing the data in chronological order.

[0367] The present invention is a system that collects voice data and message data, integrates the data, performs sentiment analysis, and generates improvement proposals. This system requires hardware and software with the functions of voice data collection, message data collection, data integration, sentiment analysis, data organization and visualization, and generation of improvement proposals.

[0368] Audio data collection

[0369] Users use a meeting application (e.g., a video conferencing application) to hold a conversation. A device (e.g., a user's PC or smartphone) records the audio during the meeting and generates an audio file. This audio file is automatically uploaded to cloud storage.

[0370] Message Data Collection

[0371] A user sends a message using a messaging application (e.g., a text messaging app), and the device receives the text message and sends it to a server via a dedicated application.

[0372] Converting audio data to text

[0373] The server receives the audio file from the cloud storage. The received audio file is input into a speech recognition engine (e.g., a speech recognition API) and converted into text data. For example, a statement such as "How will you proceed with this week's tasks?" is generated as text data.

[0374] Text data integration

[0375] The server combines the text data obtained from LINE and other messaging apps with the text data converted from voice. During the combination process, a data cleansing process is performed to remove unnecessary symbols and spaces from each data, improving the accuracy of the data.

[0376] sentiment analysis

[0377] The server inputs the cleansed text data into an emotion engine (e.g., a natural language processing API). The emotion engine recognizes the user's emotion for each piece of text and calculates an emotion score. Specifically, for questions like "How will you progress with your tasks this week?" and "When is the next meeting?", emotions such as impatience and anticipation are recognized.

[0378] Timestamp of emotion data

[0379] The server assigns a timestamp to each piece of text data and organizes it in chronological order. For example, it records that "How are we progressing with this week's tasks?" was said at 9:00 AM and "When is the next meeting?" was said at 9:10 AM.

[0380] Visualizing Emotion Data

[0381] The server converts the emotion data into a visual format such as a graph or chart, allowing users to intuitively understand their emotional fluctuations throughout the day. For example, a line graph showing the emotional fluctuations throughout the day can be generated and displayed on the user interface.

[0382] Generate improvement suggestions

[0383] The server analyzes the user's emotional tendencies based on the results of the emotion analysis. It then generates specific improvement suggestions based on the analysis results and notifies the user through the user interface. For example, a suggestion such as "Try taking deep breaths to relax before your morning meeting" may be displayed.

[0384] Specific examples

[0385] A user says during a meeting, "How are we progressing with our tasks this week?"

[0386] The device records what is said and sends the audio file to the server.

[0387] The server converts the audio file into text data such as "How will you proceed with your tasks this week?"

[0388] A user sends a message in a messaging app asking, "When is our next meeting?"

[0389] The terminal sends the message to the server.

[0390] The server combines both sets of text data and inputs them into the emotion engine.

[0391] The emotion engine recognizes "impatience" for "How are you progressing with your tasks this week?" and "When is your next meeting?"

[0392] Based on the sentiment score, the server generates a suggestion such as "Try taking some deep breaths to relax before your morning meeting."

[0393] In this way, the system of the present invention operates in cooperation with the server, the terminal, and the user, and in particular, combines the emotion engine to form a specific means for supporting the user's emotional health management.

[0394] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0395] Step 1:

[0396] Audio data collection

[0397] A user has a conversation using a meeting application (e.g., a video conferencing application). The device records the audio during the meeting and sends it to cloud storage in the form of audio data.

[0398] Input: User's voice

[0399] Output: Audio files saved in cloud storage

[0400] Specifically, when a user says, "How will you proceed with this week's tasks?", the device uploads the voice recording to cloud storage as a digital audio file.

[0401] Step 2:

[0402] Message Data Collection

[0403] A user sends a message using a messaging application (e.g., a text messaging app), and the device sends the text message to a server via a dedicated application.

[0404] Input: User's text message

[0405] Output: Text data stored on the server

[0406] Specifically, when a user sends a message such as "When is the next meeting?", the terminal sends the text message to the server.

[0407] Step 3:

[0408] Converting audio data to text

[0409] The server receives the audio file from the cloud storage and inputs it into a speech recognition engine, which analyzes the audio file and converts it into text data.

[0410] Input: Audio files stored in cloud storage

[0411] Output: Text data converted from audio

[0412] For example, the speech recognition engine generates text data such as "How will you proceed with this week's tasks?"

[0413] Step 4:

[0414] Text data integration

[0415] The server receives and integrates text data obtained from LINE and other messaging apps and text data converted from voice.

[0416] Input: Text data converted from speech and text data from messaging apps

[0417] Output: Integrated text data

[0418] The server performs data cleansing processing, removing unnecessary symbols and spaces, and formatting the data into a unified format, thereby improving the quality of the data.

[0419] Step 5:

[0420] sentiment analysis

[0421] The server inputs the cleansed text data into the emotion engine, which recognizes the user's emotion for each text and calculates an emotion score.

[0422] Input: Cleansed text data

[0423] Output: Sentiment analysis results for each text

[0424] For example, the emotion engine recognizes "impatience" in response to questions such as "How will you progress with your tasks this week?" and "When is the next meeting?"

[0425] Step 6:

[0426] Timestamp of emotion data

[0427] The server assigns a timestamp to each piece of text data and organizes the data in chronological order.

[0428] Input: Sentiment analysis results

[0429] Output: Emotion data with timestamps

[0430] For example, it records that "How will you proceed with this week's tasks?" was said at 9:00 AM, and "When is the next meeting?" was said at 9:10 AM.

[0431] Step 7:

[0432] Visualizing Emotion Data

[0433] The server converts the emotion data into visual formats such as graphs and charts, allowing users to intuitively understand their emotional fluctuations throughout the day.

[0434] Input: Emotion data with timestamps

[0435] Output: Visualized emotional fluctuation data

[0436] Specifically, it generates a line graph showing emotional fluctuations and displays it on the user interface.

[0437] Step 8:

[0438] Generate improvement suggestions

[0439] The server analyzes the user's emotional tendencies based on the results of the emotion analysis, generates specific improvement suggestions based on the analysis results, and notifies the user through the user interface.

[0440] Input: Sentiment analysis results and trends

[0441] Output: Suggested improvements to the user

[0442] As a specific action, the suggestion is to "Try taking some deep breaths to relax before your morning meeting."

[0443] Through the above processing steps, the system can integrate voice data and text messages, perform sentiment analysis, and provide appropriate improvement suggestions to users, thereby effectively supporting their emotional health management.

[0444] (Application example 2)

[0445] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0446] Autonomous vehicles require real-time monitoring of passengers' emotional states and the provision of appropriate support based on those emotions. In particular, a system that can automatically generate and implement specific suggestions and actions to reduce passenger stress and provide a comfortable travel environment is required.

[0447] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0448] In this invention, the server includes means for collecting voice data, means for converting the collected voice data into text data, means for collecting message data, means for integrating the voice data and the message data, means for sentiment analysis of the integrated data, means for generating improvement suggestions based on the sentiment analysis results, and means for collecting passenger voice data and message data, conducting sentiment analysis, and generating stress reduction suggestions. This makes it possible to grasp the passenger's real-time emotional state and generate and provide appropriate stress reduction suggestions according to the situation.

[0449] "Voice data" refers to data in which the user's voice is recorded in digital form.

[0450] "Text data" refers to character information data obtained by converting voice data into text.

[0451] "Message data" refers to data of text messages sent and received by users.

[0452] "Merge" is the process of combining the text data converted from the voice data and the message data into a single data set.

[0453] "Sentiment analysis" is the process of analyzing a user's emotional state (e.g., joy, sadness, impatience, etc.) from text data.

[0454] "Improvement suggestions" are specific actions or advice suggested to the user based on the results of sentiment analysis.

[0455] "Passenger" refers to a person using an autonomous vehicle.

[0456] "Stress reduction suggestions" are suggestions for reducing stress, taking into account the emotional state of passengers.

[0457] The present invention provides a system for collecting and integrating voice and message data, performing sentiment analysis, and generating stress reduction suggestions based on the results, which is particularly applicable to passenger emotion monitoring and stress reduction assistance for autonomous vehicles.

[0458] System configuration

[0459] To realize this system, the following hardware and software are required.

[0460] 1. Terminal

[0461] Microphones placed inside the vehicle: collect passenger voice data.

[0462] In-car infotainment system: Collects passenger text messages through messaging applications.

[0463] 2. Server

[0464] Speech recognition engine: Converts collected voice data into text data.

[0465] Data integration engine: Integrates text data converted from voice data with message data.

[0466] Sentiment Engine: Performs sentiment analysis using the integrated data.

[0467] Suggestion generation engine: Generates stress reduction suggestions based on the sentiment analysis results.

[0468] System operation procedure

[0469] Audio data collection

[0470] When passengers speak in the car, the microphones collect their words and generate an audio file. For example, if they say, "What should we do about our next meeting?", that will be recorded.

[0471] Message Data Collection

[0472] Text data is collected when a passenger sends a message using a messaging app connected to the vehicle's infotainment system. For example, if a passenger sends a message in a messaging app saying, "When is our next meeting?", that message is captured.

[0473] Data Preprocessing

[0474] The server receives the collected voice files and converts them into text data using a speech recognition engine. For example, "What should we do about the next meeting?" is converted into text "What should we do about the next meeting?"

[0475] The server combines the text data obtained from the messaging application with the text data converted from the voice data, and after a data cleansing process, it is formatted into a unified format.

[0476] sentiment analysis

[0477] The server inputs the formatted text data into the emotion engine, which calculates an emotion score (e.g., impatience, tension, etc.) based on passenger comments such as "What should we do about the next meeting?" and "When is the next meeting?"

[0478] Proposal Generation

[0479] The server generates stress reduction suggestions based on the emotion analysis results. For example, if the emotion engine recognizes "impatience," it generates a suggestion such as "Would you like to play some relaxing music?"

[0480] Suggestions are communicated to passengers through the vehicle's infotainment system, which will either play music automatically or allow passengers to select the suggestion.

[0481] Specific examples

[0482] If a passenger says, "What should we do about our next meeting?", the server converts the speech to text and analyzes that text with an emotion engine. If the emotion engine recognizes the emotion "tension," it generates a suggestion: "Take a deep breath and relax."

[0483] Similarly, if a passenger sends a message in a messaging app asking, "When is my next meeting?", sentiment analysis will be performed and stress-reducing suggestions will be generated.

[0484] Prompt Sentence Examples

[0485] The following are examples of prompts to input into the generative AI model:

[0486] Collect passenger comments, analyze their emotions, and generate stress-reducing suggestions. Below is an example of a passenger comment and the results of the emotion analysis.

[0487] Say: "What about our next meeting?"

[0488] Sentiment Analysis: Negative

[0489] In these cases, generate an appropriate suggestion, for example, "Would you like to play some relaxing music?"

[0490] This system can monitor the emotional state of passengers in autonomous vehicles in real time and provide appropriate stress reduction suggestions, allowing passengers to enjoy a more comfortable and safer travel environment.

[0491] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0492] Step 1:

[0493] Audio data collection

[0494] Input: User's spoken utterance.

[0495] The device uses the in-car microphone to collect the user's speech in real time and record it as an audio file.

[0496] Output: Collected audio data (audio files).

[0497] Step 2:

[0498] Message Data Collection

[0499] Input: A text message that a user sends using a messaging application.

[0500] The device works with the car's infotainment system to retrieve text messages from messaging applications.

[0501] Output: Collected message data (text messages).

[0502] Step 3:

[0503] Converting voice data to text

[0504] Input: Audio data (audio file).

[0505] The server uses a speech recognition engine to convert the voice data into text data.

[0506] Specifically, a speech recognition engine analyzes the collected audio files and generates corresponding text.

[0507] Output: Text data converted from audio.

[0508] Step 4:

[0509] Data integration

[0510] Input: Text data converted from audio and message data.

[0511] The server integrates the text data converted from the voice with the message data and performs a cleansing process.

[0512] The cleansing process is a process of arranging data into a unified format and removing unnecessary symbols and spaces.

[0513] Output: Consolidated clean text data.

[0514] Step 5:

[0515] sentiment analysis

[0516] Input: Integrated text data.

[0517] The server uses an emotion engine to analyze the integrated text data and calculate an emotion score.

[0518] The emotion engine identifies an emotional state, such as "impatience" or "tension," based on the text data.

[0519] Output: Emotion score and emotional state identification results.

[0520] Step 6:

[0521] Generate improvement suggestions

[0522] Input: Emotion scores and emotional state identification results.

[0523] The server generates stress reduction suggestions based on the results of the sentiment analysis.

[0524] A specific suggestion generation engine generates specific actions such as "play relaxing music" or "encourage deep breathing."

[0525] Output: Specific stress reduction suggestions.

[0526] Step 7:

[0527] Proposal Notification

[0528] Enter: stress reduction suggestions.

[0529] The device will then notify passengers of stress-reducing suggestions through the vehicle's infotainment system.

[0530] The passenger can select a suggestion or an automatic action will be initiated to implement the suggestion.

[0531] Output: Passenger notification and action taken.

[0532] Through these processing steps, the system can monitor the user's emotional state in real time and provide appropriate stress reduction suggestions, allowing passengers to enjoy a more comfortable travel environment.

[0533] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0534] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0535] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0536] [Second embodiment]

[0537] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0538] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0539] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0540] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0541] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0542] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0543] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0544] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0545] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0546] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0547] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0548] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0549] The present invention is a system that collects voice data and text data, integrates the data, performs sentiment analysis, and generates improvement suggestions for users based on emotional fluctuations. To realize this system, hardware and software with the following functions are required.

[0550] System configuration

[0551] 1. Device for collecting voice data

[0552] The user's voice is recorded by using a meeting application or the like.

[0553] 2. Device for collecting message data

[0554] Get text messages sent and received from LINE and other messaging applications.

[0555] 3. Data analysis server

[0556] A speech recognition engine for converting collected voice data into text data.

[0557] An analysis engine that integrates message data and text data converted from voice to perform sentiment analysis.

[0558] An engine for generating specific improvement suggestions for users based on the results of sentiment analysis.

[0559] System operation procedure

[0560] Data collection

[0561] Audio data collection

[0562] A user speaks during a conference using a meeting application.

[0563] The device records what is said, generates an audio file, and sends it to the server.

[0564] Message Data Collection

[0565] A user sends a message via LINE.

[0566] The device receives the LINE text message and sends it to the server.

[0567] Data Preprocessing

[0568] Converting voice data to text

[0569] The server inputs the received voice file into a voice recognition engine, which converts the voice into text data.

[0570] Text data integration

[0571] The server integrates the converted voice data with the message data obtained from LINE.

[0572] The integrated data is cleansed to remove unnecessary symbols and spaces.

[0573] sentiment analysis

[0574] Sentiment analysis of text data

[0575] The server inputs the cleansed text data into a sentiment analysis engine.

[0576] A sentiment analysis engine calculates an emotion score (e.g., happy, sad, anger) for each text block.

[0577] Data integration

[0578] Timestamp of emotion data

[0579] The server assigns a timestamp to each piece of data and organizes it in chronological order.

[0580] Visualizing Emotion Data

[0581] The server converts the emotional data into a visual format such as graphs and charts, and displays it so that users can intuitively understand the fluctuations in their emotions throughout the day.

[0582] Proposal Generation

[0583] Generate improvement suggestions

[0584] The server analyzes the user's emotional tendencies based on the emotion analysis results.

[0585] The server generates appropriate improvement suggestions for the user based on the analysis results (e.g., "Try taking deep relaxation breaths before a meeting").

[0586] Specific examples

[0587] Example 1:

[0588] A user says during a meeting, "What are your tasks for today?"

[0589] The device records what is said and sends the audio file to the server.

[0590] The server converts the audio file into text data: "What is your task today?"

[0591] A user sends a message on LINE asking, "How did the afternoon meeting go?"

[0592] The terminal sends the message to the server.

[0593] The server combines both sets of text data and feeds it into a sentiment analysis engine.

[0594] The emotion analysis detects "impatience" and generates a suggestion based on that, such as "We recommend you relax before your afternoon meeting."

[0595] In this way, the system of the present invention mainly operates in cooperation with the server, the terminal, and the user, and constitutes a specific means for supporting the user's emotional health management.

[0596] The processing flow will be explained below.

[0597] Step 1:

[0598] A user speaks up during a meeting, for example, "What are our goals for next week?"

[0599] Step 2:

[0600] The device records what the user says and generates an audio file that is temporarily stored on the device.

[0601] Step 3:

[0602] The terminal transmits the recorded audio file to the server via the meeting application.

[0603] Step 4:

[0604] The server inputs the received voice file into a voice recognition engine, which converts the voice data into text data, generating a string of characters such as "What are your goals for next week?"

[0605] Step 5:

[0606] A user sends a message on LINE saying, "What day is the meeting next week?"

[0607] Step 6:

[0608] The device retrieves message data from the LINE application and sends it to the server.

[0609] Step 7:

[0610] The server combines the received LINE message data with the text data converted from the voice. At this time, a cleansing process is performed to unify the data and remove unnecessary symbols and spaces.

[0611] Step 8:

[0612] The server inputs the formatted text data into a sentiment analysis engine, which calculates sentiment scores for the texts "What are your goals for next week?" and "What day of the week is the meeting next week?"

[0613] Step 9:

[0614] The server assigns a timestamp to each block of text based on the sentiment score obtained from the sentiment analysis engine, for example, recording that "What are your goals for next week?" was spoken at 10:00 AM and "What day is our meeting next week?" was spoken at 10:05 AM.

[0615] Step 10:

[0616] The server chronologically organizes the time-stamped emotion data and summarizes the emotional fluctuations by daily activity. The emotion data is then converted into a visual format such as a graph or chart.

[0617] Step 11:

[0618] The server analyzes the user's emotional tendencies based on the emotional data and timestamp data, and obtains analysis results such as "people often feel anxious during morning meetings."

[0619] Step 12:

[0620] Based on the analysis results, the server generates improvement suggestions for the user, such as "I suggest you meditate for five minutes before your morning meeting."

[0621] Step 13:

[0622] The user checks the suggestions from the server on a device such as a smartphone and adjusts their behavior based on the advice.

[0623] Example 1

[0624] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0625] In today's communication environment, there is a lack of methods to properly understand users' emotional fluctuations and stress levels and support their emotional health management. In particular, there is a need for a method that can more accurately assess users' emotional state and provide effective improvement suggestions by integrating and analyzing voice data and text messages.

[0626] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0627] In this invention, the server includes means for collecting voice data, means for converting the collected voice data into text data, means for collecting text messages, means for integrating the voice data and the text messages, means for preprocessing the integrated data, means for sentiment analysis of the preprocessed data, and means for generating improvement suggestions based on the sentiment analysis results. This makes it possible to comprehensively analyze a user's various communication data, appropriately evaluate the emotional state, and provide effective improvement suggestions.

[0628] "Audio data" refers to digital audio files that record the content of user speech occurring during meetings, conversations, and the like.

[0629] "Text data" refers to character information converted from voice data or character information sent and received by a message application.

[0630] A "text message" is text information that a user sends and receives using a text application (e.g., a messaging app).

[0631] "Preprocessing" is the procedure of removing unnecessary information and extracting necessary parts in order to prepare data in a format that can be analyzed.

[0632] "Sentiment analysis" is the process of analyzing a user's emotional state (e.g., joy, sadness, anger) from text data and classifying it into scores or categories.

[0633] "Improvement suggestions" are specific advice on improving behavior or status that is provided to the user based on the results of sentiment analysis.

[0634] "Fusion" refers to combining multiple data sources (e.g., voice data and text messages) into a single dataset.

[0635] A "timestamp" is a means of clarifying time-series information by adding information about the date and time when data was generated or collected.

[0636] "Visualization" is the process of transforming data into a visually understandable format such as a graph or chart.

[0637] The present invention provides a system for collecting voice data and text messages, integrating the collected data to perform sentiment analysis, and generating improvement suggestions for users based on emotional fluctuations. Specific embodiments are described below.

[0638] Hardware and Software Configuration

[0639] 1. Device for collecting audio data:

[0640] A user speaks during a meeting using a meeting application (e.g., an online conference system).

[0641] The terminal records the user's speech during the conference, generates an audio file (e.g., a WAV file), and sends it to the server.

[0642] 2. Devices for Text Message Data Collection:

[0643] A user has a conversation using a messaging application (e.g., a messaging app).

[0644] The terminal retrieves the text message from the message application and sends it to the server as a text file.

[0645] 3. Data analysis server:

[0646] The server uses speech recognition and sentiment analysis engines such as Google Cloud Speech-to-Text and IBM Watson Natural Language Understanding to convert the voice data into text data and perform further sentiment analysis.

[0647] The server includes an engine for generating specific improvement suggestions for the user based on the sentiment analysis results.

[0648] Specific examples

[0649] Example 1:

[0650] A user says, "What are your tasks for today?" during an online meeting.

[0651] The device records what is said and sends the audio file to the server.

[0652] The server inputs the audio file into the Google Cloud Speech-to-Text engine and converts it into text data: "What is your task today?"

[0653] A user sends a message in a messaging app asking, "How did the meeting this afternoon go?"

[0654] The terminal receives the message and sends it to the server.

[0655] The server integrates both sets of text data and performs a cleansing process.

[0656] The server inputs the cleansed text data into IBM Watson Natural Language Understanding for sentiment analysis.

[0657] As a result of the emotion analysis, "impatience" is detected, and based on that, an improvement suggestion is generated, such as "recommending relaxation before the afternoon meeting."

[0658] Examples of prompt statements

[0659] Example prompt sentence:

[0660] "A user might say, 'What are your tasks for today?' in an online meeting, and then later send a message in a messaging app asking, 'How did your afternoon meeting go?' Combine these data points, perform sentiment analysis, and generate appropriate improvement suggestions."

[0661] In this way, the system of the present invention constitutes a concrete means for supporting the emotional health management of users, with the server, terminal, and user working together.

[0662] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0663] Step 1:

[0664] Collection and transmission of voice data

[0665] A user speaks during a conference using an online conference system.

[0666] The terminal records the user's speech and generates an audio file (e.g., WAV format).

[0667] The terminal transmits the generated audio file to the server.

[0668] Input: User's speech in an online conference system

[0669] Data processing: Recording audio data and generating WAV format audio files

[0670] Output: Audio file sent to the server

[0671] Specific operation:

[0672] Execute a script to record speech on the online conference system, generate an audio file, and upload the generated audio file to the server using FTP or HTTP protocol.

[0673] Step 2:

[0674] Collection and transmission of message data

[0675] A user sends a text message using a messaging application.

[0676] The terminal acquires text messages sent and received from the message application and saves them as text files.

[0677] The terminal transmits the obtained text file to the server.

[0678] Input: The user's text message in the Messages application

[0679] Data processing: Obtaining text messages and generating text files in TXT format

[0680] Output: A text file that is sent to the server.

[0681] Specific operation:

[0682] When a new message is detected, a script running in the background of the Messages app extracts its contents, saves the extracted text message in a text file, and uploads it to a server.

[0683] Step 3:

[0684] Speech-to-text

[0685] The server inputs the received audio file into a speech recognition engine (e.g., Google Cloud Speech-to-Text).

[0686] The server converts the audio file into text data using a speech recognition engine.

[0687] The server stores the converted text data in temporary storage.

[0688] Input: Audio file sent to the server

[0689] Data processing: Converting audio files into text using a speech recognition engine

[0690] Output: Text data saved in temporary storage

[0691] Specific operation:

[0692] A Python script is executed on the server to send the audio file to the Google Cloud Speech-to-Text API, and the text data returned by the API is retrieved in JSON format and saved in a database on the server.

[0693] Step 4:

[0694] Text data integration and preprocessing

[0695] The server integrates the text data converted from the voice data with the text file obtained from the messaging app.

[0696] The server cleanses the integrated text data, removing unnecessary symbols and spaces.

[0697] Input: Text data converted from voice, text data obtained from messaging apps

[0698] Data processing: text data integration and cleansing

[0699] Output: Integrated text data after cleansing

[0700] Specific operation:

[0701] The integration process runs an SQL query to combine the voice text and message text, and uses regular expressions (RegEx) on the combined data to remove unnecessary symbols and spaces.

[0702] Step 5:

[0703] Conducting sentiment analysis

[0704] The server inputs the cleansed text data into a sentiment analysis engine (e.g., IBM Watson Natural Language Understanding).

[0705] The server obtains the sentiment score (e.g., happy, sad, anger) for each text block returned by the sentiment analysis engine.

[0706] Input: Text data after cleansing

[0707] Data calculation: Calculating sentiment scores using a sentiment analysis engine

[0708] Output: Text data with sentiment scores

[0709] Specific operation:

[0710] The server makes an API request to send the cleansed text data to the sentiment analysis engine, which analyzes the resulting sentiment scores and stores them in a database.

[0711] Step 6:

[0712] Timestamp and organize emotion data

[0713] The server assigns a timestamp to each piece of emotion data and organizes it in chronological order.

[0714] Input: Text data with sentiment scores

[0715] Data processing: Adding timestamps and organizing data in chronological order

[0716] Output: Emotion data with timestamps and organized in chronological order

[0717] Specific operation:

[0718] The Python Pandas library is used to assign timestamps, and a sorting algorithm is applied to organize the data in chronological order, before storing the results in a database.

[0719] Step 7:

[0720] Visualizing Emotion Data

[0721] The server converts the emotional data into a visual format such as a graph or chart, and displays it so that the user can intuitively understand the fluctuations in their emotions throughout the day.

[0722] Input: Organized emotion data

[0723] Data processing: Converting data into graphs, charts, etc.

[0724] Output: Emotion data displayed in a visually understandable format

[0725] Specific operation:

[0726] We create timeline graphs using libraries such as Matplotlib and Plotly, and display the generated graphs on a web dashboard to help users intuitively understand the fluctuations in sentiment throughout the day.

[0727] Step 8:

[0728] Generate improvement suggestions

[0729] The server analyzes the user's emotional tendency based on the emotion analysis result.

[0730] The server generates specific improvement suggestions for the user from the analysis results.

[0731] Input: Sentiment analysis results

[0732] Data calculation: analyzing emotional trends and generating improvement suggestions

[0733] Output: Improvement suggestions

[0734] Specific operation:

[0735] It uses machine learning models to learn patterns and trends from past data and generate suggestions for future improvements, saving the suggestions in a text file in natural language format and sending them to a module that notifies the user.

[0736] (Application example 1)

[0737] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0738] Conventional emotion analysis systems only collect and analyze voice and text data, making it difficult to utilize the resulting emotion information to detect risks or propose countermeasures. In particular, to improve corporate security and safety, it is necessary to quickly grasp changes in users' emotions and propose appropriate countermeasures. To solve this problem, a system is needed that can detect risks from emotion analysis results and propose countermeasures in a timely manner.

[0739] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0740] In this invention, the server includes means for collecting voice data, means for converting the collected voice data into text data, means for collecting message data, means for integrating the voice data and the message data, means for sentiment analysis of the integrated data, means for generating improvement proposals based on the sentiment analysis results, and means for detecting risks from the collected data and proposing appropriate responses. This makes it possible to grasp the emotional state of the user in real time, detect potential risks early, and propose appropriate measures.

[0741] A "means for collecting voice data" is a device or software that records a user's voice and stores it as digital data.

[0742] The "means for converting collected voice data into text data" refers to a device or software that converts voice data into text using voice recognition technology.

[0743] A "means for collecting message data" is a device or software that obtains text data sent and received from a messaging application.

[0744] "Means for integrating voice data and message data" refers to a device or software that unifies data collected in different formats into one format.

[0745] A "means for sentiment analysis of integrated data" is a device or software that analyzes and evaluates sentiment from collected text data.

[0746] The "means for generating improvement suggestions based on the results of sentiment analysis" is a device or software that provides a specific action plan or advice to the user based on the results of sentiment analysis.

[0747] "Means for detecting risks from collected data and proposing appropriate responses" refers to devices or software that analyze emotional data, identify potential risks, and suggest necessary measures.

[0748] The "means for assigning timestamps and organizing the results of sentiment analysis in chronological order" refers to a device or software that assigns time information to each piece of data and arranges it in chronological order.

[0749] "Means for converting the results of sentiment analysis into a visual format such as a graph or chart" refers to a device or software that visually displays the analysis results so that they can be intuitively understood.

[0750] The "means for warning the user of the occurrence of a risk" refers to a device or software that notifies the user when a risk increases based on the analysis results.

[0751] To implement this invention, a system with the following functions is required. First, as a means for collecting voice data, a microphone on a smartphone or smart glasses is used to record the user's voice in real time. As a means for converting collected voice data into text data, voice recognition technology is used. Google Speech-to-Text is a suitable software.

[0752] Additionally, the "means of collecting message data" involves using APIs to obtain text data from messaging applications used within the company (e.g., Slack, Microsoft Teams). The terminals collecting this data must be appropriate devices connected to the server.

[0753] Next, the different formats of data are unified into one format using a "means for integrating voice data and message data." This allows the data to be analyzed using a "means for sentiment analysis of the integrated data" to calculate an emotional score. For this part, a sentiment analysis engine such as Amazon Comprehend is used.

[0754] Furthermore, the "means for generating improvement proposals based on the results of sentiment analysis" proposes specific measures to users based on the obtained sentiment data. This system is capable of detecting risks from the collected data and proposing appropriate responses in a timely manner.

[0755] The "means of assigning timestamps and organizing the sentiment analysis results in chronological order" involves assigning time information to each piece of data and organizing it in chronological order. This allows for an intuitive understanding of sentiment fluctuations. The "means of converting the sentiment analysis results into visual formats such as graphs and charts and alerting users to emerging risks" involves visually displaying them using D3.js and Grafana.

[0756] As a specific example, audio spoken by a user during a meeting is collected using a smartphone microphone and sent to a server. The server then uses Google Speech-to-Text to convert the audio into text data. In parallel, text data sent and received by the user via a messaging application is also sent to the server, and both sets of data are integrated. Sentiment analysis is then performed using Amazon Comprehend, and risks are detected based on the results. For example, specific improvement suggestions are displayed, such as, "A high stress level has been detected from comments made during the meeting. We recommend that you take a break to relax."

[0757] An example prompt is, "Generate an application that uses Google Speech-to-Text to convert employees' real-time voice data into text data, perform sentiment analysis, detect security risks early, and generate improvement suggestions."

[0758] In this way, the entire system works together to grasp the user's emotional state in real time, detect potential risks early, and propose appropriate countermeasures.

[0759] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0760] Step 1:

[0761] Audio data collection

[0762] Users communicate using the microphone on their smartphones or smart glasses. The device collects voice data in real time through the microphone and stores it in a digital file format. The input is the user's voice, and the output is a voice data file.

[0763] Step 2:

[0764] Converting audio data to text

[0765] The device sends the collected voice data to the server, which then uses a voice recognition engine (e.g., Google Speech-to-Text) to convert the voice data into text data. The input is a voice data file, and the output is text data.

[0766] Step 3:

[0767] Message Data Collection

[0768] The device acquires text data from messaging applications used within the company (e.g., Slack, Microsoft Teams), and uses an API to send the message data to the server. The input is the text message from the messaging application, and the output is the collected message data.

[0769] Step 4:

[0770] Data integration

[0771] The server integrates the text data converted from the voice with the message data obtained from the messaging application. Data integration is a process for unifying data of different formats into a single format. The input is the text data from the voice and the message data, and the output is the integrated text data.

[0772] Step 5:

[0773] sentiment analysis

[0774] The server inputs the integrated text data into a sentiment analysis engine (e.g., Amazon Comprehend). The sentiment analysis engine calculates an emotional score (e.g., happy, sad, or angry) from the text data. The input is the integrated text data, and the output is the emotional score.

[0775] Step 6:

[0776] Risk detection and response proposals

[0777] The server runs an algorithm to detect potential risks from the sentiment analysis results. If a risk is detected, the server generates a response suggestion (e.g., "High stress levels detected. We recommend taking a break to relax."). The input is the sentiment score, and the output is the risk detection result and a response suggestion.

[0778] Step 7:

[0779] Organizing data chronologically

[0780] The server assigns a timestamp to each voice and message data and organizes the emotion analysis results in chronological order, making it easier to understand emotion fluctuations throughout the day. The input is the emotion score and the original data, and the output is the time-stamped data.

[0781] Step 8:

[0782] Visualization of sentiment analysis results

[0783] The server converts the sentiment analysis results into visual formats such as graphs and charts and displays them to the user. This process uses visualization tools such as D3.js and Grafana. The input is time-stamped sentiment data, and the output is visualized graphs and charts.

[0784] Step 9:

[0785] Risk warning notification

[0786] If a risk is detected, the server sends a notification to the user to warn them of the risk. This notification can be in the form of a push notification to a smartphone or other device. The input is the risk detection result, and the output is a warning notification to the user.

[0787] By following these steps, users can understand their own emotional state in real time, detect potential risks early, and take appropriate measures.

[0788] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0789] This invention combines a system that collects voice data and message data, integrates the data, performs emotion analysis, and generates improvement proposals, with an emotion engine that recognizes the user's emotions. To realize this system, hardware and software with the following functions are required.

[0790] System configuration

[0791] 1. Device for collecting voice data

[0792] A terminal that uses a meeting application or the like to record user speech.

[0793] 2. Device for collecting message data

[0794] A device that retrieves text messages from a messaging application (e.g., LINE).

[0795] 3. Data analysis server

[0796] A speech recognition engine for converting collected voice data into text data.

[0797] An analysis engine that integrates message data and text data converted from voice to perform sentiment analysis.

[0798] Processing to recognize user emotions using an emotion engine.

[0799] An engine that generates specific improvement suggestions for users based on the results of sentiment analysis.

[0800] System operation procedure

[0801] Data collection

[0802] Audio data collection

[0803] A user speaks during a meeting using a meeting application, for example, "How are we progressing with this week's tasks?"

[0804] The device records the user's speech and generates an audio file.

[0805] Message Data Collection

[0806] A user sends a message on LINE asking, "When is the next meeting?"

[0807] The device receives the LINE text message.

[0808] Data Preprocessing

[0809] Converting voice data to text

[0810] The server receives the audio file and inputs it into a speech recognition engine, which converts the audio data into text data. For example, it generates text data such as "How will you proceed with this week's tasks?"

[0811] Text data integration

[0812] The server combines the text data obtained from LINE with the text data converted from the voice data. The data is then formatted into a unified format and a cleansing process is performed to remove unnecessary symbols and spaces.

[0813] sentiment analysis

[0814] Sentiment analysis of text data

[0815] The server inputs the formatted text data into the emotion engine, which directly recognizes the user's emotions from the text "How will you proceed with your tasks this week?" and "When is the next meeting?" and calculates an emotion score (e.g., joy, sadness, impatience, etc.).

[0816] Data integration

[0817] Timestamp of emotion data

[0818] The server assigns a timestamp to each piece of data and organizes it in chronological order. For example, it records that "How are we progressing with our tasks this week?" was said at 9:00 AM and "When is our next meeting?" was said at 9:10 AM.

[0819] Visualizing Emotion Data

[0820] The server converts the emotional data into a visual format such as graphs and charts, and displays it so that users can intuitively understand the fluctuations in their emotions throughout the day.

[0821] Proposal Generation

[0822] Generate improvement suggestions

[0823] The server analyzes the user's emotional tendencies based on the results of the emotion analysis, and obtains an analysis result such as "people often feel anxious during morning meetings."

[0824] The server generates appropriate improvement suggestions for the user based on the analysis results (e.g., "I suggest you meditate for five minutes before your morning meeting").

[0825] Specific examples

[0826] Example 1:

[0827] A user says during a meeting, "How are we progressing with our tasks this week?"

[0828] The device records what is said and sends the audio file to the server.

[0829] The server converts the audio file into text data such as "How will you proceed with your tasks this week?"

[0830] A user sends a message on LINE asking, "When is the next meeting?"

[0831] The terminal sends the message to the server.

[0832] The server combines both sets of text data and inputs them into the emotion engine.

[0833] The emotion engine recognizes the emotion "impatient" for "How are you progressing with your tasks this week?" and "When is your next meeting?"

[0834] Based on the sentiment score, the server generates a suggestion such as "Try taking some deep breaths to relax before your morning meeting."

[0835] In this way, the system of the present invention operates in cooperation with the server, the terminal, and the user, and in particular, combines the emotion engine to form a specific means for supporting the user's emotional health management.

[0836] The processing flow will be explained below.

[0837] Step 1:

[0838] A user says during a meeting, "What are our goals for next week?"

[0839] Step 2:

[0840] The device records what the user says and generates an audio file that is temporarily stored on the device.

[0841] Step 3:

[0842] The terminal transmits the recorded audio file to the server via the meeting application.

[0843] Step 4:

[0844] The server inputs the received voice file into a speech recognition engine, which converts the voice data into text data and generates the string "What are your goals for next week?"

[0845] Step 5:

[0846] A user sends a message on LINE saying, "What day is the meeting next week?"

[0847] Step 6:

[0848] The device retrieves message data from the LINE application and sends it to the server.

[0849] Step 7:

[0850] The server combines the received LINE message data with the text data converted from the voice. At this time, a cleansing process is performed to unify the data and remove unnecessary symbols and spaces.

[0851] Step 8:

[0852] The server inputs the formatted text data into the emotion engine, which calculates emotion scores for the texts "What are your goals for next week?" and "What day of the week is the meeting next week?"

[0853] Step 9:

[0854] The server recognizes the emotion of each text based on the emotion score obtained from the emotion engine. For example, "What are your goals for next week?" will get an emotion score of "Interest," and "What day of the week is the meeting next week?" will get an emotion score of "Impatience."

[0855] Step 10:

[0856] The server assigns a timestamp to the sentiment score and organizes the data chronologically, for example, recording that "What are your goals for next week?" was said at 10:00 AM and "What day is the meeting next week?" at 10:05 AM.

[0857] Step 11:

[0858] The server converts the time-stamped emotion data into visual formats such as graphs and charts, allowing users to intuitively understand the fluctuations of emotions throughout the day.

[0859] Step 12:

[0860] The server analyzes the user's emotional tendencies based on the emotional data and timestamp data. For example, it determines whether the user frequently feels impatient during certain times of the day.

[0861] Step 13:

[0862] The server generates improvement suggestions for the user based on the analysis results, such as "Try taking deep breaths to relax before your morning meeting."

[0863] Step 14:

[0864] The user checks the suggestions from the server on a device such as a smartphone and adjusts their behavior based on the advice.

[0865] Example 2

[0866] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0867] Conventional systems collect and analyze voice data and message data separately, making it difficult to comprehensively and accurately recognize user emotions. Furthermore, the collected data is often insufficiently cleansed, hindering accurate emotion analysis. Furthermore, it is difficult to provide specific improvement suggestions to users based on the results of emotion analysis, and there is a lack of a way to visually understand daily emotional fluctuations.

[0868] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0869] In this invention, the server includes means for collecting voice data, means for converting the collected voice data into text data, means for collecting message data, means for integrating the voice data and the message data, means for sentiment analysis of the integrated data, means for generating improvement suggestions based on the sentiment analysis results, means for data cleansing to remove unnecessary symbols and spaces from each data, means for generating suggestions to support the user's emotional health management based on the sentiment analysis results, and means for organizing the emotional data in timestamp order. This enables the integrated collection and analysis of voice and message data to accurately recognize the user's emotions. Furthermore, the cleansing process improves data accuracy, and appropriate improvement suggestions can be provided to support the user's emotional health. Organizing the emotional data in timestamp order allows for a visual understanding of emotional fluctuations.

[0870] A "means for collecting voice data" is a device or application that records a user's speech or voice and stores it as a digital audio file.

[0871] The "means for converting collected voice data into text data" refers to a voice recognition engine or software for analyzing voice files and converting their contents into text data in sentence format.

[0872] A "means for collecting message data" is a device or application that captures and stores text messages or chat messages.

[0873] A "means for integrating voice and message data" is software or a process for consolidating data collected from different formats and sources into a single format and managing it in a unified manner.

[0874] "Means for sentiment analysis of the integrated data" refers to an emotion engine or software for analyzing the integrated text data and determining a user's emotion based on the content of the text.

[0875] The "means for generating improvement suggestions based on the results of sentiment analysis" is software or an engine for generating feedback and advice for users based on the results of sentiment analysis.

[0876] A "data cleansing method that removes unnecessary symbols and spaces from each piece of data" is a process or software that automatically removes unnecessary symbols and spaces from collected data to improve the quality of the data.

[0877] The "means for generating suggestions to support the user's emotional health management based on the results of sentiment analysis" is software or an engine for utilizing the results of sentiment analysis to present specific advice and schedules for improving the user's emotional health.

[0878] "Means for organizing emotion data in timestamp order" refers to software or a process for adding time information to collected emotion data and organizing and storing the data in chronological order.

[0879] The present invention is a system that collects voice data and message data, integrates the data, performs sentiment analysis, and generates improvement proposals. This system requires hardware and software with the functions of voice data collection, message data collection, data integration, sentiment analysis, data organization and visualization, and generation of improvement proposals.

[0880] Audio data collection

[0881] Users use a meeting application (e.g., a video conferencing application) to hold a conversation. A device (e.g., a user's PC or smartphone) records the audio during the meeting and generates an audio file. This audio file is automatically uploaded to cloud storage.

[0882] Message Data Collection

[0883] A user sends a message using a messaging application (e.g., a text messaging app), and the device receives the text message and sends it to a server via a dedicated application.

[0884] Converting audio data to text

[0885] The server receives the audio file from the cloud storage. The received audio file is input into a speech recognition engine (e.g., a speech recognition API) and converted into text data. For example, a statement such as "How will you proceed with this week's tasks?" is generated as text data.

[0886] Text data integration

[0887] The server combines the text data obtained from LINE and other messaging apps with the text data converted from voice. During the combination process, a data cleansing process is performed to remove unnecessary symbols and spaces from each data, improving the accuracy of the data.

[0888] sentiment analysis

[0889] The server inputs the cleansed text data into an emotion engine (e.g., a natural language processing API). The emotion engine recognizes the user's emotion for each piece of text and calculates an emotion score. Specifically, for questions like "How will you progress with your tasks this week?" and "When is the next meeting?", emotions such as impatience and anticipation are recognized.

[0890] Timestamp of emotion data

[0891] The server assigns a timestamp to each piece of text data and organizes it in chronological order. For example, it records that "How are we progressing with this week's tasks?" was said at 9:00 AM and "When is the next meeting?" was said at 9:10 AM.

[0892] Visualizing Emotion Data

[0893] The server converts the emotion data into a visual format such as a graph or chart, allowing users to intuitively understand their emotional fluctuations throughout the day. For example, a line graph showing the emotional fluctuations throughout the day can be generated and displayed on the user interface.

[0894] Generate improvement suggestions

[0895] The server analyzes the user's emotional tendencies based on the results of the emotion analysis. It then generates specific improvement suggestions based on the analysis results and notifies the user through the user interface. For example, a suggestion such as "Try taking deep breaths to relax before your morning meeting" may be displayed.

[0896] Specific examples

[0897] A user says during a meeting, "How are we progressing with our tasks this week?"

[0898] The device records what is said and sends the audio file to the server.

[0899] The server converts the audio file into text data such as "How will you proceed with your tasks this week?"

[0900] A user sends a message in a messaging app asking, "When is our next meeting?"

[0901] The terminal sends the message to the server.

[0902] The server combines both sets of text data and inputs them into the emotion engine.

[0903] The emotion engine recognizes "impatience" for "How are you progressing with your tasks this week?" and "When is your next meeting?"

[0904] Based on the sentiment score, the server generates a suggestion such as "Try taking some deep breaths to relax before your morning meeting."

[0905] In this way, the system of the present invention operates in cooperation with the server, the terminal, and the user, and in particular, combines the emotion engine to form a specific means for supporting the user's emotional health management.

[0906] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0907] Step 1:

[0908] Audio data collection

[0909] A user has a conversation using a meeting application (e.g., a video conferencing application). The device records the audio during the meeting and sends it to cloud storage in the form of audio data.

[0910] Input: User's voice

[0911] Output: Audio files saved in cloud storage

[0912] Specifically, when a user says, "How will you proceed with this week's tasks?", the device uploads the voice recording to cloud storage as a digital audio file.

[0913] Step 2:

[0914] Message Data Collection

[0915] A user sends a message using a messaging application (e.g., a text messaging app), and the device sends the text message to a server via a dedicated application.

[0916] Input: User's text message

[0917] Output: Text data stored on the server

[0918] Specifically, when a user sends a message such as "When is the next meeting?", the terminal sends the text message to the server.

[0919] Step 3:

[0920] Converting audio data to text

[0921] The server receives the audio file from the cloud storage and inputs it into a speech recognition engine, which analyzes the audio file and converts it into text data.

[0922] Input: Audio files stored in cloud storage

[0923] Output: Text data converted from audio

[0924] For example, the speech recognition engine generates text data such as "How will you proceed with this week's tasks?"

[0925] Step 4:

[0926] Text data integration

[0927] The server receives and integrates text data obtained from LINE and other messaging apps and text data converted from voice.

[0928] Input: Text data converted from speech and text data from messaging apps

[0929] Output: Integrated text data

[0930] The server performs data cleansing processing, removing unnecessary symbols and spaces, and formatting the data into a unified format, thereby improving the quality of the data.

[0931] Step 5:

[0932] sentiment analysis

[0933] The server inputs the cleansed text data into the emotion engine, which recognizes the user's emotion for each text and calculates an emotion score.

[0934] Input: Cleansed text data

[0935] Output: Sentiment analysis results for each text

[0936] For example, the emotion engine recognizes "impatience" in response to questions such as "How will you progress with your tasks this week?" and "When is the next meeting?"

[0937] Step 6:

[0938] Timestamp of emotion data

[0939] The server assigns a timestamp to each piece of text data and organizes the data in chronological order.

[0940] Input: Sentiment analysis results

[0941] Output: Emotion data with timestamps

[0942] For example, it records that "How will you proceed with this week's tasks?" was said at 9:00 AM, and "When is the next meeting?" was said at 9:10 AM.

[0943] Step 7:

[0944] Visualizing Emotion Data

[0945] The server converts the emotion data into visual formats such as graphs and charts, allowing users to intuitively understand their emotional fluctuations throughout the day.

[0946] Input: Emotion data with timestamps

[0947] Output: Visualized emotional fluctuation data

[0948] Specifically, it generates a line graph showing emotional fluctuations and displays it on the user interface.

[0949] Step 8:

[0950] Generate improvement suggestions

[0951] The server analyzes the user's emotional tendencies based on the results of the emotion analysis, generates specific improvement suggestions based on the analysis results, and notifies the user through the user interface.

[0952] Input: Sentiment analysis results and trends

[0953] Output: Suggested improvements to the user

[0954] As a specific action, the suggestion is to "Try taking some deep breaths to relax before your morning meeting."

[0955] Through the above processing steps, the system can integrate voice data and text messages, perform sentiment analysis, and provide appropriate improvement suggestions to users, thereby effectively supporting their emotional health management.

[0956] (Application example 2)

[0957] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0958] Autonomous vehicles require real-time monitoring of passengers' emotional states and the provision of appropriate support based on those emotions. In particular, a system that can automatically generate and implement specific suggestions and actions to reduce passenger stress and provide a comfortable travel environment is required.

[0959] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0960] In this invention, the server includes means for collecting voice data, means for converting the collected voice data into text data, means for collecting message data, means for integrating the voice data and the message data, means for sentiment analysis of the integrated data, means for generating improvement suggestions based on the sentiment analysis results, and means for collecting passenger voice data and message data, conducting sentiment analysis, and generating stress reduction suggestions. This makes it possible to grasp the passenger's real-time emotional state and generate and provide appropriate stress reduction suggestions according to the situation.

[0961] "Voice data" refers to data in which the user's voice is recorded in digital form.

[0962] "Text data" refers to character information data obtained by converting voice data into text.

[0963] "Message data" refers to data of text messages sent and received by users.

[0964] "Merge" is the process of combining the text data converted from the voice data and the message data into a single data set.

[0965] "Sentiment analysis" is the process of analyzing a user's emotional state (e.g., joy, sadness, impatience, etc.) from text data.

[0966] "Improvement suggestions" are specific actions or advice suggested to the user based on the results of sentiment analysis.

[0967] "Passenger" refers to a person using an autonomous vehicle.

[0968] "Stress reduction suggestions" are suggestions for reducing stress, taking into account the emotional state of passengers.

[0969] The present invention provides a system for collecting and integrating voice and message data, performing sentiment analysis, and generating stress reduction suggestions based on the results, which is particularly applicable to passenger emotion monitoring and stress reduction assistance for autonomous vehicles.

[0970] System configuration

[0971] To realize this system, the following hardware and software are required.

[0972] 1. Terminal

[0973] Microphones placed inside the vehicle: collect passenger voice data.

[0974] In-car infotainment system: Collects passenger text messages through messaging applications.

[0975] 2. Server

[0976] Speech recognition engine: Converts collected voice data into text data.

[0977] Data integration engine: Integrates text data converted from voice data with message data.

[0978] Sentiment Engine: Performs sentiment analysis using the integrated data.

[0979] Suggestion generation engine: Generates stress reduction suggestions based on the sentiment analysis results.

[0980] System operation procedure

[0981] Audio data collection

[0982] When passengers speak in the car, the microphones collect their words and generate an audio file. For example, if they say, "What should we do about our next meeting?", that will be recorded.

[0983] Message Data Collection

[0984] Text data is collected when a passenger sends a message using a messaging app connected to the vehicle's infotainment system. For example, if a passenger sends a message in a messaging app saying, "When is our next meeting?", that message is captured.

[0985] Data Preprocessing

[0986] The server receives the collected voice files and converts them into text data using a speech recognition engine. For example, "What should we do about the next meeting?" is converted into text "What should we do about the next meeting?"

[0987] The server combines the text data obtained from the messaging application with the text data converted from the voice data, and after a data cleansing process, it is formatted into a unified format.

[0988] sentiment analysis

[0989] The server inputs the formatted text data into the emotion engine, which calculates an emotion score (e.g., impatience, tension, etc.) based on passenger comments such as "What should we do about the next meeting?" and "When is the next meeting?"

[0990] Proposal Generation

[0991] The server generates stress reduction suggestions based on the emotion analysis results. For example, if the emotion engine recognizes "impatience," it generates a suggestion such as "Would you like to play some relaxing music?"

[0992] Suggestions are communicated to passengers through the vehicle's infotainment system, which will either play music automatically or allow passengers to select the suggestion.

[0993] Specific examples

[0994] If a passenger says, "What should we do about our next meeting?", the server converts the speech to text and analyzes that text with an emotion engine. If the emotion engine recognizes the emotion "tension," it generates a suggestion: "Take a deep breath and relax."

[0995] Similarly, if a passenger sends a message in a messaging app asking, "When is my next meeting?", sentiment analysis will be performed and stress-reducing suggestions will be generated.

[0996] Prompt Sentence Examples

[0997] The following are examples of prompts to input into the generative AI model:

[0998] Collect passenger comments, analyze their emotions, and generate stress-reducing suggestions. Below is an example of a passenger comment and the results of the emotion analysis.

[0999] Say: "What about our next meeting?"

[1000] Sentiment Analysis: Negative

[1001] In these cases, generate an appropriate suggestion, for example, "Would you like to play some relaxing music?"

[1002] This system can monitor the emotional state of passengers in autonomous vehicles in real time and provide appropriate stress reduction suggestions, allowing passengers to enjoy a more comfortable and safer travel environment.

[1003] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1004] Step 1:

[1005] Audio data collection

[1006] Input: User's spoken utterance.

[1007] The device uses the in-car microphone to collect the user's speech in real time and record it as an audio file.

[1008] Output: Collected audio data (audio files).

[1009] Step 2:

[1010] Message Data Collection

[1011] Input: A text message that a user sends using a messaging application.

[1012] The device works with the car's infotainment system to retrieve text messages from messaging applications.

[1013] Output: Collected message data (text messages).

[1014] Step 3:

[1015] Converting voice data to text

[1016] Input: Audio data (audio file).

[1017] The server uses a speech recognition engine to convert the voice data into text data.

[1018] Specifically, a speech recognition engine analyzes the collected audio files and generates corresponding text.

[1019] Output: Text data converted from audio.

[1020] Step 4:

[1021] Data integration

[1022] Input: Text data converted from audio and message data.

[1023] The server integrates the text data converted from the voice with the message data and performs a cleansing process.

[1024] The cleansing process is a process of arranging data into a unified format and removing unnecessary symbols and spaces.

[1025] Output: Consolidated clean text data.

[1026] Step 5:

[1027] sentiment analysis

[1028] Input: Integrated text data.

[1029] The server uses an emotion engine to analyze the integrated text data and calculate an emotion score.

[1030] The emotion engine identifies an emotional state, such as "impatience" or "tension," based on the text data.

[1031] Output: Emotion score and emotional state identification results.

[1032] Step 6:

[1033] Generate improvement suggestions

[1034] Input: Emotion scores and emotional state identification results.

[1035] The server generates stress reduction suggestions based on the results of the sentiment analysis.

[1036] A specific suggestion generation engine generates specific actions such as "play relaxing music" or "encourage deep breathing."

[1037] Output: Specific stress reduction suggestions.

[1038] Step 7:

[1039] Proposal Notification

[1040] Enter: stress reduction suggestions.

[1041] The device will then notify passengers of stress-reducing suggestions through the vehicle's infotainment system.

[1042] The passenger can select a suggestion or an automatic action will be initiated to implement the suggestion.

[1043] Output: Passenger notification and action taken.

[1044] Through these processing steps, the system can monitor the user's emotional state in real time and provide appropriate stress reduction suggestions, allowing passengers to enjoy a more comfortable travel environment.

[1045] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1046] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1047] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1048] [Third embodiment]

[1049] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1050] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1051] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1052] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1053] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1054] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1055] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1056] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1057] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1058] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1059] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1060] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1061] The present invention is a system that collects voice data and text data, integrates the data, performs sentiment analysis, and generates improvement suggestions for users based on emotional fluctuations. To realize this system, hardware and software with the following functions are required.

[1062] System configuration

[1063] 1. Device for collecting voice data

[1064] The user's voice is recorded by using a meeting application or the like.

[1065] 2. Device for collecting message data

[1066] Get text messages sent and received from LINE and other messaging applications.

[1067] 3. Data analysis server

[1068] A speech recognition engine for converting collected voice data into text data.

[1069] An analysis engine that integrates message data and text data converted from voice to perform sentiment analysis.

[1070] An engine for generating specific improvement suggestions for users based on the results of sentiment analysis.

[1071] System operation procedure

[1072] Data collection

[1073] Audio data collection

[1074] A user speaks during a conference using a meeting application.

[1075] The device records what is said, generates an audio file, and sends it to the server.

[1076] Message Data Collection

[1077] A user sends a message via LINE.

[1078] The device receives the LINE text message and sends it to the server.

[1079] Data Preprocessing

[1080] Converting voice data to text

[1081] The server inputs the received voice file into a voice recognition engine, which converts the voice into text data.

[1082] Text data integration

[1083] The server integrates the converted voice data with the message data obtained from LINE.

[1084] The integrated data is cleansed to remove unnecessary symbols and spaces.

[1085] sentiment analysis

[1086] Sentiment analysis of text data

[1087] The server inputs the cleansed text data into a sentiment analysis engine.

[1088] A sentiment analysis engine calculates an emotion score (e.g., happy, sad, anger) for each text block.

[1089] Data integration

[1090] Timestamp of emotion data

[1091] The server assigns a timestamp to each piece of data and organizes it in chronological order.

[1092] Visualizing Emotion Data

[1093] The server converts the emotional data into a visual format such as graphs and charts, and displays it so that users can intuitively understand the fluctuations in their emotions throughout the day.

[1094] Proposal Generation

[1095] Generate improvement suggestions

[1096] The server analyzes the user's emotional tendencies based on the emotion analysis results.

[1097] The server generates appropriate improvement suggestions for the user based on the analysis results (e.g., "Try taking deep relaxation breaths before a meeting").

[1098] Specific examples

[1099] Example 1:

[1100] A user says during a meeting, "What are your tasks for today?"

[1101] The device records what is said and sends the audio file to the server.

[1102] The server converts the audio file into text data: "What is your task today?"

[1103] A user sends a message on LINE asking, "How did the afternoon meeting go?"

[1104] The terminal sends the message to the server.

[1105] The server combines both sets of text data and feeds it into a sentiment analysis engine.

[1106] The emotion analysis detects "impatience" and generates a suggestion based on that, such as "We recommend you relax before your afternoon meeting."

[1107] In this way, the system of the present invention mainly operates in cooperation with the server, the terminal, and the user, and constitutes a specific means for supporting the user's emotional health management.

[1108] The processing flow will be explained below.

[1109] Step 1:

[1110] A user speaks up during a meeting, for example, "What are our goals for next week?"

[1111] Step 2:

[1112] The device records what the user says and generates an audio file that is temporarily stored on the device.

[1113] Step 3:

[1114] The terminal transmits the recorded audio file to the server via the meeting application.

[1115] Step 4:

[1116] The server inputs the received voice file into a voice recognition engine, which converts the voice data into text data, generating a string of characters such as "What are your goals for next week?"

[1117] Step 5:

[1118] A user sends a message on LINE saying, "What day is the meeting next week?"

[1119] Step 6:

[1120] The device retrieves message data from the LINE application and sends it to the server.

[1121] Step 7:

[1122] The server combines the received LINE message data with the text data converted from the voice. At this time, a cleansing process is performed to unify the data and remove unnecessary symbols and spaces.

[1123] Step 8:

[1124] The server inputs the formatted text data into a sentiment analysis engine, which calculates sentiment scores for the texts "What are your goals for next week?" and "What day of the week is the meeting next week?"

[1125] Step 9:

[1126] The server assigns a timestamp to each block of text based on the sentiment score obtained from the sentiment analysis engine, for example, recording that "What are your goals for next week?" was spoken at 10:00 AM and "What day is our meeting next week?" was spoken at 10:05 AM.

[1127] Step 10:

[1128] The server chronologically organizes the time-stamped emotion data and summarizes the emotional fluctuations by daily activity. The emotion data is then converted into a visual format such as a graph or chart.

[1129] Step 11:

[1130] The server analyzes the user's emotional tendencies based on the emotional data and timestamp data, and obtains analysis results such as "people often feel anxious during morning meetings."

[1131] Step 12:

[1132] Based on the analysis results, the server generates improvement suggestions for the user, such as "I suggest you meditate for five minutes before your morning meeting."

[1133] Step 13:

[1134] The user checks the suggestions from the server on a device such as a smartphone and adjusts their behavior based on the advice.

[1135] Example 1

[1136] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1137] In today's communication environment, there is a lack of methods to properly understand users' emotional fluctuations and stress levels and support their emotional health management. In particular, there is a need for a method that can more accurately assess users' emotional state and provide effective improvement suggestions by integrating and analyzing voice data and text messages.

[1138] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1139] In this invention, the server includes means for collecting voice data, means for converting the collected voice data into text data, means for collecting text messages, means for integrating the voice data and the text messages, means for preprocessing the integrated data, means for sentiment analysis of the preprocessed data, and means for generating improvement suggestions based on the sentiment analysis results. This makes it possible to comprehensively analyze a user's various communication data, appropriately evaluate the emotional state, and provide effective improvement suggestions.

[1140] "Audio data" refers to digital audio files that record the content of user speech occurring during meetings, conversations, and the like.

[1141] "Text data" refers to character information converted from voice data or character information sent and received by a message application.

[1142] A "text message" is text information that a user sends and receives using a text application (e.g., a messaging app).

[1143] "Preprocessing" is the procedure of removing unnecessary information and extracting necessary parts in order to prepare data in a format that can be analyzed.

[1144] "Sentiment analysis" is the process of analyzing a user's emotional state (e.g., joy, sadness, anger) from text data and classifying it into scores or categories.

[1145] "Improvement suggestions" are specific advice on improving behavior or status that is provided to the user based on the results of sentiment analysis.

[1146] "Fusion" refers to combining multiple data sources (e.g., voice data and text messages) into a single dataset.

[1147] A "timestamp" is a means of clarifying time-series information by adding information about the date and time when data was generated or collected.

[1148] "Visualization" is the process of transforming data into a visually understandable format such as a graph or chart.

[1149] The present invention provides a system for collecting voice data and text messages, integrating the collected data to perform sentiment analysis, and generating improvement suggestions for users based on emotional fluctuations. Specific embodiments are described below.

[1150] Hardware and Software Configuration

[1151] 1. Device for collecting audio data:

[1152] A user speaks during a meeting using a meeting application (e.g., an online conference system).

[1153] The terminal records the user's speech during the conference, generates an audio file (e.g., a WAV file), and sends it to the server.

[1154] 2. Devices for Text Message Data Collection:

[1155] A user has a conversation using a messaging application (e.g., a messaging app).

[1156] The terminal retrieves the text message from the message application and sends it to the server as a text file.

[1157] 3. Data analysis server:

[1158] The server uses speech recognition and sentiment analysis engines such as Google Cloud Speech-to-Text and IBM Watson Natural Language Understanding to convert the voice data into text data and perform further sentiment analysis.

[1159] The server includes an engine for generating specific improvement suggestions for the user based on the sentiment analysis results.

[1160] Specific examples

[1161] Example 1:

[1162] A user says, "What are your tasks for today?" during an online meeting.

[1163] The device records what is said and sends the audio file to the server.

[1164] The server inputs the audio file into the Google Cloud Speech-to-Text engine and converts it into text data: "What is your task today?"

[1165] A user sends a message in a messaging app asking, "How did the meeting this afternoon go?"

[1166] The terminal receives the message and sends it to the server.

[1167] The server integrates both sets of text data and performs a cleansing process.

[1168] The server inputs the cleansed text data into IBM Watson Natural Language Understanding for sentiment analysis.

[1169] As a result of the emotion analysis, "impatience" is detected, and based on that, an improvement suggestion is generated, such as "recommending relaxation before the afternoon meeting."

[1170] Examples of prompt statements

[1171] Example prompt sentence:

[1172] "A user might say, 'What are your tasks for today?' in an online meeting, and then later send a message in a messaging app asking, 'How did your afternoon meeting go?' Combine these data points, perform sentiment analysis, and generate appropriate improvement suggestions."

[1173] In this way, the system of the present invention constitutes a concrete means for supporting the emotional health management of users, with the server, terminal, and user working together.

[1174] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1175] Step 1:

[1176] Collection and transmission of voice data

[1177] A user speaks during a conference using an online conference system.

[1178] The terminal records the user's speech and generates an audio file (e.g., WAV format).

[1179] The terminal transmits the generated audio file to the server.

[1180] Input: User's speech in an online conference system

[1181] Data processing: Recording audio data and generating WAV format audio files

[1182] Output: Audio file sent to the server

[1183] Specific operation:

[1184] Execute a script to record speech on the online conference system, generate an audio file, and upload the generated audio file to the server using FTP or HTTP protocol.

[1185] Step 2:

[1186] Collection and transmission of message data

[1187] A user sends a text message using a messaging application.

[1188] The terminal acquires text messages sent and received from the message application and saves them as text files.

[1189] The terminal transmits the obtained text file to the server.

[1190] Input: The user's text message in the Messages application

[1191] Data processing: Obtaining text messages and generating text files in TXT format

[1192] Output: A text file that is sent to the server.

[1193] Specific operation:

[1194] When a new message is detected, a script running in the background of the Messages app extracts its contents, saves the extracted text message in a text file, and uploads it to a server.

[1195] Step 3:

[1196] Speech-to-text

[1197] The server inputs the received audio file into a speech recognition engine (e.g., Google Cloud Speech-to-Text).

[1198] The server converts the audio file into text data using a speech recognition engine.

[1199] The server stores the converted text data in temporary storage.

[1200] Input: Audio file sent to the server

[1201] Data processing: Converting audio files into text using a speech recognition engine

[1202] Output: Text data saved in temporary storage

[1203] Specific operation:

[1204] A Python script is executed on the server to send the audio file to the Google Cloud Speech-to-Text API, and the text data returned by the API is retrieved in JSON format and saved in a database on the server.

[1205] Step 4:

[1206] Text data integration and preprocessing

[1207] The server integrates the text data converted from the voice data with the text file obtained from the messaging app.

[1208] The server cleanses the integrated text data, removing unnecessary symbols and spaces.

[1209] Input: Text data converted from voice, text data obtained from messaging apps

[1210] Data processing: text data integration and cleansing

[1211] Output: Integrated text data after cleansing

[1212] Specific operation:

[1213] The integration process runs an SQL query to combine the voice text and message text, and uses regular expressions (RegEx) on the combined data to remove unnecessary symbols and spaces.

[1214] Step 5:

[1215] Conducting sentiment analysis

[1216] The server inputs the cleansed text data into a sentiment analysis engine (e.g., IBM Watson Natural Language Understanding).

[1217] The server obtains the sentiment score (e.g., happy, sad, anger) for each text block returned by the sentiment analysis engine.

[1218] Input: Text data after cleansing

[1219] Data calculation: Calculating sentiment scores using a sentiment analysis engine

[1220] Output: Text data with sentiment scores

[1221] Specific operation:

[1222] The server makes an API request to send the cleansed text data to the sentiment analysis engine, which analyzes the resulting sentiment scores and stores them in a database.

[1223] Step 6:

[1224] Timestamp and organize emotion data

[1225] The server assigns a timestamp to each piece of emotion data and organizes it in chronological order.

[1226] Input: Text data with sentiment scores

[1227] Data processing: Adding timestamps and organizing data in chronological order

[1228] Output: Emotion data with timestamps and organized in chronological order

[1229] Specific operation:

[1230] The Python Pandas library is used to assign timestamps, and a sorting algorithm is applied to organize the data in chronological order, before storing the results in a database.

[1231] Step 7:

[1232] Visualizing Emotion Data

[1233] The server converts the emotional data into a visual format such as a graph or chart, and displays it so that the user can intuitively understand the fluctuations in their emotions throughout the day.

[1234] Input: Organized emotion data

[1235] Data processing: Converting data into graphs, charts, etc.

[1236] Output: Emotion data displayed in a visually understandable format

[1237] Specific operation:

[1238] We create timeline graphs using libraries such as Matplotlib and Plotly, and display the generated graphs on a web dashboard to help users intuitively understand the fluctuations in sentiment throughout the day.

[1239] Step 8:

[1240] Generate improvement suggestions

[1241] The server analyzes the user's emotional tendency based on the emotion analysis result.

[1242] The server generates specific improvement suggestions for the user from the analysis results.

[1243] Input: Sentiment analysis results

[1244] Data calculation: analyzing emotional trends and generating improvement suggestions

[1245] Output: Improvement suggestions

[1246] Specific operation:

[1247] It uses machine learning models to learn patterns and trends from past data and generate suggestions for future improvements, saving the suggestions in a text file in natural language format and sending them to a module that notifies the user.

[1248] (Application example 1)

[1249] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1250] Conventional emotion analysis systems only collect and analyze voice and text data, making it difficult to utilize the resulting emotion information to detect risks or propose countermeasures. In particular, to improve corporate security and safety, it is necessary to quickly grasp changes in users' emotions and propose appropriate countermeasures. To solve this problem, a system is needed that can detect risks from emotion analysis results and propose countermeasures in a timely manner.

[1251] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1252] In this invention, the server includes means for collecting voice data, means for converting the collected voice data into text data, means for collecting message data, means for integrating the voice data and the message data, means for sentiment analysis of the integrated data, means for generating improvement proposals based on the sentiment analysis results, and means for detecting risks from the collected data and proposing appropriate responses. This makes it possible to grasp the emotional state of the user in real time, detect potential risks early, and propose appropriate measures.

[1253] A "means for collecting voice data" is a device or software that records a user's voice and stores it as digital data.

[1254] The "means for converting collected voice data into text data" refers to a device or software that converts voice data into text using voice recognition technology.

[1255] A "means for collecting message data" is a device or software that obtains text data sent and received from a messaging application.

[1256] "Means for integrating voice data and message data" refers to a device or software that unifies data collected in different formats into one format.

[1257] A "means for sentiment analysis of integrated data" is a device or software that analyzes and evaluates sentiment from collected text data.

[1258] The "means for generating improvement suggestions based on the results of sentiment analysis" is a device or software that provides a specific action plan or advice to the user based on the results of sentiment analysis.

[1259] "Means for detecting risks from collected data and proposing appropriate responses" refers to devices or software that analyze emotional data, identify potential risks, and suggest necessary measures.

[1260] The "means for assigning timestamps and organizing the results of sentiment analysis in chronological order" refers to a device or software that assigns time information to each piece of data and arranges it in chronological order.

[1261] "Means for converting the results of sentiment analysis into a visual format such as a graph or chart" refers to a device or software that visually displays the analysis results so that they can be intuitively understood.

[1262] The "means for warning the user of the occurrence of a risk" refers to a device or software that notifies the user when a risk increases based on the analysis results.

[1263] To implement this invention, a system with the following functions is required. First, as a means for collecting voice data, a microphone on a smartphone or smart glasses is used to record the user's voice in real time. As a means for converting collected voice data into text data, voice recognition technology is used. Google Speech-to-Text is a suitable software.

[1264] Additionally, the "means of collecting message data" involves using APIs to obtain text data from messaging applications used within the company (e.g., Slack, Microsoft Teams). The terminals collecting this data must be appropriate devices connected to the server.

[1265] Next, the different formats of data are unified into one format using a "means for integrating voice data and message data." This allows the data to be analyzed using a "means for sentiment analysis of the integrated data" to calculate an emotional score. For this part, a sentiment analysis engine such as Amazon Comprehend is used.

[1266] Furthermore, the "means for generating improvement proposals based on the results of sentiment analysis" proposes specific measures to users based on the obtained sentiment data. This system is capable of detecting risks from the collected data and proposing appropriate responses in a timely manner.

[1267] The "means of assigning timestamps and organizing the sentiment analysis results in chronological order" involves assigning time information to each piece of data and organizing it in chronological order. This allows for an intuitive understanding of sentiment fluctuations. The "means of converting the sentiment analysis results into visual formats such as graphs and charts and alerting users to emerging risks" involves visually displaying them using D3.js and Grafana.

[1268] As a specific example, audio spoken by a user during a meeting is collected using a smartphone microphone and sent to a server. The server then uses Google Speech-to-Text to convert the audio into text data. In parallel, text data sent and received by the user via a messaging application is also sent to the server, and both sets of data are integrated. Sentiment analysis is then performed using Amazon Comprehend, and risks are detected based on the results. For example, specific improvement suggestions are displayed, such as, "A high stress level has been detected from comments made during the meeting. We recommend that you take a break to relax."

[1269] An example prompt is, "Generate an application that uses Google Speech-to-Text to convert employees' real-time voice data into text data, perform sentiment analysis, detect security risks early, and generate improvement suggestions."

[1270] In this way, the entire system works together to grasp the user's emotional state in real time, detect potential risks early, and propose appropriate countermeasures.

[1271] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1272] Step 1:

[1273] Audio data collection

[1274] Users communicate using the microphone on their smartphones or smart glasses. The device collects voice data in real time through the microphone and stores it in a digital file format. The input is the user's voice, and the output is a voice data file.

[1275] Step 2:

[1276] Converting audio data to text

[1277] The device sends the collected voice data to the server, which then uses a voice recognition engine (e.g., Google Speech-to-Text) to convert the voice data into text data. The input is a voice data file, and the output is text data.

[1278] Step 3:

[1279] Message Data Collection

[1280] The device acquires text data from messaging applications used within the company (e.g., Slack, Microsoft Teams), and uses an API to send the message data to the server. The input is the text message from the messaging application, and the output is the collected message data.

[1281] Step 4:

[1282] Data integration

[1283] The server integrates the text data converted from the voice with the message data obtained from the messaging application. Data integration is a process for unifying data of different formats into a single format. The input is the text data from the voice and the message data, and the output is the integrated text data.

[1284] Step 5:

[1285] sentiment analysis

[1286] The server inputs the integrated text data into a sentiment analysis engine (e.g., Amazon Comprehend). The sentiment analysis engine calculates an emotional score (e.g., happy, sad, or angry) from the text data. The input is the integrated text data, and the output is the emotional score.

[1287] Step 6:

[1288] Risk detection and response proposals

[1289] The server runs an algorithm to detect potential risks from the sentiment analysis results. If a risk is detected, the server generates a response suggestion (e.g., "High stress levels detected. We recommend taking a break to relax."). The input is the sentiment score, and the output is the risk detection result and a response suggestion.

[1290] Step 7:

[1291] Organizing data chronologically

[1292] The server assigns a timestamp to each voice and message data and organizes the emotion analysis results in chronological order, making it easier to understand emotion fluctuations throughout the day. The input is the emotion score and the original data, and the output is the time-stamped data.

[1293] Step 8:

[1294] Visualization of sentiment analysis results

[1295] The server converts the sentiment analysis results into visual formats such as graphs and charts and displays them to the user. This process uses visualization tools such as D3.js and Grafana. The input is time-stamped sentiment data, and the output is visualized graphs and charts.

[1296] Step 9:

[1297] Risk warning notification

[1298] If a risk is detected, the server sends a notification to the user to warn them of the risk. This notification can be in the form of a push notification to a smartphone or other device. The input is the risk detection result, and the output is a warning notification to the user.

[1299] By following these steps, users can understand their own emotional state in real time, detect potential risks early, and take appropriate measures.

[1300] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1301] This invention combines a system that collects voice data and message data, integrates the data, performs emotion analysis, and generates improvement proposals, with an emotion engine that recognizes the user's emotions. To realize this system, hardware and software with the following functions are required.

[1302] System configuration

[1303] 1. Device for collecting voice data

[1304] A terminal that uses a meeting application or the like to record user speech.

[1305] 2. Device for collecting message data

[1306] A device that retrieves text messages from a messaging application (e.g., LINE).

[1307] 3. Data analysis server

[1308] A speech recognition engine for converting collected voice data into text data.

[1309] An analysis engine that integrates message data and text data converted from voice to perform sentiment analysis.

[1310] Processing to recognize user emotions using an emotion engine.

[1311] An engine that generates specific improvement suggestions for users based on the results of sentiment analysis.

[1312] System operation procedure

[1313] Data collection

[1314] Audio data collection

[1315] A user speaks during a meeting using a meeting application, for example, "How are we progressing with this week's tasks?"

[1316] The device records the user's speech and generates an audio file.

[1317] Message Data Collection

[1318] A user sends a message on LINE asking, "When is the next meeting?"

[1319] The device receives the LINE text message.

[1320] Data Preprocessing

[1321] Converting voice data to text

[1322] The server receives the audio file and inputs it into a speech recognition engine, which converts the audio data into text data. For example, it generates text data such as "How will you proceed with this week's tasks?"

[1323] Text data integration

[1324] The server combines the text data obtained from LINE with the text data converted from the voice data. The data is then formatted into a unified format and a cleansing process is performed to remove unnecessary symbols and spaces.

[1325] sentiment analysis

[1326] Sentiment analysis of text data

[1327] The server inputs the formatted text data into the emotion engine, which directly recognizes the user's emotions from the text "How will you proceed with your tasks this week?" and "When is the next meeting?" and calculates an emotion score (e.g., joy, sadness, impatience, etc.).

[1328] Data integration

[1329] Timestamp of emotion data

[1330] The server assigns a timestamp to each piece of data and organizes it in chronological order. For example, it records that "How are we progressing with our tasks this week?" was said at 9:00 AM and "When is our next meeting?" was said at 9:10 AM.

[1331] Visualizing Emotion Data

[1332] The server converts the emotional data into a visual format such as graphs and charts, and displays it so that users can intuitively understand the fluctuations in their emotions throughout the day.

[1333] Proposal Generation

[1334] Generate improvement suggestions

[1335] The server analyzes the user's emotional tendencies based on the results of the emotion analysis, and obtains an analysis result such as "people often feel anxious during morning meetings."

[1336] The server generates appropriate improvement suggestions for the user based on the analysis results (e.g., "I suggest you meditate for five minutes before your morning meeting").

[1337] Specific examples

[1338] Example 1:

[1339] A user says during a meeting, "How are we progressing with our tasks this week?"

[1340] The device records what is said and sends the audio file to the server.

[1341] The server converts the audio file into text data such as "How will you proceed with your tasks this week?"

[1342] A user sends a message on LINE asking, "When is the next meeting?"

[1343] The terminal sends the message to the server.

[1344] The server combines both sets of text data and inputs them into the emotion engine.

[1345] The emotion engine recognizes the emotion "impatient" for "How are you progressing with your tasks this week?" and "When is your next meeting?"

[1346] Based on the sentiment score, the server generates a suggestion such as "Try taking some deep breaths to relax before your morning meeting."

[1347] In this way, the system of the present invention operates in cooperation with the server, the terminal, and the user, and in particular, combines the emotion engine to form a specific means for supporting the user's emotional health management.

[1348] The processing flow will be explained below.

[1349] Step 1:

[1350] A user says during a meeting, "What are our goals for next week?"

[1351] Step 2:

[1352] The device records what the user says and generates an audio file that is temporarily stored on the device.

[1353] Step 3:

[1354] The terminal transmits the recorded audio file to the server via the meeting application.

[1355] Step 4:

[1356] The server inputs the received voice file into a speech recognition engine, which converts the voice data into text data and generates the string "What are your goals for next week?"

[1357] Step 5:

[1358] A user sends a message on LINE saying, "What day is the meeting next week?"

[1359] Step 6:

[1360] The device retrieves message data from the LINE application and sends it to the server.

[1361] Step 7:

[1362] The server combines the received LINE message data with the text data converted from the voice. At this time, a cleansing process is performed to unify the data and remove unnecessary symbols and spaces.

[1363] Step 8:

[1364] The server inputs the formatted text data into the emotion engine, which calculates emotion scores for the texts "What are your goals for next week?" and "What day of the week is the meeting next week?"

[1365] Step 9:

[1366] The server recognizes the emotion of each text based on the emotion score obtained from the emotion engine. For example, "What are your goals for next week?" will get an emotion score of "Interest," and "What day of the week is the meeting next week?" will get an emotion score of "Impatience."

[1367] Step 10:

[1368] The server assigns a timestamp to the sentiment score and organizes the data chronologically, for example, recording that "What are your goals for next week?" was said at 10:00 AM and "What day is the meeting next week?" at 10:05 AM.

[1369] Step 11:

[1370] The server converts the time-stamped emotion data into visual formats such as graphs and charts, allowing users to intuitively understand the fluctuations of emotions throughout the day.

[1371] Step 12:

[1372] The server analyzes the user's emotional tendencies based on the emotional data and timestamp data. For example, it determines whether the user frequently feels impatient during certain times of the day.

[1373] Step 13:

[1374] The server generates improvement suggestions for the user based on the analysis results, such as "Try taking deep breaths to relax before your morning meeting."

[1375] Step 14:

[1376] The user checks the suggestions from the server on a device such as a smartphone and adjusts their behavior based on the advice.

[1377] Example 2

[1378] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1379] Conventional systems collect and analyze voice data and message data separately, making it difficult to comprehensively and accurately recognize user emotions. Furthermore, the collected data is often insufficiently cleansed, hindering accurate emotion analysis. Furthermore, it is difficult to provide specific improvement suggestions to users based on the results of emotion analysis, and there is a lack of a way to visually understand daily emotional fluctuations.

[1380] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1381] In this invention, the server includes means for collecting voice data, means for converting the collected voice data into text data, means for collecting message data, means for integrating the voice data and the message data, means for sentiment analysis of the integrated data, means for generating improvement suggestions based on the sentiment analysis results, means for data cleansing to remove unnecessary symbols and spaces from each data, means for generating suggestions to support the user's emotional health management based on the sentiment analysis results, and means for organizing the emotional data in timestamp order. This enables the integrated collection and analysis of voice and message data to accurately recognize the user's emotions. Furthermore, the cleansing process improves data accuracy, and appropriate improvement suggestions can be provided to support the user's emotional health. Organizing the emotional data in timestamp order allows for a visual understanding of emotional fluctuations.

[1382] A "means for collecting voice data" is a device or application that records a user's speech or voice and stores it as a digital audio file.

[1383] The "means for converting collected voice data into text data" refers to a voice recognition engine or software for analyzing voice files and converting their contents into text data in sentence format.

[1384] A "means for collecting message data" is a device or application that captures and stores text messages or chat messages.

[1385] A "means for integrating voice and message data" is software or a process for consolidating data collected from different formats and sources into a single format and managing it in a unified manner.

[1386] "Means for sentiment analysis of the integrated data" refers to an emotion engine or software for analyzing the integrated text data and determining a user's emotion based on the content of the text.

[1387] The "means for generating improvement suggestions based on the results of sentiment analysis" is software or an engine for generating feedback and advice for users based on the results of sentiment analysis.

[1388] A "data cleansing method that removes unnecessary symbols and spaces from each piece of data" is a process or software that automatically removes unnecessary symbols and spaces from collected data to improve the quality of the data.

[1389] The "means for generating suggestions to support the user's emotional health management based on the results of sentiment analysis" is software or an engine for utilizing the results of sentiment analysis to present specific advice and schedules for improving the user's emotional health.

[1390] "Means for organizing emotion data in timestamp order" refers to software or a process for adding time information to collected emotion data and organizing and storing the data in chronological order.

[1391] The present invention is a system that collects voice data and message data, integrates the data, performs sentiment analysis, and generates improvement proposals. This system requires hardware and software with the functions of voice data collection, message data collection, data integration, sentiment analysis, data organization and visualization, and generation of improvement proposals.

[1392] Audio data collection

[1393] Users use a meeting application (e.g., a video conferencing application) to hold a conversation. A device (e.g., a user's PC or smartphone) records the audio during the meeting and generates an audio file. This audio file is automatically uploaded to cloud storage.

[1394] Message Data Collection

[1395] A user sends a message using a messaging application (e.g., a text messaging app), and the device receives the text message and sends it to a server via a dedicated application.

[1396] Converting audio data to text

[1397] The server receives the audio file from the cloud storage. The received audio file is input into a speech recognition engine (e.g., a speech recognition API) and converted into text data. For example, a statement such as "How will you proceed with this week's tasks?" is generated as text data.

[1398] Text data integration

[1399] The server combines the text data obtained from LINE and other messaging apps with the text data converted from voice. During the combination process, a data cleansing process is performed to remove unnecessary symbols and spaces from each data, improving the accuracy of the data.

[1400] sentiment analysis

[1401] The server inputs the cleansed text data into an emotion engine (e.g., a natural language processing API). The emotion engine recognizes the user's emotion for each piece of text and calculates an emotion score. Specifically, for questions like "How will you progress with your tasks this week?" and "When is the next meeting?", emotions such as impatience and anticipation are recognized.

[1402] Timestamp of emotion data

[1403] The server assigns a timestamp to each piece of text data and organizes it in chronological order. For example, it records that "How are we progressing with this week's tasks?" was said at 9:00 AM and "When is the next meeting?" was said at 9:10 AM.

[1404] Visualizing Emotion Data

[1405] The server converts the emotion data into a visual format such as a graph or chart, allowing users to intuitively understand their emotional fluctuations throughout the day. For example, a line graph showing the emotional fluctuations throughout the day can be generated and displayed on the user interface.

[1406] Generate improvement suggestions

[1407] The server analyzes the user's emotional tendencies based on the results of the emotion analysis. It then generates specific improvement suggestions based on the analysis results and notifies the user through the user interface. For example, a suggestion such as "Try taking deep breaths to relax before your morning meeting" may be displayed.

[1408] Specific examples

[1409] A user says during a meeting, "How are we progressing with our tasks this week?"

[1410] The device records what is said and sends the audio file to the server.

[1411] The server converts the audio file into text data such as "How will you proceed with your tasks this week?"

[1412] A user sends a message in a messaging app asking, "When is our next meeting?"

[1413] The terminal sends the message to the server.

[1414] The server combines both sets of text data and inputs them into the emotion engine.

[1415] The emotion engine recognizes "impatience" for "How are you progressing with your tasks this week?" and "When is your next meeting?"

[1416] Based on the sentiment score, the server generates a suggestion such as "Try taking some deep breaths to relax before your morning meeting."

[1417] In this way, the system of the present invention operates in cooperation with the server, the terminal, and the user, and in particular, combines the emotion engine to form a specific means for supporting the user's emotional health management.

[1418] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1419] Step 1:

[1420] Audio data collection

[1421] A user has a conversation using a meeting application (e.g., a video conferencing application). The device records the audio during the meeting and sends it to cloud storage in the form of audio data.

[1422] Input: User's voice

[1423] Output: Audio files saved in cloud storage

[1424] Specifically, when a user says, "How will you proceed with this week's tasks?", the device uploads the voice recording to cloud storage as a digital audio file.

[1425] Step 2:

[1426] Message Data Collection

[1427] A user sends a message using a messaging application (e.g., a text messaging app), and the device sends the text message to a server via a dedicated application.

[1428] Input: User's text message

[1429] Output: Text data stored on the server

[1430] Specifically, when a user sends a message such as "When is the next meeting?", the terminal sends the text message to the server.

[1431] Step 3:

[1432] Converting audio data to text

[1433] The server receives the audio file from the cloud storage and inputs it into a speech recognition engine, which analyzes the audio file and converts it into text data.

[1434] Input: Audio files stored in cloud storage

[1435] Output: Text data converted from audio

[1436] For example, the speech recognition engine generates text data such as "How will you proceed with this week's tasks?"

[1437] Step 4:

[1438] Text data integration

[1439] The server receives and integrates text data obtained from LINE and other messaging apps and text data converted from voice.

[1440] Input: Text data converted from speech and text data from messaging apps

[1441] Output: Integrated text data

[1442] The server performs data cleansing processing, removing unnecessary symbols and spaces, and formatting the data into a unified format, thereby improving the quality of the data.

[1443] Step 5:

[1444] sentiment analysis

[1445] The server inputs the cleansed text data into the emotion engine, which recognizes the user's emotion for each text and calculates an emotion score.

[1446] Input: Cleansed text data

[1447] Output: Sentiment analysis results for each text

[1448] For example, the emotion engine recognizes "impatience" in response to questions such as "How will you progress with your tasks this week?" and "When is the next meeting?"

[1449] Step 6:

[1450] Timestamp of emotion data

[1451] The server assigns a timestamp to each piece of text data and organizes the data in chronological order.

[1452] Input: Sentiment analysis results

[1453] Output: Emotion data with timestamps

[1454] For example, it records that "How will you proceed with this week's tasks?" was said at 9:00 AM, and "When is the next meeting?" was said at 9:10 AM.

[1455] Step 7:

[1456] Visualizing Emotion Data

[1457] The server converts the emotion data into visual formats such as graphs and charts, allowing users to intuitively understand their emotional fluctuations throughout the day.

[1458] Input: Emotion data with timestamps

[1459] Output: Visualized emotional fluctuation data

[1460] Specifically, it generates a line graph showing emotional fluctuations and displays it on the user interface.

[1461] Step 8:

[1462] Generate improvement suggestions

[1463] The server analyzes the user's emotional tendencies based on the results of the emotion analysis, generates specific improvement suggestions based on the analysis results, and notifies the user through the user interface.

[1464] Input: Sentiment analysis results and trends

[1465] Output: Suggested improvements to the user

[1466] As a specific action, the suggestion is to "Try taking some deep breaths to relax before your morning meeting."

[1467] Through the above processing steps, the system can integrate voice data and text messages, perform sentiment analysis, and provide appropriate improvement suggestions to users, thereby effectively supporting their emotional health management.

[1468] (Application example 2)

[1469] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1470] Autonomous vehicles require real-time monitoring of passengers' emotional states and the provision of appropriate support based on those emotions. In particular, a system that can automatically generate and implement specific suggestions and actions to reduce passenger stress and provide a comfortable travel environment is required.

[1471] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1472] In this invention, the server includes means for collecting voice data, means for converting the collected voice data into text data, means for collecting message data, means for integrating the voice data and the message data, means for sentiment analysis of the integrated data, means for generating improvement suggestions based on the sentiment analysis results, and means for collecting passenger voice data and message data, conducting sentiment analysis, and generating stress reduction suggestions. This makes it possible to grasp the passenger's real-time emotional state and generate and provide appropriate stress reduction suggestions according to the situation.

[1473] "Voice data" refers to data in which the user's voice is recorded in digital form.

[1474] "Text data" refers to character information data obtained by converting voice data into text.

[1475] "Message data" refers to data of text messages sent and received by users.

[1476] "Merge" is the process of combining the text data converted from the voice data and the message data into a single data set.

[1477] "Sentiment analysis" is the process of analyzing a user's emotional state (e.g., joy, sadness, impatience, etc.) from text data.

[1478] "Improvement suggestions" are specific actions or advice suggested to the user based on the results of sentiment analysis.

[1479] "Passenger" refers to a person using an autonomous vehicle.

[1480] "Stress reduction suggestions" are suggestions for reducing stress, taking into account the emotional state of passengers.

[1481] The present invention provides a system for collecting and integrating voice and message data, performing sentiment analysis, and generating stress reduction suggestions based on the results, which is particularly applicable to passenger emotion monitoring and stress reduction assistance for autonomous vehicles.

[1482] System configuration

[1483] To realize this system, the following hardware and software are required.

[1484] 1. Terminal

[1485] Microphones placed inside the vehicle: collect passenger voice data.

[1486] In-car infotainment system: Collects passenger text messages through messaging applications.

[1487] 2. Server

[1488] Speech recognition engine: Converts collected voice data into text data.

[1489] Data integration engine: Integrates text data converted from voice data with message data.

[1490] Sentiment Engine: Performs sentiment analysis using the integrated data.

[1491] Suggestion generation engine: Generates stress reduction suggestions based on the sentiment analysis results.

[1492] System operation procedure

[1493] Audio data collection

[1494] When passengers speak in the car, the microphones collect their words and generate an audio file. For example, if they say, "What should we do about our next meeting?", that will be recorded.

[1495] Message Data Collection

[1496] Text data is collected when a passenger sends a message using a messaging app connected to the vehicle's infotainment system. For example, if a passenger sends a message in a messaging app saying, "When is our next meeting?", that message is captured.

[1497] Data Preprocessing

[1498] The server receives the collected voice files and converts them into text data using a speech recognition engine. For example, "What should we do about the next meeting?" is converted into text "What should we do about the next meeting?"

[1499] The server combines the text data obtained from the messaging application with the text data converted from the voice data, and after a data cleansing process, it is formatted into a unified format.

[1500] sentiment analysis

[1501] The server inputs the formatted text data into the emotion engine, which calculates an emotion score (e.g., impatience, tension, etc.) based on passenger comments such as "What should we do about the next meeting?" and "When is the next meeting?"

[1502] Proposal Generation

[1503] The server generates stress reduction suggestions based on the emotion analysis results. For example, if the emotion engine recognizes "impatience," it generates a suggestion such as "Would you like to play some relaxing music?"

[1504] Suggestions are communicated to passengers through the vehicle's infotainment system, which will either play music automatically or allow passengers to select the suggestion.

[1505] Specific examples

[1506] If a passenger says, "What should we do about our next meeting?", the server converts the speech to text and analyzes that text with an emotion engine. If the emotion engine recognizes the emotion "tension," it generates a suggestion: "Take a deep breath and relax."

[1507] Similarly, if a passenger sends a message in a messaging app asking, "When is my next meeting?", sentiment analysis will be performed and stress-reducing suggestions will be generated.

[1508] Prompt Sentence Examples

[1509] The following are examples of prompts to input into the generative AI model:

[1510] Collect passenger comments, analyze their emotions, and generate stress-reducing suggestions. Below is an example of a passenger comment and the results of the emotion analysis.

[1511] Say: "What about our next meeting?"

[1512] Sentiment Analysis: Negative

[1513] In these cases, generate an appropriate suggestion, for example, "Would you like to play some relaxing music?"

[1514] This system can monitor the emotional state of passengers in autonomous vehicles in real time and provide appropriate stress reduction suggestions, allowing passengers to enjoy a more comfortable and safer travel environment.

[1515] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1516] Step 1:

[1517] Audio data collection

[1518] Input: User's spoken utterance.

[1519] The device uses the in-car microphone to collect the user's speech in real time and record it as an audio file.

[1520] Output: Collected audio data (audio files).

[1521] Step 2:

[1522] Message Data Collection

[1523] Input: A text message that a user sends using a messaging application.

[1524] The device works with the car's infotainment system to retrieve text messages from messaging applications.

[1525] Output: Collected message data (text messages).

[1526] Step 3:

[1527] Converting voice data to text

[1528] Input: Audio data (audio file).

[1529] The server uses a speech recognition engine to convert the voice data into text data.

[1530] Specifically, a speech recognition engine analyzes the collected audio files and generates corresponding text.

[1531] Output: Text data converted from audio.

[1532] Step 4:

[1533] Data integration

[1534] Input: Text data converted from audio and message data.

[1535] The server integrates the text data converted from the voice with the message data and performs a cleansing process.

[1536] The cleansing process is a process of arranging data into a unified format and removing unnecessary symbols and spaces.

[1537] Output: Consolidated clean text data.

[1538] Step 5:

[1539] sentiment analysis

[1540] Input: Integrated text data.

[1541] The server uses an emotion engine to analyze the integrated text data and calculate an emotion score.

[1542] The emotion engine identifies an emotional state, such as "impatience" or "tension," based on the text data.

[1543] Output: Emotion score and emotional state identification results.

[1544] Step 6:

[1545] Generate improvement suggestions

[1546] Input: Emotion scores and emotional state identification results.

[1547] The server generates stress reduction suggestions based on the results of the sentiment analysis.

[1548] A specific suggestion generation engine generates specific actions such as "play relaxing music" or "encourage deep breathing."

[1549] Output: Specific stress reduction suggestions.

[1550] Step 7:

[1551] Proposal Notification

[1552] Enter: stress reduction suggestions.

[1553] The device will then notify passengers of stress-reducing suggestions through the vehicle's infotainment system.

[1554] The passenger can select a suggestion or an automatic action will be initiated to implement the suggestion.

[1555] Output: Passenger notification and action taken.

[1556] Through these processing steps, the system can monitor the user's emotional state in real time and provide appropriate stress reduction suggestions, allowing passengers to enjoy a more comfortable travel environment.

[1557] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1558] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1559] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1560] [Fourth embodiment]

[1561] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1562] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1563] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1564] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1565] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1566] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1567] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1568] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1569] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1570] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1571] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1572] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1573] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1574] The present invention is a system that collects voice data and text data, integrates the data, performs sentiment analysis, and generates improvement suggestions for users based on emotional fluctuations. To realize this system, hardware and software with the following functions are required.

[1575] System configuration

[1576] 1. Device for collecting voice data

[1577] The user's voice is recorded by using a meeting application or the like.

[1578] 2. Device for collecting message data

[1579] Get text messages sent and received from LINE and other messaging applications.

[1580] 3. Data analysis server

[1581] A speech recognition engine for converting collected voice data into text data.

[1582] An analysis engine that integrates message data and text data converted from voice to perform sentiment analysis.

[1583] An engine for generating specific improvement suggestions for users based on the results of sentiment analysis.

[1584] System operation procedure

[1585] Data collection

[1586] Audio data collection

[1587] A user speaks during a conference using a meeting application.

[1588] The device records what is said, generates an audio file, and sends it to the server.

[1589] Message Data Collection

[1590] A user sends a message via LINE.

[1591] The device receives the LINE text message and sends it to the server.

[1592] Data Preprocessing

[1593] Converting voice data to text

[1594] The server inputs the received voice file into a voice recognition engine, which converts the voice into text data.

[1595] Text data integration

[1596] The server integrates the converted voice data with the message data obtained from LINE.

[1597] The integrated data is cleansed to remove unnecessary symbols and spaces.

[1598] sentiment analysis

[1599] Sentiment analysis of text data

[1600] The server inputs the cleansed text data into a sentiment analysis engine.

[1601] A sentiment analysis engine calculates an emotion score (e.g., happy, sad, anger) for each text block.

[1602] Data integration

[1603] Timestamp of emotion data

[1604] The server assigns a timestamp to each piece of data and organizes it in chronological order.

[1605] Visualizing Emotion Data

[1606] The server converts the emotional data into a visual format such as graphs and charts, and displays it so that users can intuitively understand the fluctuations in their emotions throughout the day.

[1607] Proposal Generation

[1608] Generate improvement suggestions

[1609] The server analyzes the user's emotional tendencies based on the emotion analysis results.

[1610] The server generates appropriate improvement suggestions for the user based on the analysis results (e.g., "Try taking deep relaxation breaths before a meeting").

[1611] Specific examples

[1612] Example 1:

[1613] A user says during a meeting, "What are your tasks for today?"

[1614] The device records what is said and sends the audio file to the server.

[1615] The server converts the audio file into text data: "What is your task today?"

[1616] A user sends a message on LINE asking, "How did the afternoon meeting go?"

[1617] The terminal sends the message to the server.

[1618] The server combines both sets of text data and feeds it into a sentiment analysis engine.

[1619] The emotion analysis detects "impatience" and generates a suggestion based on that, such as "We recommend you relax before your afternoon meeting."

[1620] In this way, the system of the present invention mainly operates in cooperation with the server, the terminal, and the user, and constitutes a specific means for supporting the user's emotional health management.

[1621] The processing flow will be explained below.

[1622] Step 1:

[1623] A user speaks up during a meeting, for example, "What are our goals for next week?"

[1624] Step 2:

[1625] The device records what the user says and generates an audio file that is temporarily stored on the device.

[1626] Step 3:

[1627] The terminal transmits the recorded audio file to the server via the meeting application.

[1628] Step 4:

[1629] The server inputs the received voice file into a voice recognition engine, which converts the voice data into text data, generating a string of characters such as "What are your goals for next week?"

[1630] Step 5:

[1631] A user sends a message on LINE saying, "What day is the meeting next week?"

[1632] Step 6:

[1633] The device retrieves message data from the LINE application and sends it to the server.

[1634] Step 7:

[1635] The server combines the received LINE message data with the text data converted from the voice. At this time, a cleansing process is performed to unify the data and remove unnecessary symbols and spaces.

[1636] Step 8:

[1637] The server inputs the formatted text data into a sentiment analysis engine, which calculates sentiment scores for the texts "What are your goals for next week?" and "What day of the week is the meeting next week?"

[1638] Step 9:

[1639] The server assigns a timestamp to each block of text based on the sentiment score obtained from the sentiment analysis engine, for example, recording that "What are your goals for next week?" was spoken at 10:00 AM and "What day is our meeting next week?" was spoken at 10:05 AM.

[1640] Step 10:

[1641] The server chronologically organizes the time-stamped emotion data and summarizes the emotional fluctuations by daily activity. The emotion data is then converted into a visual format such as a graph or chart.

[1642] Step 11:

[1643] The server analyzes the user's emotional tendencies based on the emotional data and timestamp data, and obtains analysis results such as "people often feel anxious during morning meetings."

[1644] Step 12:

[1645] Based on the analysis results, the server generates improvement suggestions for the user, such as "I suggest you meditate for five minutes before your morning meeting."

[1646] Step 13:

[1647] The user checks the suggestions from the server on a device such as a smartphone and adjusts their behavior based on the advice.

[1648] Example 1

[1649] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1650] In today's communication environment, there is a lack of methods to properly understand users' emotional fluctuations and stress levels and support their emotional health management. In particular, there is a need for a method that can more accurately assess users' emotional state and provide effective improvement suggestions by integrating and analyzing voice data and text messages.

[1651] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1652] In this invention, the server includes means for collecting voice data, means for converting the collected voice data into text data, means for collecting text messages, means for integrating the voice data and the text messages, means for preprocessing the integrated data, means for sentiment analysis of the preprocessed data, and means for generating improvement suggestions based on the sentiment analysis results. This makes it possible to comprehensively analyze a user's various communication data, appropriately evaluate the emotional state, and provide effective improvement suggestions.

[1653] "Audio data" refers to digital audio files that record the content of user speech occurring during meetings, conversations, and the like.

[1654] "Text data" refers to character information converted from voice data or character information sent and received by a message application.

[1655] A "text message" is text information that a user sends and receives using a text application (e.g., a messaging app).

[1656] "Preprocessing" is the procedure of removing unnecessary information and extracting necessary parts in order to prepare data in a format that can be analyzed.

[1657] "Sentiment analysis" is the process of analyzing a user's emotional state (e.g., joy, sadness, anger) from text data and classifying it into scores or categories.

[1658] "Improvement suggestions" are specific advice on improving behavior or status that is provided to the user based on the results of sentiment analysis.

[1659] "Fusion" refers to combining multiple data sources (e.g., voice data and text messages) into a single dataset.

[1660] A "timestamp" is a means of clarifying time-series information by adding information about the date and time when data was generated or collected.

[1661] "Visualization" is the process of transforming data into a visually understandable format such as a graph or chart.

[1662] The present invention provides a system for collecting voice data and text messages, integrating the collected data to perform sentiment analysis, and generating improvement suggestions for users based on emotional fluctuations. Specific embodiments are described below.

[1663] Hardware and Software Configuration

[1664] 1. Device for collecting audio data:

[1665] A user speaks during a meeting using a meeting application (e.g., an online conference system).

[1666] The terminal records the user's speech during the conference, generates an audio file (e.g., a WAV file), and sends it to the server.

[1667] 2. Devices for Text Message Data Collection:

[1668] A user has a conversation using a messaging application (e.g., a messaging app).

[1669] The terminal retrieves the text message from the message application and sends it to the server as a text file.

[1670] 3. Data analysis server:

[1671] The server uses speech recognition and sentiment analysis engines such as Google Cloud Speech-to-Text and IBM Watson Natural Language Understanding to convert the voice data into text data and perform further sentiment analysis.

[1672] The server includes an engine for generating specific improvement suggestions for the user based on the sentiment analysis results.

[1673] Specific examples

[1674] Example 1:

[1675] A user says, "What are your tasks for today?" during an online meeting.

[1676] The device records what is said and sends the audio file to the server.

[1677] The server inputs the audio file into the Google Cloud Speech-to-Text engine and converts it into text data: "What is your task today?"

[1678] A user sends a message in a messaging app asking, "How did the meeting this afternoon go?"

[1679] The terminal receives the message and sends it to the server.

[1680] The server integrates both sets of text data and performs a cleansing process.

[1681] The server inputs the cleansed text data into IBM Watson Natural Language Understanding for sentiment analysis.

[1682] As a result of the emotion analysis, "impatience" is detected, and based on that, an improvement suggestion is generated, such as "recommending relaxation before the afternoon meeting."

[1683] Examples of prompt statements

[1684] Example prompt sentence:

[1685] "A user might say, 'What are your tasks for today?' in an online meeting, and then later send a message in a messaging app asking, 'How did your afternoon meeting go?' Combine these data points, perform sentiment analysis, and generate appropriate improvement suggestions."

[1686] In this way, the system of the present invention constitutes a concrete means for supporting the emotional health management of users, with the server, terminal, and user working together.

[1687] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1688] Step 1:

[1689] Collection and transmission of voice data

[1690] A user speaks during a conference using an online conference system.

[1691] The terminal records the user's speech and generates an audio file (e.g., WAV format).

[1692] The terminal transmits the generated audio file to the server.

[1693] Input: User's speech in an online conference system

[1694] Data processing: Recording audio data and generating WAV format audio files

[1695] Output: Audio file sent to the server

[1696] Specific operation:

[1697] Execute a script to record speech on the online conference system, generate an audio file, and upload the generated audio file to the server using FTP or HTTP protocol.

[1698] Step 2:

[1699] Collection and transmission of message data

[1700] A user sends a text message using a messaging application.

[1701] The terminal acquires text messages sent and received from the message application and saves them as text files.

[1702] The terminal transmits the obtained text file to the server.

[1703] Input: The user's text message in the Messages application

[1704] Data processing: Obtaining text messages and generating text files in TXT format

[1705] Output: A text file that is sent to the server.

[1706] Specific operation:

[1707] When a new message is detected, a script running in the background of the Messages app extracts its contents, saves the extracted text message in a text file, and uploads it to a server.

[1708] Step 3:

[1709] Speech-to-text

[1710] The server inputs the received audio file into a speech recognition engine (e.g., Google Cloud Speech-to-Text).

[1711] The server converts the audio file into text data using a speech recognition engine.

[1712] The server stores the converted text data in temporary storage.

[1713] Input: Audio file sent to the server

[1714] Data processing: Converting audio files into text using a speech recognition engine

[1715] Output: Text data saved in temporary storage

[1716] Specific operation:

[1717] A Python script is executed on the server to send the audio file to the Google Cloud Speech-to-Text API, and the text data returned by the API is retrieved in JSON format and saved in a database on the server.

[1718] Step 4:

[1719] Text data integration and preprocessing

[1720] The server integrates the text data converted from the voice data with the text file obtained from the messaging app.

[1721] The server cleanses the integrated text data, removing unnecessary symbols and spaces.

[1722] Input: Text data converted from voice, text data obtained from messaging apps

[1723] Data processing: text data integration and cleansing

[1724] Output: Integrated text data after cleansing

[1725] Specific operation:

[1726] The integration process runs an SQL query to combine the voice text and message text, and uses regular expressions (RegEx) on the combined data to remove unnecessary symbols and spaces.

[1727] Step 5:

[1728] Conducting sentiment analysis

[1729] The server inputs the cleansed text data into a sentiment analysis engine (e.g., IBM Watson Natural Language Understanding).

[1730] The server obtains the sentiment score (e.g., happy, sad, anger) for each text block returned by the sentiment analysis engine.

[1731] Input: Text data after cleansing

[1732] Data calculation: Calculating sentiment scores using a sentiment analysis engine

[1733] Output: Text data with sentiment scores

[1734] Specific operation:

[1735] The server makes an API request to send the cleansed text data to the sentiment analysis engine, which analyzes the resulting sentiment scores and stores them in a database.

[1736] Step 6:

[1737] Timestamp and organize emotion data

[1738] The server assigns a timestamp to each piece of emotion data and organizes it in chronological order.

[1739] Input: Text data with sentiment scores

[1740] Data processing: Adding timestamps and organizing data in chronological order

[1741] Output: Emotion data with timestamps and organized in chronological order

[1742] Specific operation:

[1743] The Python Pandas library is used to assign timestamps, and a sorting algorithm is applied to organize the data in chronological order, before storing the results in a database.

[1744] Step 7:

[1745] Visualizing Emotion Data

[1746] The server converts the emotional data into a visual format such as a graph or chart, and displays it so that the user can intuitively understand the fluctuations in their emotions throughout the day.

[1747] Input: Organized emotion data

[1748] Data processing: Converting data into graphs, charts, etc.

[1749] Output: Emotion data displayed in a visually understandable format

[1750] Specific operation:

[1751] We create timeline graphs using libraries such as Matplotlib and Plotly, and display the generated graphs on a web dashboard to help users intuitively understand the fluctuations in sentiment throughout the day.

[1752] Step 8:

[1753] Generate improvement suggestions

[1754] The server analyzes the user's emotional tendency based on the emotion analysis result.

[1755] The server generates specific improvement suggestions for the user from the analysis results.

[1756] Input: Sentiment analysis results

[1757] Data calculation: analyzing emotional trends and generating improvement suggestions

[1758] Output: Improvement suggestions

[1759] Specific operation:

[1760] It uses machine learning models to learn patterns and trends from past data and generate suggestions for future improvements, saving the suggestions in a text file in natural language format and sending them to a module that notifies the user.

[1761] (Application example 1)

[1762] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1763] Conventional emotion analysis systems only collect and analyze voice and text data, making it difficult to utilize the resulting emotion information to detect risks or propose countermeasures. In particular, to improve corporate security and safety, it is necessary to quickly grasp changes in users' emotions and propose appropriate countermeasures. To solve this problem, a system is needed that can detect risks from emotion analysis results and propose countermeasures in a timely manner.

[1764] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1765] In this invention, the server includes means for collecting voice data, means for converting the collected voice data into text data, means for collecting message data, means for integrating the voice data and the message data, means for sentiment analysis of the integrated data, means for generating improvement proposals based on the sentiment analysis results, and means for detecting risks from the collected data and proposing appropriate responses. This makes it possible to grasp the emotional state of the user in real time, detect potential risks early, and propose appropriate measures.

[1766] A "means for collecting voice data" is a device or software that records a user's voice and stores it as digital data.

[1767] The "means for converting collected voice data into text data" refers to a device or software that converts voice data into text using voice recognition technology.

[1768] A "means for collecting message data" is a device or software that obtains text data sent and received from a messaging application.

[1769] "Means for integrating voice data and message data" refers to a device or software that unifies data collected in different formats into one format.

[1770] A "means for sentiment analysis of integrated data" is a device or software that analyzes and evaluates sentiment from collected text data.

[1771] The "means for generating improvement suggestions based on the results of sentiment analysis" is a device or software that provides a specific action plan or advice to the user based on the results of sentiment analysis.

[1772] "Means for detecting risks from collected data and proposing appropriate responses" refers to devices or software that analyze emotional data, identify potential risks, and suggest necessary measures.

[1773] The "means for assigning timestamps and organizing the results of sentiment analysis in chronological order" refers to a device or software that assigns time information to each piece of data and arranges it in chronological order.

[1774] "Means for converting the results of sentiment analysis into a visual format such as a graph or chart" refers to a device or software that visually displays the analysis results so that they can be intuitively understood.

[1775] The "means for warning the user of the occurrence of a risk" refers to a device or software that notifies the user when a risk increases based on the analysis results.

[1776] To implement this invention, a system with the following functions is required. First, as a means for collecting voice data, a microphone on a smartphone or smart glasses is used to record the user's voice in real time. As a means for converting collected voice data into text data, voice recognition technology is used. Google Speech-to-Text is a suitable software.

[1777] Additionally, the "means of collecting message data" involves using APIs to obtain text data from messaging applications used within the company (e.g., Slack, Microsoft Teams). The terminals collecting this data must be appropriate devices connected to the server.

[1778] Next, the different formats of data are unified into one format using a "means for integrating voice data and message data." This allows the data to be analyzed using a "means for sentiment analysis of the integrated data" to calculate an emotional score. For this part, a sentiment analysis engine such as Amazon Comprehend is used.

[1779] Furthermore, the "means for generating improvement proposals based on the results of sentiment analysis" proposes specific measures to users based on the obtained sentiment data. This system is capable of detecting risks from the collected data and proposing appropriate responses in a timely manner.

[1780] The "means of assigning timestamps and organizing the sentiment analysis results in chronological order" involves assigning time information to each piece of data and organizing it in chronological order. This allows for an intuitive understanding of sentiment fluctuations. The "means of converting the sentiment analysis results into visual formats such as graphs and charts and alerting users to emerging risks" involves visually displaying them using D3.js and Grafana.

[1781] As a specific example, audio spoken by a user during a meeting is collected using a smartphone microphone and sent to a server. The server then uses Google Speech-to-Text to convert the audio into text data. In parallel, text data sent and received by the user via a messaging application is also sent to the server, and both sets of data are integrated. Sentiment analysis is then performed using Amazon Comprehend, and risks are detected based on the results. For example, specific improvement suggestions are displayed, such as, "A high stress level has been detected from comments made during the meeting. We recommend that you take a break to relax."

[1782] An example prompt is, "Generate an application that uses Google Speech-to-Text to convert employees' real-time voice data into text data, perform sentiment analysis, detect security risks early, and generate improvement suggestions."

[1783] In this way, the entire system works together to grasp the user's emotional state in real time, detect potential risks early, and propose appropriate countermeasures.

[1784] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1785] Step 1:

[1786] Audio data collection

[1787] Users communicate using the microphone on their smartphones or smart glasses. The device collects voice data in real time through the microphone and stores it in a digital file format. The input is the user's voice, and the output is a voice data file.

[1788] Step 2:

[1789] Converting audio data to text

[1790] The device sends the collected voice data to the server, which then uses a voice recognition engine (e.g., Google Speech-to-Text) to convert the voice data into text data. The input is a voice data file, and the output is text data.

[1791] Step 3:

[1792] Message Data Collection

[1793] The device acquires text data from messaging applications used within the company (e.g., Slack, Microsoft Teams), and uses an API to send the message data to the server. The input is the text message from the messaging application, and the output is the collected message data.

[1794] Step 4:

[1795] Data integration

[1796] The server integrates the text data converted from the voice with the message data obtained from the messaging application. Data integration is a process for unifying data of different formats into a single format. The input is the text data from the voice and the message data, and the output is the integrated text data.

[1797] Step 5:

[1798] sentiment analysis

[1799] The server inputs the integrated text data into a sentiment analysis engine (e.g., Amazon Comprehend). The sentiment analysis engine calculates an emotional score (e.g., happy, sad, or angry) from the text data. The input is the integrated text data, and the output is the emotional score.

[1800] Step 6:

[1801] Risk detection and response proposals

[1802] The server runs an algorithm to detect potential risks from the sentiment analysis results. If a risk is detected, the server generates a response suggestion (e.g., "High stress levels detected. We recommend taking a break to relax."). The input is the sentiment score, and the output is the risk detection result and a response suggestion.

[1803] Step 7:

[1804] Organizing data chronologically

[1805] The server assigns a timestamp to each voice and message data and organizes the emotion analysis results in chronological order, making it easier to understand emotion fluctuations throughout the day. The input is the emotion score and the original data, and the output is the time-stamped data.

[1806] Step 8:

[1807] Visualization of sentiment analysis results

[1808] The server converts the sentiment analysis results into visual formats such as graphs and charts and displays them to the user. This process uses visualization tools such as D3.js and Grafana. The input is time-stamped sentiment data, and the output is visualized graphs and charts.

[1809] Step 9:

[1810] Risk warning notification

[1811] If a risk is detected, the server sends a notification to the user to warn them of the risk. This notification can be in the form of a push notification to a smartphone or other device. The input is the risk detection result, and the output is a warning notification to the user.

[1812] By following these steps, users can understand their own emotional state in real time, detect potential risks early, and take appropriate measures.

[1813] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1814] This invention combines a system that collects voice data and message data, integrates the data, performs emotion analysis, and generates improvement proposals, with an emotion engine that recognizes the user's emotions. To realize this system, hardware and software with the following functions are required.

[1815] System configuration

[1816] 1. Device for collecting voice data

[1817] A terminal that uses a meeting application or the like to record user speech.

[1818] 2. Device for collecting message data

[1819] A device that retrieves text messages from a messaging application (e.g., LINE).

[1820] 3. Data analysis server

[1821] A speech recognition engine for converting collected voice data into text data.

[1822] An analysis engine that integrates message data and text data converted from voice to perform sentiment analysis.

[1823] Processing to recognize user emotions using an emotion engine.

[1824] An engine that generates specific improvement suggestions for users based on the results of sentiment analysis.

[1825] System operation procedure

[1826] Data collection

[1827] Audio data collection

[1828] A user speaks during a meeting using a meeting application, for example, "How are we progressing with this week's tasks?"

[1829] The device records the user's speech and generates an audio file.

[1830] Message Data Collection

[1831] A user sends a message on LINE asking, "When is the next meeting?"

[1832] The device receives the LINE text message.

[1833] Data Preprocessing

[1834] Converting voice data to text

[1835] The server receives the audio file and inputs it into a speech recognition engine, which converts the audio data into text data. For example, it generates text data such as "How will you proceed with this week's tasks?"

[1836] Text data integration

[1837] The server combines the text data obtained from LINE with the text data converted from the voice data. The data is then formatted into a unified format and a cleansing process is performed to remove unnecessary symbols and spaces.

[1838] sentiment analysis

[1839] Sentiment analysis of text data

[1840] The server inputs the formatted text data into the emotion engine, which directly recognizes the user's emotions from the text "How will you proceed with your tasks this week?" and "When is the next meeting?" and calculates an emotion score (e.g., joy, sadness, impatience, etc.).

[1841] Data integration

[1842] Timestamp of emotion data

[1843] The server assigns a timestamp to each piece of data and organizes it in chronological order. For example, it records that "How are we progressing with our tasks this week?" was said at 9:00 AM and "When is our next meeting?" was said at 9:10 AM.

[1844] Visualizing Emotion Data

[1845] The server converts the emotional data into a visual format such as graphs and charts, and displays it so that users can intuitively understand the fluctuations in their emotions throughout the day.

[1846] Proposal Generation

[1847] Generate improvement suggestions

[1848] The server analyzes the user's emotional tendencies based on the results of the emotion analysis, and obtains an analysis result such as "people often feel anxious during morning meetings."

[1849] The server generates appropriate improvement suggestions for the user based on the analysis results (e.g., "I suggest you meditate for five minutes before your morning meeting").

[1850] Specific examples

[1851] Example 1:

[1852] A user says during a meeting, "How are we progressing with our tasks this week?"

[1853] The device records what is said and sends the audio file to the server.

[1854] The server converts the audio file into text data such as "How will you proceed with your tasks this week?"

[1855] A user sends a message on LINE asking, "When is the next meeting?"

[1856] The terminal sends the message to the server.

[1857] The server combines both sets of text data and inputs them into the emotion engine.

[1858] The emotion engine recognizes the emotion "impatient" for "How are you progressing with your tasks this week?" and "When is your next meeting?"

[1859] Based on the sentiment score, the server generates a suggestion such as "Try taking some deep breaths to relax before your morning meeting."

[1860] In this way, the system of the present invention operates in cooperation with the server, the terminal, and the user, and in particular, combines the emotion engine to form a specific means for supporting the user's emotional health management.

[1861] The processing flow will be explained below.

[1862] Step 1:

[1863] A user says during a meeting, "What are our goals for next week?"

[1864] Step 2:

[1865] The device records what the user says and generates an audio file that is temporarily stored on the device.

[1866] Step 3:

[1867] The terminal transmits the recorded audio file to the server via the meeting application.

[1868] Step 4:

[1869] The server inputs the received voice file into a speech recognition engine, which converts the voice data into text data and generates the string "What are your goals for next week?"

[1870] Step 5:

[1871] A user sends a message on LINE saying, "What day is the meeting next week?"

[1872] Step 6:

[1873] The device retrieves message data from the LINE application and sends it to the server.

[1874] Step 7:

[1875] The server combines the received LINE message data with the text data converted from the voice. At this time, a cleansing process is performed to unify the data and remove unnecessary symbols and spaces.

[1876] Step 8:

[1877] The server inputs the formatted text data into the emotion engine, which calculates emotion scores for the texts "What are your goals for next week?" and "What day of the week is the meeting next week?"

[1878] Step 9:

[1879] The server recognizes the emotion of each text based on the emotion score obtained from the emotion engine. For example, "What are your goals for next week?" will get an emotion score of "Interest," and "What day of the week is the meeting next week?" will get an emotion score of "Impatience."

[1880] Step 10:

[1881] The server assigns a timestamp to the sentiment score and organizes the data chronologically, for example, recording that "What are your goals for next week?" was said at 10:00 AM and "What day is the meeting next week?" at 10:05 AM.

[1882] Step 11:

[1883] The server converts the time-stamped emotion data into visual formats such as graphs and charts, allowing users to intuitively understand the fluctuations of emotions throughout the day.

[1884] Step 12:

[1885] The server analyzes the user's emotional tendencies based on the emotional data and timestamp data. For example, it determines whether the user frequently feels impatient during certain times of the day.

[1886] Step 13:

[1887] The server generates improvement suggestions for the user based on the analysis results, such as "Try taking deep breaths to relax before your morning meeting."

[1888] Step 14:

[1889] The user checks the suggestions from the server on a device such as a smartphone and adjusts their behavior based on the advice.

[1890] Example 2

[1891] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1892] Conventional systems collect and analyze voice data and message data separately, making it difficult to comprehensively and accurately recognize user emotions. Furthermore, the collected data is often insufficiently cleansed, hindering accurate emotion analysis. Furthermore, it is difficult to provide specific improvement suggestions to users based on the results of emotion analysis, and there is a lack of a way to visually understand daily emotional fluctuations.

[1893] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1894] In this invention, the server includes means for collecting voice data, means for converting the collected voice data into text data, means for collecting message data, means for integrating the voice data and the message data, means for sentiment analysis of the integrated data, means for generating improvement suggestions based on the sentiment analysis results, means for data cleansing to remove unnecessary symbols and spaces from each data, means for generating suggestions to support the user's emotional health management based on the sentiment analysis results, and means for organizing the emotional data in timestamp order. This enables the integrated collection and analysis of voice and message data to accurately recognize the user's emotions. Furthermore, the cleansing process improves data accuracy, and appropriate improvement suggestions can be provided to support the user's emotional health. Organizing the emotional data in timestamp order allows for a visual understanding of emotional fluctuations.

[1895] A "means for collecting voice data" is a device or application that records a user's speech or voice and stores it as a digital audio file.

[1896] The "means for converting collected voice data into text data" refers to a voice recognition engine or software for analyzing voice files and converting their contents into text data in sentence format.

[1897] A "means for collecting message data" is a device or application that captures and stores text messages or chat messages.

[1898] A "means for integrating voice and message data" is software or a process for consolidating data collected from different formats and sources into a single format and managing it in a unified manner.

[1899] "Means for sentiment analysis of the integrated data" refers to an emotion engine or software for analyzing the integrated text data and determining a user's emotion based on the content of the text.

[1900] The "means for generating improvement suggestions based on the results of sentiment analysis" is software or an engine for generating feedback and advice for users based on the results of sentiment analysis.

[1901] A "data cleansing method that removes unnecessary symbols and spaces from each piece of data" is a process or software that automatically removes unnecessary symbols and spaces from collected data to improve the quality of the data.

[1902] The "means for generating suggestions to support the user's emotional health management based on the results of sentiment analysis" is software or an engine for utilizing the results of sentiment analysis to present specific advice and schedules for improving the user's emotional health.

[1903] "Means for organizing emotion data in timestamp order" refers to software or a process for adding time information to collected emotion data and organizing and storing the data in chronological order.

[1904] The present invention is a system that collects voice data and message data, integrates the data, performs sentiment analysis, and generates improvement proposals. This system requires hardware and software with the functions of voice data collection, message data collection, data integration, sentiment analysis, data organization and visualization, and generation of improvement proposals.

[1905] Audio data collection

[1906] Users use a meeting application (e.g., a video conferencing application) to hold a conversation. A device (e.g., a user's PC or smartphone) records the audio during the meeting and generates an audio file. This audio file is automatically uploaded to cloud storage.

[1907] Message Data Collection

[1908] A user sends a message using a messaging application (e.g., a text messaging app), and the device receives the text message and sends it to a server via a dedicated application.

[1909] Converting audio data to text

[1910] The server receives the audio file from the cloud storage. The received audio file is input into a speech recognition engine (e.g., a speech recognition API) and converted into text data. For example, a statement such as "How will you proceed with this week's tasks?" is generated as text data.

[1911] Text data integration

[1912] The server combines the text data obtained from LINE and other messaging apps with the text data converted from voice. During the combination process, a data cleansing process is performed to remove unnecessary symbols and spaces from each data, improving the accuracy of the data.

[1913] sentiment analysis

[1914] The server inputs the cleansed text data into an emotion engine (e.g., a natural language processing API). The emotion engine recognizes the user's emotion for each piece of text and calculates an emotion score. Specifically, for questions like "How will you progress with your tasks this week?" and "When is the next meeting?", emotions such as impatience and anticipation are recognized.

[1915] Timestamp of emotion data

[1916] The server assigns a timestamp to each piece of text data and organizes it in chronological order. For example, it records that "How are we progressing with this week's tasks?" was said at 9:00 AM and "When is the next meeting?" was said at 9:10 AM.

[1917] Visualizing Emotion Data

[1918] The server converts the emotion data into a visual format such as a graph or chart, allowing users to intuitively understand their emotional fluctuations throughout the day. For example, a line graph showing the emotional fluctuations throughout the day can be generated and displayed on the user interface.

[1919] Generate improvement suggestions

[1920] The server analyzes the user's emotional tendencies based on the results of the emotion analysis. It then generates specific improvement suggestions based on the analysis results and notifies the user through the user interface. For example, a suggestion such as "Try taking deep breaths to relax before your morning meeting" may be displayed.

[1921] Specific examples

[1922] A user says during a meeting, "How are we progressing with our tasks this week?"

[1923] The device records what is said and sends the audio file to the server.

[1924] The server converts the audio file into text data such as "How will you proceed with your tasks this week?"

[1925] A user sends a message in a messaging app asking, "When is our next meeting?"

[1926] The terminal sends the message to the server.

[1927] The server combines both sets of text data and inputs them into the emotion engine.

[1928] The emotion engine recognizes "impatience" for "How are you progressing with your tasks this week?" and "When is your next meeting?"

[1929] Based on the sentiment score, the server generates a suggestion such as "Try taking some deep breaths to relax before your morning meeting."

[1930] In this way, the system of the present invention operates in cooperation with the server, the terminal, and the user, and in particular, combines the emotion engine to form a specific means for supporting the user's emotional health management.

[1931] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1932] Step 1:

[1933] Audio data collection

[1934] A user has a conversation using a meeting application (e.g., a video conferencing application). The device records the audio during the meeting and sends it to cloud storage in the form of audio data.

[1935] Input: User's voice

[1936] Output: Audio files saved in cloud storage

[1937] Specifically, when a user says, "How will you proceed with this week's tasks?", the device uploads the voice recording to cloud storage as a digital audio file.

[1938] Step 2:

[1939] Message Data Collection

[1940] A user sends a message using a messaging application (e.g., a text messaging app), and the device sends the text message to a server via a dedicated application.

[1941] Input: User's text message

[1942] Output: Text data stored on the server

[1943] Specifically, when a user sends a message such as "When is the next meeting?", the terminal sends the text message to the server.

[1944] Step 3:

[1945] Converting audio data to text

[1946] The server receives the audio file from the cloud storage and inputs it into a speech recognition engine, which analyzes the audio file and converts it into text data.

[1947] Input: Audio files stored in cloud storage

[1948] Output: Text data converted from audio

[1949] For example, the speech recognition engine generates text data such as "How will you proceed with this week's tasks?"

[1950] Step 4:

[1951] Text data integration

[1952] The server receives and integrates text data obtained from LINE and other messaging apps and text data converted from voice.

[1953] Input: Text data converted from speech and text data from messaging apps

[1954] Output: Integrated text data

[1955] The server performs data cleansing processing, removing unnecessary symbols and spaces, and formatting the data into a unified format, thereby improving the quality of the data.

[1956] Step 5:

[1957] sentiment analysis

[1958] The server inputs the cleansed text data into the emotion engine, which recognizes the user's emotion for each text and calculates an emotion score.

[1959] Input: Cleansed text data

[1960] Output: Sentiment analysis results for each text

[1961] For example, the emotion engine recognizes "impatience" in response to questions such as "How will you progress with your tasks this week?" and "When is the next meeting?"

[1962] Step 6:

[1963] Timestamp of emotion data

[1964] The server assigns a timestamp to each piece of text data and organizes the data in chronological order.

[1965] Input: Sentiment analysis results

[1966] Output: Emotion data with timestamps

[1967] For example, it records that "How will you proceed with this week's tasks?" was said at 9:00 AM, and "When is the next meeting?" was said at 9:10 AM.

[1968] Step 7:

[1969] Visualizing Emotion Data

[1970] The server converts the emotion data into visual formats such as graphs and charts, allowing users to intuitively understand their emotional fluctuations throughout the day.

[1971] Input: Emotion data with timestamps

[1972] Output: Visualized emotional fluctuation data

[1973] Specifically, it generates a line graph showing emotional fluctuations and displays it on the user interface.

[1974] Step 8:

[1975] Generate improvement suggestions

[1976] The server analyzes the user's emotional tendencies based on the results of the emotion analysis, generates specific improvement suggestions based on the analysis results, and notifies the user through the user interface.

[1977] Input: Sentiment analysis results and trends

[1978] Output: Suggested improvements to the user

[1979] As a specific action, the suggestion is to "Try taking some deep breaths to relax before your morning meeting."

[1980] Through the above processing steps, the system can integrate voice data and text messages, perform sentiment analysis, and provide appropriate improvement suggestions to users, thereby effectively supporting their emotional health management.

[1981] (Application example 2)

[1982] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1983] Autonomous vehicles require real-time monitoring of passengers' emotional states and the provision of appropriate support based on those emotions. In particular, a system that can automatically generate and implement specific suggestions and actions to reduce passenger stress and provide a comfortable travel environment is required.

[1984] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1985] In this invention, the server includes means for collecting voice data, means for converting the collected voice data into text data, means for collecting message data, means for integrating the voice data and the message data, means for sentiment analysis of the integrated data, means for generating improvement suggestions based on the sentiment analysis results, and means for collecting passenger voice data and message data, conducting sentiment analysis, and generating stress reduction suggestions. This makes it possible to grasp the passenger's real-time emotional state and generate and provide appropriate stress reduction suggestions according to the situation.

[1986] "Voice data" refers to data in which the user's voice is recorded in digital form.

[1987] "Text data" refers to character information data obtained by converting voice data into text.

[1988] "Message data" refers to data of text messages sent and received by users.

[1989] "Merge" is the process of combining the text data converted from the voice data and the message data into a single data set.

[1990] "Sentiment analysis" is the process of analyzing a user's emotional state (e.g., joy, sadness, impatience, etc.) from text data.

[1991] "Improvement suggestions" are specific actions or advice suggested to the user based on the results of sentiment analysis.

[1992] "Passenger" refers to a person using an autonomous vehicle.

[1993] "Stress reduction suggestions" are suggestions for reducing stress, taking into account the emotional state of passengers.

[1994] The present invention provides a system for collecting and integrating voice and message data, performing sentiment analysis, and generating stress reduction suggestions based on the results, which is particularly applicable to passenger emotion monitoring and stress reduction assistance for autonomous vehicles.

[1995] System configuration

[1996] To realize this system, the following hardware and software are required.

[1997] 1. Terminal

[1998] Microphones placed inside the vehicle: collect passenger voice data.

[1999] In-car infotainment system: Collects passenger text messages through messaging applications.

[2000] 2. Server

[2001] Speech recognition engine: Converts collected voice data into text data.

[2002] Data integration engine: Integrates text data converted from voice data with message data.

[2003] Sentiment Engine: Performs sentiment analysis using the integrated data.

[2004] Suggestion generation engine: Generates stress reduction suggestions based on the sentiment analysis results.

[2005] System operation procedure

[2006] Audio data collection

[2007] When passengers speak in the car, the microphones collect their words and generate an audio file. For example, if they say, "What should we do about our next meeting?", that will be recorded.

[2008] Message Data Collection

[2009] Text data is collected when a passenger sends a message using a messaging app connected to the vehicle's infotainment system. For example, if a passenger sends a message in a messaging app saying, "When is our next meeting?", that message is captured.

[2010] Data Preprocessing

[2011] The server receives the collected voice files and converts them into text data using a speech recognition engine. For example, "What should we do about the next meeting?" is converted into text "What should we do about the next meeting?"

[2012] The server combines the text data obtained from the messaging application with the text data converted from the voice data, and after a data cleansing process, it is formatted into a unified format.

[2013] sentiment analysis

[2014] The server inputs the formatted text data into the emotion engine, which calculates an emotion score (e.g., impatience, tension, etc.) based on passenger comments such as "What should we do about the next meeting?" and "When is the next meeting?"

[2015] Proposal Generation

[2016] The server generates stress reduction suggestions based on the emotion analysis results. For example, if the emotion engine recognizes "impatience," it generates a suggestion such as "Would you like to play some relaxing music?"

[2017] Suggestions are communicated to passengers through the vehicle's infotainment system, which will either play music automatically or allow passengers to select the suggestion.

[2018] Specific examples

[2019] If a passenger says, "What should we do about our next meeting?", the server converts the speech to text and analyzes that text with an emotion engine. If the emotion engine recognizes the emotion "tension," it generates a suggestion: "Take a deep breath and relax."

[2020] Similarly, if a passenger sends a message in a messaging app asking, "When is my next meeting?", sentiment analysis will be performed and stress-reducing suggestions will be generated.

[2021] Prompt Sentence Examples

[2022] The following are examples of prompts to input into the generative AI model:

[2023] Collect passenger comments, analyze their emotions, and generate stress-reducing suggestions. Below is an example of a passenger comment and the results of the emotion analysis.

[2024] Say: "What about our next meeting?"

[2025] Sentiment Analysis: Negative

[2026] In these cases, generate an appropriate suggestion, for example, "Would you like to play some relaxing music?"

[2027] This system can monitor the emotional state of passengers in autonomous vehicles in real time and provide appropriate stress reduction suggestions, allowing passengers to enjoy a more comfortable and safer travel environment.

[2028] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2029] Step 1:

[2030] Audio data collection

[2031] Input: User's spoken utterance.

[2032] The device uses the in-car microphone to collect the user's speech in real time and record it as an audio file.

[2033] Output: Collected audio data (audio files).

[2034] Step 2:

[2035] Message Data Collection

[2036] Input: A text message that a user sends using a messaging application.

[2037] The device works with the car's infotainment system to retrieve text messages from messaging applications.

[2038] Output: Collected message data (text messages).

[2039] Step 3:

[2040] Converting voice data to text

[2041] Input: Audio data (audio file).

[2042] The server uses a speech recognition engine to convert the voice data into text data.

[2043] Specifically, a speech recognition engine analyzes the collected audio files and generates corresponding text.

[2044] Output: Text data converted from audio.

[2045] Step 4:

[2046] Data integration

[2047] Input: Text data converted from audio and message data.

[2048] The server integrates the text data converted from the voice with the message data and performs a cleansing process.

[2049] The cleansing process is a process of arranging data into a unified format and removing unnecessary symbols and spaces.

[2050] Output: Consolidated clean text data.

[2051] Step 5:

[2052] sentiment analysis

[2053] Input: Integrated text data.

[2054] The server uses an emotion engine to analyze the integrated text data and calculate an emotion score.

[2055] The emotion engine identifies an emotional state, such as "impatience" or "tension," based on the text data.

[2056] Output: Emotion score and emotional state identification results.

[2057] Step 6:

[2058] Generate improvement suggestions

[2059] Input: Emotion scores and emotional state identification results.

[2060] The server generates stress reduction suggestions based on the results of the sentiment analysis.

[2061] A specific suggestion generation engine generates specific actions such as "play relaxing music" or "encourage deep breathing."

[2062] Output: Specific stress reduction suggestions.

[2063] Step 7:

[2064] Proposal Notification

[2065] Enter: stress reduction suggestions.

[2066] The device will then notify passengers of stress-reducing suggestions through the vehicle's infotainment system.

[2067] The passenger can select a suggestion or an automatic action will be initiated to implement the suggestion.

[2068] Output: Passenger notification and action taken.

[2069] Through these processing steps, the system can monitor the user's emotional state in real time and provide appropriate stress reduction suggestions, allowing passengers to enjoy a more comfortable travel environment.

[2070] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2071] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2072] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2073] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2074] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2075] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2076] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2077] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2078] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2079] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2080] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2081] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2082] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2083] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2084] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2085] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2086] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2087] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2088] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[2089] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[2090] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[2091] The following is further disclosed regarding the above embodiment.

[2092] (Claim 1)

[2093] means for collecting audio data;

[2094] A means for converting the collected voice data into text data;

[2095] a means for collecting message data;

[2096] means for integrating the voice data and the message data;

[2097] a means for sentiment analysis of the integrated data;

[2098] means for generating improvement suggestions based on the sentiment analysis results;

[2099] A system including:

[2100] (Claim 2)

[2101] 10. The system of claim 1, further comprising means for assigning a timestamp to each piece of voice data and message data, and for chronologically organizing the sentiment analysis results.

[2102] (Claim 3)

[2103] 10. The system of claim 1, further comprising means for converting the sentiment analysis results into a visual format such as a graph or chart.

[2104] "Example 1"

[2105] (Claim 1)

[2106] means for collecting audio data;

[2107] A means for converting the collected voice data into text data;

[2108] a means for collecting text messages;

[2109] a means for integrating voice data with text messages;

[2110] a means for preprocessing the integrated data;

[2111] means for sentiment analyzing the preprocessed data;

[2112] means for generating improvement suggestions based on the sentiment analysis results;

[2113] A system including:

[2114] (Claim 2)

[2115] 10. The system of claim 1, further comprising means for assigning a timestamp to each voice data and text message and for chronologically organizing the sentiment analysis results.

[2116] (Claim 3)

[2117] 10. The system of claim 1, further comprising means for converting the sentiment analysis results into a visual format such as a graph or chart.

[2118] "Application Example 1"

[2119] (Claim 1)

[2120] means for collecting audio data;

[2121] A means for converting the collected voice data into text data;

[2122] a means for collecting message data;

[2123] means for integrating the voice data and the message data;

[2124] a means for sentiment analysis of the integrated data;

[2125] means for generating improvement suggestions based on the sentiment analysis results;

[2126] A means of detecting risks from collected data and proposing appropriate responses;

[2127] A system including:

[2128] (Claim 2)

[2129] 10. The system of claim 1, further comprising means for assigning a timestamp to each piece of voice data and message data, and for chronologically organizing the sentiment analysis results.

[2130] (Claim 3)

[2131] 10. The system of claim 1, further comprising means for converting the sentiment analysis results into a visual format such as a graph or chart and alerting a user to an emerging risk.

[2132] "Example 2: Combining Emotion Engines"

[2133] (Claim 1)

[2134] means for collecting audio data;

[2135] A means for converting the collected voice data into text data;

[2136] a means for collecting message data;

[2137] means for integrating the voice data and the message data;

[2138] a means for sentiment analysis of the integrated data;

[2139] means for generating improvement suggestions based on the sentiment analysis results;

[2140] A data cleansing method for removing unnecessary symbols and spaces from each data item;

[2141] means for generating suggestions to support the user's emotional well-being management based on the sentiment analysis results;

[2142] A means for organizing emotion data in timestamp order;

[2143] A system including:

[2144] (Claim 2)

[2145] 10. The system of claim 1, further comprising means for assigning a timestamp to each piece of voice data and message data, and for chronologically organizing the sentiment analysis results.

[2146] (Claim 3)

[2147] 10. The system of claim 1, further comprising means for converting the sentiment analysis results into a visual format such as a graph or chart.

[2148] "Application example 2 when combining emotion engines"

[2149] (Claim 1)

[2150] means for collecting audio data;

[2151] A means for converting the collected voice data into text data;

[2152] a means for collecting message data;

[2153] means for integrating the voice data and the message data;

[2154] a means for sentiment analysis of the integrated data;

[2155] means for generating improvement suggestions based on the sentiment analysis results;

[2156] A means for collecting passenger voice data and message data, and generating stress reduction suggestions after performing sentiment analysis;

[2157] A system including:

[2158] (Claim 2)

[2159] 10. The system of claim 1, further comprising means for assigning a timestamp to each piece of voice data and message data, and for chronologically organizing the sentiment analysis results.

[2160] (Claim 3)

[2161] 10. The system of claim 1, further comprising means for converting the sentiment analysis results into a visual format such as a graph or chart, and means for providing real-time stress reduction suggestions based on the passenger's emotional state. [Explanation of symbols]

[2162] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for collecting audio data; A means for converting the collected voice data into text data; a means for collecting message data; means for integrating the voice data and the message data; a means for sentiment analysis of the integrated data; means for generating improvement suggestions based on the sentiment analysis results; A system including:

2. 2. The system according to claim 1, further comprising means for assigning a timestamp to each piece of voice data and message data, and for chronologically organizing the sentiment analysis results.

3. The system of claim 1 , further comprising means for converting the sentiment analysis results into a visual format such as a graph or chart.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A